CellType

The agentic drug company. We simulate human biology.

Hiring — 2 openYC-W26Healthcare -> Drug Discovery and DeliveryEarly

What CellType does

CellType is building an agentic drug company. AI agents that run the full drug discovery pipeline on top of biological foundation models that simulate human biology. Our core technology, developed with Google DeepMind, has already discovered and validated a new cancer treatment signal. We’re working with Top 10 pharma.

2 open roles

Founding Platform Engineer, Data & ML Systems
New York, NY, USFull-time3+ years$145K - $250K0.50% - 2.00% equityVisa: US citizen/visa only
What the role involves

About CellType CellType is building foundation models and agent systems for biology. We believe the next major advances in biotech AI will come from rich biological data, strong model systems, and reliable infrastructure working together. We work with pharma and biotech partners on problems such as preclinical-to-clinical translation, response prediction, biomarker discovery, and scientific reasoning across complex biological datasets. We are building the core intelligence layer for biology, and that requires a world-class data and ML platform. About the role We are hiring a Founding Platform Engineer to build the infrastructure backbone behind our training, evaluation, and inference stack. We are looking for someone who can build the systems that make biological data usable for model development at speed and at scale: ingestion, indexing, search, retrieval, dataset interfaces, reproducibility, validation, orchestration, observability, and distributed performance. You will work on the full path from raw data to training-ready datasets to reliable production workflows. The right person will make it dramatically easier for the rest of the team to build, evaluate, and ship models. What you'll do Build and maintain data infrastructure for model training, evaluation, and inference Design and scale high-performance inference serving systems for biological foundation models Design standardized dataset interfaces so biological data is consistent, discoverable, and easy to use across the team Build ingestion and processing pipelines for public, proprietary, and customer datasets Build indexing, search, and retrieval systems that make large datasets queryable and useful in practice Establish safeguards and validation systems so datasets are reproducible, versioned, and trustworthy once standardized Improve throughput, latency, and reliability of distributed data loading and ML pipelines Profile and eliminate performance bottlenecks across GPU, networking, and storage layers Automate fault detection and recovery for serving and training systems Build internal tools for dataset inspection, debugging, quality control, and operational visibility Partner closely with ML engineers and researchers so the platform fits real workflows rather than abstract platform ideals Help define how we handle permissions, privacy, compliance boundaries, and operational rigor for sensitive biological and customer data You may be a fit if you Have deep experience in backend, infrastructure, distributed systems, or data platform engineering Have built scalable data pipelines or stateful distributed systems in production Have experience building or operating large-scale inference or training systems Have a deep understanding of GPU execution constraints, memory trade-offs, and data-loading bottlenecks around training workloads Have experience with dataset infrastructure for large-scale ML systems, training pipelines, or inference-adjacent systems Have worked with multimodal or very large datasets that cannot simply fit in memory Have hands-on experience with data indexing, search, or retrieval infrastructure, and understand how to make large datasets discoverable, queryable, and usable in practice Can reason about system-level trade-offs between latency, throughput, and cost Have experience working with privacy-sensitive or compliance-sensitive data systems Have built internal developer tools for ML or data teams Have a track record of owning critical production infrastructure Are comfortable designing APIs, modular abstractions, and internal platform interfaces with strong attention to user experience Have strong instincts around reliability, reproducibility, and operational simplicity Are comfortable with cloud infrastructure, containers, Kubernetes, Infrastructure-as-Code, CI/CD, and observability Produce maintainable code and make pragmatic architecture decisions under time pressure Thrive in a small team where ownership is broad and priorities can change

Founding Research Engineer, Model Training
New York, NY, USFull-time3+ years$150K - $250K0.50% - 2.00% equityVisa: Will sponsor
What the role involves

Founding Research Engineer, Model Training Location: New York City Type: Full-time About CellType CellType is building foundation models and agent systems for biology. We believe the next major advances in biotech AI will come from models trained to reason over biological data, experiments, and translational outcomes, not from lightweight wrappers around generic models. We work with pharma and biotech partners on problems such as preclinical-to-clinical translation, response prediction, biomarker discovery, and scientific reasoning across complex biological datasets. Our core technology was originally developed at Yale in collaboration with Google DeepMind, and has been published at top ML venues including ICML. We are building the core intelligence layer for biology. About the role We are hiring a Founding Research Engineer to build and scale the systems that improve our models. This role sits at the boundary of research and engineering. You will work on training, post-training, evaluation, performance optimization, and the systems needed to support all of that. You should be excited by both novel model development and the operational reality of making training systems run reliably. What you'll do Build and improve training and post-training systems for biological foundation models and agentic model workflows Design and run experiments across supervised fine-tuning, reinforcement learning, tool use, evaluation, and model behavior optimization Build and maintain distributed RL and post-training infrastructure Improve reliability of rollout, evaluation, and reward pipelines Own critical parts of the model training stack, including performance, reliability, observability, and debugging Investigate and resolve issues across the full stack, from training dynamics and evaluation infrastructure to distributed systems and hardware bottlenecks Profile and eliminate performance bottlenecks across GPU, networking, and storage layers Build clean abstractions for experiments, model evaluation, and distributed training workflows Improve training efficiency, stability, and throughput Work closely with founders and domain experts to translate biological problems into model tasks, environments, and evaluation frameworks Help turn research improvements into real product and customer advantage You may be a fit if you Have hands-on experience training or materially improving serious LLM or generative ML systems Have strong software engineering and distributed systems fundamentals Have deep experience with Python and modern ML frameworks such as PyTorch, JAX, or equivalent systems Have experience with reinforcement learning or post-training methods Have built evaluation systems for tool-using or open-ended models Have a deep understanding of GPU execution constraints and memory trade-offs Have experience debugging performance issues in production ML systems Can reason about system-level trade-offs between latency, throughput, and cost Have a track record of owning critical production infrastructure Can balance research exploration with engineering implementation Have experience with distributed systems, large-scale training, or performance-sensitive ML workloads Care about code quality, testing, performance, and maintainability Are comfortable in a small team where priorities move toward whatever is most important Communicate clearly and collaborate well under both normal and high-pressure conditions Want broad ownership rather than a narrow role boundary This role will directly shape the quality and speed of CellType's core model systems. The right person will help determine not only how good our models become, but how fast we can improve them and how confidently we can deploy them. If you want to work on difficult model problems with real scientific and commercial consequences, we'd love to talk.

Distributed SystemsMachine LearningReinforcement learning (RL)

Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.

Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.

Questions and experiences

Nobody has asked anything about CellType yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.

Reviewed before it appears. Do not include anything that identifies you or anyone else.

Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.

jobo is a browser extension. Open this on a computer to install it.