
Cedana
Fast, reliable, reproducible AI with GPU live migration
What Cedana does
Cedana (YC S23) brings hyperscaler and frontier-lab orchestration capabilities for AI workflows. Our core capability is live migration for CPUs and GPUs workloads. This increases cost savings up to 80%, accelerates time to first token 2-10x, and enables stateful reliability of training jobs even through catastrophic GPU failures. We've integrated our solution into K8s, and support Kueue and Slurm for training distributed jobs, and Kserve for serving inference. OpenAI, Meta and Microsoft have flavors of these capabilities internally and we’re bringing them to everyone. Our vision is to transform cloud compute into a real-time, arbitraged commodity. https://www.cedana.ai
2 open roles
What the role involves
Introducing Cedana The Problem AI and HPC infrastructure suffers from scarcity and high costs, so when failures happen they are costly in terms of time and money. Cluster productivity directly determines research output and revenue. Achieving high utilization and throughput is increasingly challenging due to the complexity of workloads, hardware, and operations. Cedana’s Solution Cedana maximizes AI+HPC cluster utilization and reliability with automated GPU checkpointing infrastructure. We enable transparent and fast migration of GPU workloads across instances, without losing work. Workloads automatically migrate to achieve new levels of reliability and throughput while accelerating time to results. Our system is at the kernel/OS level, requiring no code or config changes, and works seamlessly with Kubernetes, SLURM, and NVIDIA Dynamo. Today, we're deploying into leading inference platforms, neoclouds, enterprise, and research clusters. The Team Cedana's founding team has spent over a decade making computation run fast, productively, and reliably for AI. Our research appears in NeurIPS and CVPR. We published some of the earliest formal methods for guaranteeing convergence in distributed training. At Shopify we've deployed warehouse automation and robot fleets building behavior trees, fleet control planes, and OTA infrastructure that performs reliably over constrained networks. We bring repeat founder experience having built and exited a healthcare AI company. The Role What you’ll own As a Forward Deployed Engineer at Cedana, you’ll lead and own technical engagement from end to end. You’ll engage with customers to understand and deploy in their environments: from production SLURM at a university, bare-metal Kubernetes at an inference provider, hybrid setup at a Fortune 100 Pharma enterprise. You’ll rapidly understand their key pain points, and use Cedana to solve their problems. For each customer you own everything from the OS up: SLURM plugins, Kubernetes operators, node configuration, networking, and observability. This role will expose you to the cutting edge of AI and HPC infrastructure, working with the world’s leading research and commercial customers to deliver a breakthrough solution. What You'll Do Engineer solutions at client sites: Lead customer integrations. Install, configure, and deploy Cedana into SLURM, Kubernetes, and Dynamo environments. Drive product innovation from the field: Identify technical gaps while embedded with clients, then provide product feedback for new capabilities that become core product features. Measure and optimize platform performance: Measure reliability, throughput, and performance using our internal tools. Design and implement policy-based migration automations to optimize reliability, throughput, and performance Own critical deployments: Ensure our platform performs reliably for clients' critical operations, debugging issues across the full stack. Debug install issues against unfamiliar customer infrastructure, and escalate to engineering when necessary. Improve scalability: Build and own the internal installation playbook so that the second customer in each segment is onboarded faster than the first. Respect our customers: Understand how to make their lives easier and minimize their time and overhead. What we are looking for Team management experience. Requires strong project and time management skills, delivering milestones on time, and effective 3-10 years of software engineering experience with a track record of configuring and managing SLURM deployments. A multi-month enterprise or research deployment you led end-to-end, from scoping through signoff. You write effective status updates to keep your team updated and on schedule. Production experience in standing up SLURM in a customer or research environment. You've configured slurmctld, slurmdbd, accounting, cgroup integration, and GPU resource selection. Strong Linux fundamentals of systemd, cgroups v2, namespaces, network
What the role involves
Introducing Cedana The Problem AI and HPC infrastructure suffer from scarcity and high costs, so when failures occur, they are costly in time and money. Cluster productivity directly determines research output and revenue. Achieving high utilization and throughput is increasingly challenging due to the complexity of workloads, hardware, and operations. Cedana’s Solution Cedana maximizes AI+HPC cluster utilization and reliability with automated GPU checkpointing infrastructure. We enable transparent, fast migration of GPU workloads across instances without losing work. Workloads automatically migrate to achieve new levels of reliability and throughput while accelerating time to results. Our system is at the kernel/OS level, requiring no code or config changes, and works seamlessly with Kubernetes, SLURM, and NVIDIA Dynamo. Today, we're deploying into leading inference platforms, neoclouds, enterprise, and research clusters. The Team Cedana's founding team has spent over a decade making computation run fast, productively, and reliably for AI. Our research appears in NeurIPS and CVPR. We published some of the earliest formal methods for guaranteeing convergence in distributed training. At Shopify, we've developed a control plane for robotics fleets used in warehouse automation. We bring repeat founder experience, having built and exited a Series B healthcare AI company. Backed by Y Combinator, Initialized Capital, Pebblebed (founders of OpenAI and Facebook AI Research), Keith Adams (engineer #20 at VMware, founded HHVM and FAIR at Facebook, Chief Architect at Slack), Venture Guides, Garry Tan, and Gokul Rajaram. The Role What you’ll own You will own the core components of our test and validation for kernel development. This includes measurement and optimization. Over time, you will contribute to critical core components of our system, with a focus on new capabilities, performance, and reliability. What You'll Do Validate and test automation: Our engineers contribute to and develop testing capabilities. This will be a core part of your initial work to establish your understanding of how our system works. Reliability and testing is a key part of our culture. Measure and optimize platform performance: Get Cedana to the theoretical maximum performance by understanding fundamental bottlenecks. Measure reliability, throughput and performance using our internal tools. Design and contribute to key system components : Our solution touches all the major aspects of the OS, kernel, GPU, and CPU. Write design papers to outline your vision for a specific capability and then lead implementation. Educate our team: Our team is our best learning resource and we continually educate each other. We huddle and co-pair as needed. Excellent communication: You enjoy writing concise and articulate design papers on your code and experiments. You respond to slacks and emails within our internal SLAs. What we are looking for 5-10 years of software engineering experience with Linux Kernel. Kernel-level depth: reads kernel and driver source, and root-causes defects at the kernel/driver/syscall boundary rather than only reproducing them at the surface. Comfortable operating below the abstraction line, at the hardware/software interface. Proven expert-level mastery of at least one performance- or correctness-critical Linux systems domain (networking data plane, virtualization/hypervisor, storage and block I/O, scheduling, memory management, or real-time/determinism), plus demonstrated ability to ramp into an unfamiliar low-level subsystem quickly. Depth in one hard domain is a valuable signal that you can reach depth in the next; we are hiring the descent capability, not the specific subsystem. This included some combination of: OVS, DPDK, SR-IOV, RDMA, or high-performance packet processing CRIU + QEMU live migration + VFIO runc / containerd / OCI / namespaces / cgroups v2 / overlayfs Performance and latency engineering: characterizes throughput, jitter,
Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.
Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.
Questions and experiences
Nobody has asked anything about Cedana yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.
Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.