Cumulus Labs

The Fastest Multimodal Inference OS

YC-W26B2B -> InfrastructureEarly

What Cumulus Labs does

Cumulus Labs lets engineering teams ship AI in production without needing a dedicated ML platform team. Right now, companies building AI products are forced to stitch together separate vendors for routing, observability, evaluation, fine-tuning, and inference. This fragmented approach is brittle, expensive, and is a common reason enterprises fail with AI. We replace that entire stack with a single unified platform. Developers can keep their existing code while instantly upgrading to a unified platform that handles routing, semantic caching, continuous shadow evaluation, simulated data, and one-click fine-tuning. Behind the platform is Ion, our proprietary inference engine running on a custom NVIDIA Grace GPU fleet. Ion uses in-house custom GPU kernels to deliver 30 to 50 percent more throughput than standard vLLM or SGLang, giving our customers SOTA inference economics.

No open roles recorded right now.

That does not mean they are not hiring — plenty of roles never reach a job board. The careers page below is the earliest place a new opening appears.

Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.

Questions and experiences

Nobody has asked anything about Cumulus Labs yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.

Reviewed before it appears. Do not include anything that identifies you or anyone else.

Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.

jobo is a browser extension. Open this on a computer to install it.