
Exla
An SDK to run transformer models anywhere
What Exla does
Exla aggressively quantizes AI models to minimize memory usage and maximize inference speed. Whether you're deploying LLMs, VLMs, VLAs, or custom models, Exla reduces memory footprint by up to 80% and accelerates inference by 3–20x - all with just a few lines of code. https://cal.com/exla-ai/schedule
Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.
Questions and experiences
Nobody has asked anything about Exla yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.
Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.