LiteLLM

Call every LLM API like it's OpenAI [100+ LLMs]

Hiring — 8 openYC-W23Early

What LiteLLM does

LiteLLM is an open-source LLM Gateway with 18K+ stars on GitHub and trusted by companies like Rocket Money, Samsara, Lemonade, and Adobe. LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund

8 open roles

Backend LLM Engineer
San Francisco, CA, US / Remote (MX)Full-time1+ years$120K - $180K0.25% - 0.75% equityVisa: US citizenship/visa not required
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We’re rapidly expanding and seeking our 6th Engineer focused on owning ‘excellence’ for unified API’s across Core LLM’s (openai/gemini/anthropic models). What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We just hit $6M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. Why do companies use LiteLLM enterprise Companies use LiteLLM Enterprise once they put LiteLLM into production and need enterprise features like Prometheus metrics (production monitoring) and need to give LLM access to a large number of people with SSO (secure sign on) or JWT (JSON Web Tokens) What you will be working on Skills: Python, LLM APIs, FastAPI, High-throughput/low-latency As the Backend LLM Engineer, you'll be responsible for ensuring LiteLLM unifies the format for calling LLM APIs in the broader OpenAI + Anthropic spec. This involves writing transformations to convert API requests from OpenAI/Anthropic spec to various LLM provider formats, building provider-agnostic unification functionality (e.g. session management across non-openai models for /v1/responses API, etc.). You'll work directly with the CEO and CTO on critical projects including: Adding support for Anthropic and Bedrock Anthropic 'thinking' parameter Handling provider-specific quirks like OpenAI o-1 streaming limitations Maintaining ‘excellent’ unified API’s across /v1/messages, /v1/responses, /chat/completions for OpenAI/Gemini/Anthropic models across Azure, OpenAI API, Bedrock Invoke, Bedrock Converse, Vertex AI, Google AI Studio Implementing cost tracking and logging for Anthropic API What is our tech stack The tech stack includes Python, FastAPI, Redis, Postgres. Who we are looking for 1-2 years of backend/full-stack experience with production systems Passion for open source and user engagement Experience working with the OpenAI api (understand the difference between /chat/completions and /responses, and can speak to API-specific nuances) Strong work ethic and ability to thrive in small teams Eagerness to talk to users and help solve real problems

Backend MCP Engineer
San Francisco, CA, US / Remote (US)Full-time1+ years$120K - $180K0.25% - 0.75% equityVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We’re rapidly expanding and seeking our 6th Engineer focused on owning ‘excellence’ for MCP’s on LiteLLM. What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We just hit $7M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. Why do companies use LiteLLM enterprise Companies use LiteLLM Enterprise once they put LiteLLM into production and need enterprise features like Prometheus metrics (production monitoring) and need to give LLM access to a large number of people with SSO (secure sign on) or JWT (JSON Web Tokens) What you will be working on Skills: Python, MCP, AI infrastructure, FastAPI As the Backend MCP Engineer, you'll be responsible for implementing MCP server support, building tool orchestration layers, designing protocol for external tool integration, enabling function calling across multiple LLM providers, and creating SDK for MCP server discovery and connection. You'll work directly with the CEO and CTO on critical projects including: Adding MCP protocol support to LiteLLM gateway Building unified tool calling interface across providers Implementing session management for stateful agents Creating examples/docs for MCP + LiteLLM integration What is our tech stack Core: Python, FastAPI, MCP, Redis, Postgres. LLM Integration: OpenAI SDK, Anthropic SDK, AWS Bedrock, Vertex AI Protocol Layer: JSON-RPC, WebSockets, Server-Sent Events (SSE) Agent Tooling: Model Context Protocol (MCP), function calling, tool schemas Infrastructure: Docker, Kubernetes, Prometheus, GitHub Actions You'll work with: Multiple LLM provider APIs (Anthropic, OpenAI, Google, AWS) MCP protocol implementation (client + server) High-throughput async systems (10K+ req/sec) Open source community (34K+ GitHub stars) What’s so exciting about this role? LiteLLM is at the intersection of 3 critical AI infrastructure layers: 1. LLM Gateway - Call any LLM with one API (our core strength) 2. MCP Gateway - Give any LLM access to any tool (emerging need) 3. Agent Gateway - Enable agents to communicate with other agents/llm’s/tools You'll help us become the unified infrastructure layer that connects:  Applications ↔ LiteLLM ↔ LLM Providers (OpenAI, Anthropic, Bedrock) LLMs ↔ LiteLLM ↔ MCP Servers (databases, APIs, internal tools)  Agents ↔ LiteLLM ↔ MCP Servers (databases, APIs, internal tools) + LLMs This means working on cutting-edge problems like: How do we route tool calls across providers with different specs? How do we make MCP servers work seamlessly with any LLM? How do we build the "Stripe of AI infrastructure"? If you're excited about building the foundational layer that every AI application will use, this is for you. Who we are looking for 1-2 years of backend/full-stack experience with production systems Passion for open source and user engagement Experience working with the OpenAI api (understand the difference between /chat/completions and /responses, and can speak to API-specific nuances) Strong work ethic and ability to thrive in small teams Eagerness to talk to users and help solve real problems

Founding Account Executive
San Francisco, CA, USFull-time3+ years$100K - $200K0.05% - 0.50% equityVisa: Will sponsor
What the role involves

About LiteLLM LiteLLM is the world’s most popular AI Gateway used by the largest companies (Adobe, Netflix, NASA, etc.) in the world to give their developers access to LLMs and adjacent services (MCP’s, Vector Stores, etc.). About the role We're looking for a Founding Account Executive with experience in selling developer tools (e.g. gitlab, datadog, etc.) to large companies (>1k+ employees). You’ll own the sales cycle from qualified opportunity to close, working closely with the founders and early sales team. As the first sales hire, you will be working closely with the Founders (ex-Microsoft, ex-Coinbase) to help us transition from a founder-led sales motion to a scalable sales motion as we continue to grow at an exponential rate in 2025. Key Responsibilities Own the Sales Process: Manage the entire sales cycle from qualified lead to close, including discovery, demo, proposal, and negotiation. Drive Revenue: Work with the founders to help LiteLLM hit it’s revenue goals. Customer Feedback Loop: Collaborate with product and engineering by sharing insights from customer conversations. Build Repeatable Processes: Establish repeatable sales processes that will scale as the company grows. Build the Sales Team: As we grow, you will also be responsible for building the future sales team by helping define KPIs, hiring criteria, and culture. Why Work At LiteLLM? You love working with developers (30k+ Github stars) You want to work hard (need to be always on - calls on weekends/post-dinner/etc.) You want to learn how AI is transforming businesses (Businesses run ALL LLM calls through LiteLLM) You enjoy building things from scratch (defining GTM, sales playbook, hiring) Pay Transparency The annual base salary range for this role is $100,000 - $175,000, with additional commission and on-target earnings Top-tier Medical, Dental, and Vision benefits, including FSA Qualification 2-4+ years of experience in full-cycle B2B sales, ideally in SaaS or tech startups. Experience selling to regulated industries (e.g. insurance, financial services) is a strong plus. Proven track record of meeting or exceeding quota. Excellent communication and storytelling skills, with the ability to tailor messaging to diverse buyer personas. Self-starter who thrives in an early-stage environment and is excited to help build sales processes from the ground up Strong organizational skills, with the ability to manage multiple opportunities through various stages of the sales pipeline. Comfortable with CRM tools, sales automation platforms, and a fast-paced, feedback-driven culture.

Founding Reliability & Performance Engineer
San Francisco, CA, USFull-time3+ years$200K - $270K0.25% - 0.75% equityVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source AI gateway (36K+ GitHub stars) that routes hundreds of millions of LLM API calls daily for companies like NASA, Adobe, Netflix, Stripe, and Nvidia. We're at $7M ARR, 10 people, YC W23. When LiteLLM goes down, our customers' entire AI stack goes down. We need someone who makes sure that doesn't happen. You'd be the first dedicated reliability hire. You'll own reliability, performance, and production stability end-to-end. Nobody will tell you how to do it What this job actually is We'll be straight with you: this role is roughly 60% operational reliability and 40% deep performance engineering. On any given week you might be: Hunting a memory leak in our async streaming handler that causes OOMs after 4 hours under load Fixing a race condition where PodLockManager releases another pod's lock Profiling why update_database() does 7 deep copies per request in the spend tracking hot path Helping a Fortune 500 customer debug why their 20-pod deployment is exhausting Postgres connections Building soak tests that catch degradation before a release goes out Reviewing a PR that touches the request hot path and saying "this will add 50ms at P99, here's why" If you're looking for a pure optimization role where you sit in a profiler all day — this isn't it. If you want to own production health for one of the most widely deployed AI infrastructure projects in the world — keep reading. Why this matters We route traffic for some of the largest AI deployments on the planet. One customer is scaling from 20M to 200M daily AI calls through our gateway. Another has 150K users hitting us daily. When we ship a bad release, it doesn't just break a dashboard — it breaks production AI systems at companies you've heard of. The problems here are genuinely hard: Memory management in long-running Python async services — our proxy handles thousands of concurrent streaming connections. HTTP client sessions, response iterators, and background tasks all need careful lifecycle management. Database at scale — spend logging, auth, and rate limiting all interact with Postgres. At 100K+ requests/day, naive patterns fall apart. 100+ provider surface area — we translate between OpenAI, Anthropic, Bedrock, Vertex, and 100+ other APIs. Each has unique streaming behavior. A refactor that fixes one provider can break three others. You won't run out of interesting problems. What you'll own Production reliability On-call for critical issues (shared rotation with the team, not solo) Incident response and blameless post-mortems Customer escalation support for enterprise deployments Making the proxy self-healing when DB/Redis is temporarily unavailable Performance engineering Memory leak detection and prevention (soak tests, CI integration) Hot path optimization — our target is <10ms overhead at 5K+ RPS P50/P95/P99 latency benchmarks that block releases on regression Profiling and fixing bottlenecks (Pydantic validation, connection pools, async task scheduling) Observability & release safety Structured logging, distributed tracing, correlation IDs Prometheus metrics that are actually accurate and actionable Building toward canary deployments and automated rollback SLO definition and tracking for enterprise customers Who you are Must have: 2+ years of experience running Python services in production, with real exposure to debugging things that break at scale Strong understanding of Python async internals — asyncio event loop, aiohttp/httpx session management, connection pooling Experience debugging production memory leaks, OOMs, or latency degradation (bonus if you've used memray, py-spy, or tracemalloc) Solid PostgreSQL knowledge — connection pool tuning, query optimization, understanding how DB operations on the request path degrade under load Comfort with Kubernetes at an operational level — pod lifecycle, resource limits, health probes You've been on-call before and you didn't hate it Strong signals: You've worked on a proxy, API ga

KubernetesPostgreSQLPythonRedis
Site Reliability Engineer
San Francisco, CA, USFull-time1+ years$150K - $200K0.25% - 0.75% equityVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We're rapidly expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy. What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format. We just hit $6M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. Why do companies use LiteLLM Enterprise Companies use LiteLLM Enterprise once they put LiteLLM into production and need enterprise features like Prometheus metrics (production monitoring) and need to give LLM access to a large number of people with SSO (secure sign on) or JWT (JSON Web Tokens). What you will be working on Skills: Python, FastAPI, PostgreSQL, Redis, Kubernetes, Prometheus, performance profiling As the SRE, you'll own the reliability and performance of the LiteLLM proxy in production. Our users run LiteLLM as a critical gateway handling millions of LLM requests — when it goes down, their entire AI stack goes down. You'll work directly with the CEO and CTO on critical projects including: Fixing OOM issues — e.g. Prisma Query Engine unable to recover from OOMKill in K8s deployments, unbounded in-memory buffers in spend log transactions Solving database connection problems — e.g. database query limits getting reached under load, spend logs loading extremely slowly, Prisma connection pool exhaustion Fixing race conditions and deadlocks — e.g. max_parallel_requests deadlocking API keys after provider timeouts (counter never released, Redis reset required), PodLockManager releasing another pod's lock, in-memory cache increment race conditions Performance optimization — e.g. update_database() doing 7 deep copies per request in the spend tracking hot path, health check fan-out overloading startup Improving Redis/cache reliability — e.g. budget limiter reading stale Redis data, cache sync issues between in-memory and Redis layers Production monitoring — making Prometheus metrics accurate (fixing missing/inf budget metrics), adding alerting, improving observability for multi-pod deployments Making the proxy self-healing — graceful degradation when DB/Redis is temporarily unavailable, connection retry logic, proper health checks What is our tech stack The tech stack includes Python, FastAPI, Redis, Postgres, Prisma ORM, Kubernetes, Prometheus, Docker. Who we are looking for 1-4 years of experience running Python services in production at scale Experience debugging OOMs, memory leaks, connection pool issues, and race conditions Comfortable with PostgreSQL (query optimization, connection pooling, PgBouncer) and Redis Kubernetes experience — you've dealt with pod restarts, resource limits, health probes, and multi-replica coordination Familiarity with Prometheus/Grafana for monitoring and alerting Passion for open source and user engagement Strong work ethic and ability to thrive in small teams Eagerness to talk to users and help solve real problems — our GitHub issues are full of production debugging sessions and you'd be jumping into those directly

KubernetesPostgreSQLPythonRedisDocker
Technical Support Engineer
San Francisco, CA, USFull-timeAny (new grads ok)$80K - $100KVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 28K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We’re rapidly expanding and seeking a performance engineer to help scale the platform to handle 5K RPS (Requests per second). We’re based in San Francisco. What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We just hit $2.5M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. About the Role We’re looking for a Technical Support Engineer to help our customers troubleshoot issues, optimize their setup, and get the most value out of our product. You’ll be the first line of defense for incoming technical questions, ensuring timely resolution and a great customer experience. This role is ideal for someone who enjoys solving complex technical problems, communicating clearly with both technical and non-technical users, and continuously improving internal processes and documentation. Responsibilities Diagnose and troubleshoot technical problems across on LiteLLM. This can be on the backend or frontend. Reproduce issues and create clear, actionable bug reports for engineering. Submit small pull requests for simple fixes, documentation updates, or configuration changes. Escalate bugs to product engineering and close the loop with the customer once shipped. Maintain a high standard of customer empathy—clear communication, active listening, and ownership of problems until resolution. Keep support documentation up to date. Why Work At LiteLLM? You love being on the front lines with customers at a fast-growing startup, quickly scoping their issues, leading troubleshooting discussions, and driving them to successful outcomes. You love working with developers (30k+ Github stars) You want to work hard (966 work culture) You want to learn how AI is transforming businesses (Businesses run ALL LLM calls through LiteLLM)

Backend Performance Engineer
San Francisco, CA, US / Remote (US)Full-timeAny (new grads ok)$150 - $2000.50% - 3.00% equityVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 28K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We’re rapidly expanding and seeking a performance engineer to help scale the platform to handle 5K RPS (Requests per second). We’re based in San Francisco. What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We just hit $2.5M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. About the Role We're hiring a Python performance engineer to own maximizing throughput, minimizing latency and ensuring our platform is reliable in production. Roadmap for Performance Engineer: By end of this year our RPS and latency overhead should be at parity with industry benchmarks. Cover stream + non-stream for /chat/completions, /completions, /embeddings, /realtime, /audio/transcriptions Reduce e2e overhead latency for cache misses. Currently at 100ms-500ms - ensure we meet industry standards. Reduce e2e overhead latency for cache hits - ensure we meet industry benchmarks. Ensure overhead latency scales well when other components are added to the platform - e.g Redis, Redis Cluster, DB, Non-Admin Virtual Keys Ensure overhead latency scales well with payload size - 1MB prompt with streaming should be sub 100ms Address customer specific and pipeline specific latency issues. e.g. Enterprise customers reporting high overhead - this person should be able to debug these issues, get on support calls and help address any environment specific settings. Address paying customer memory leaks Enterprise clients have ongoing memory leaks that need resolution Longer term - should add coverage over new endpoints - /realtime, /audio/transcriptions/, /audio/speech

PythonRust
Forward Deployed Engineer
San Francisco, CA, USFull-timeAny (new grads ok)$80K - $120KVisa: US citizen/visa only
What the role involves

TLDR LiteLLM is an open-source LLM Gateway with 28K+ stars on GitHub and trusted by companies like NASA, Rocket Money, Samsara, Lemonade, and Adobe. We’re rapidly expanding and seeking a performance engineer to help scale the platform to handle 5K RPS (Requests per second). We’re based in San Francisco. What is LiteLLM LiteLLM provides an open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format We just hit $2.5M ARR and have raised a $1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund. You can find more information on our website, Github and Technical Documentation. About the Role We're looking for a Forward Deployed Engineer to embed with our key customers, helping them successfully deploy and scale LiteLLM in production. You'll work directly at customer sites (remotely), troubleshooting complex technical issues, optimizing their infrastructure, and ensuring they extract maximum value from the platform. This role is ideal for someone who thrives in dynamic, customer-facing environments, enjoys solving production-level challenges in real-time, and can translate customer needs into actionable product improvements. Responsibilities Deploy and configure LiteLLM in customer environments, ensuring optimal performance and reliability Diagnose and resolve complex technical issues across the full stack—from infrastructure and backend to frontend integrations Work remotely with customers during critical implementations, migrations, or scaling initiatives Reproduce customer issues in their specific environments and create detailed technical reports for engineering Build custom integrations, scripts, or tooling to meet unique customer requirements Submit pull requests for bug fixes, feature enhancements, documentation improvements, and configuration optimizations Act as the technical voice of the customer—gathering feedback, identifying patterns, and advocating for product improvements Maintain close collaboration with product engineering to ensure customer-reported issues are resolved and communicated back effectively Develop and maintain customer-facing technical documentation, deployment guides, and best practices Own customer relationships from a technical perspective, building trust through deep expertise and responsiveness Why Work At LiteLLM? You thrive working directly with customers at a fast-growing startup—scoping complex technical challenges, leading implementation discussions, and driving them to production success You love working closely with developers (30K+ Github stars) and being at the intersection of product and customer You want to work hard (966 work culture) and move fast in a high-impact role You want to see firsthand how AI is transforming businesses—our customers run ALL their LLM calls through LiteLLM You enjoy the autonomy and variety of working across multiple customer environments and technical stacks

Node.jsPostgreSQLPythonDocker

Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.

Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.

Questions and experiences

Nobody has asked anything about LiteLLM yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.

Reviewed before it appears. Do not include anything that identifies you or anyone else.

Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.

jobo is a browser extension. Open this on a computer to install it.