
What the role involves
Goal: 99.99% uptime
We serve custom inference stacks that have irregular GPU load.
We're looking for people that have done genuinely amazing work in infrastructure that are interested in a challenge, working with both traditional infrastructure such as load balancers, NLB, etc., as well as very different infrastructure around inference engines and GPU loads.
This is a role that will inherently require deep experience with inference engines.
Contributions to vLLM, SGLang, trtllm, or inference frameworks a plus.
Every role at Morph comes with unlimited tokens on claude code/codex
About Morph
Specialized inference optimization for codegen Fast Open-Source models: Deepseek v4 flash, qwen 397b, at 200+ tps Specialized models for applying edits, code search, compaction, and model routing
Full Morph profile