
What the role involves
Qualifications
CUDA + GPU inference optimization
vLLM, SGLang, or TensorRT-LLM experience
KV caching, paged attention, batching, token streaming, etc.
Distributed compute (with GPUs is a super plus)
No degree required
Company
Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.
Role
Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.
Day to day responsibilities
- Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
- Conducting model performance reviews
- Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
- Sometimes write kernels and, yes, occasional tasteful shitposting
What they ask for
Torch/PyTorchCUDA
About Luminal
Luminal builds an ML framework and compiler that generates GPU code. Our stack 10x's model speed while simplifying deployment and cutting idle GPU costs Github: https://github.com/luminal-ai/luminal Discord: https://discord.gg/APjuwHAbGy
Full Luminal profileOther roles at Luminal
Similar Engineering, Machine learning elsewhere
Residency (Summer 2026)Adam · San Francisco, CA, USFounding Engineer - Full StackAndy AI · San Francisco, CA, USFounding Engineer - Healthcare AIAndy AI · San Francisco, CA, USMember of Technical Staff (applied)Anthrogen · San Francisco, CA, USApplied AIArtisan · San Francisco, CA, US / New York, NY, USData ScientistAxross Pte Ltd · MY / Kuala Lumpur, Federal Territory of Kuala Lumpur, MY / Remote (MY; ID; SG)AI / ML EngineerBland AI · San Francisco, CA, USMachine Learning EngineerBlank Bio · San Francisco, CA, US