Cloud Inference Engineer

at Luminal — Making AI run fast on any hardware.

San Francisco, CA, USFull-timeAny (new grads ok)$150K - $350K0.50% - 2.00% equityYC-S25

What the role involves

Qualifications

CUDA + GPU inference optimization

vLLM, SGLang, or TensorRT-LLM experience

KV caching, paged attention, batching, token streaming, etc.

Distributed compute (with GPUs is a super plus)

No degree required

Company

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Role

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

Day to day responsibilities

  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels and, yes, occasional tasteful shitposting

What they ask for

Torch/PyTorchCUDA

About Luminal

Luminal builds an ML framework and compiler that generates GPU code. Our stack 10x's model speed while simplifying deployment and cutting idle GPU costs Github: https://github.com/luminal-ai/luminal Discord: https://discord.gg/APjuwHAbGy

Full Luminal profile

Other roles at Luminal

Similar Engineering, Machine learning elsewhere

jobo is a browser extension. Open this on a computer to install it.