
What Wafer does
Wafer builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference. Our product is serverless and dedicated inference for the world’s fastest open source LLMs, achieved by Wafer's autonomous performance engineers.
3 open roles
What the role involves
About the Role Wafer's mission is to maximize intelligence per watt, by building AI that optimizes AI itself. Our journey starts with GPU kernels, but will expand into every corner of ML systems and AI infrastructure. We're a small team (4 people) backed by Fifty Years, Y Combinator, Jeff Dean, and Woj Zaremba (co-founder of OpenAI), and we're looking for engineers who want to work at the intersection of AI agents and systems programming. You'll work directly with the founding team to build the systems that power our GPU optimization platform, from the agent framework that iterates on kernels, to the profiling infrastructure that connects to NCU and ROCprofiler, to the compiler tooling that analyzes PTX and SASS. What You'll Do Build and improve our framework for GPU kernel optimization (multi-turn tool use, state management, reward signals) Develop integrations with GPU profilers and compiler toolchains Design the architecture for remote GPU execution across cloud GPUs Work on trace analysis systems that help the agent diagnose performance bottlenecks Ship features that engineers use daily, and that optimizes infrastructure that runs the world's AI (PyTorch, vLLM, NVIDIA, AMD, etc.) What We Look For You're a strong fit if you: Have deep technical intuition and can learn new domains quickly Are comfortable working across the stack Can ship production code fast while maintaining quality Want to work on some of the most interesting AI infra problems at a small company with no bullshit + ship fast culture. Very nice to have: GPU programming experience (CUDA, HIP, Triton) Experience with profiling tools or compiler internals Background in AI/ML research or agent systems Publications or open-source work in relevant areas
What the role involves
Join our team to build the future of inference, GPU optimization and AI infrastructure. You'll work directly with the team to define our technical direction and build the core systems that power our GPU optimization platform. What You'll Do Build scalable infrastructure for AI model training and inference Lead technical decisions and architecture choices What We Look For Core Technical Expertise GPU Fundamentals: Deep understanding of GPU architectures, CUDA programming, and parallel computing patterns. Deep Learning Frameworks: Proficiency in PyTorch, TensorFlow, or JAX, particularly for GPU-accelerated workloads. LLM/AI Knowledge: Strong grounding in large language models (training, fine-tuning, prompting, evaluation). Systems Engineering: Proficiency in C++, Python, and possibly Rust/Go for building tooling around CUDA.
What the role involves
Join our team to build the future of inference, GPU optimization and AI infrastructure. You'll work directly with the team to define our technical direction and build the core systems that power our GPU optimization platform. What You'll Do Build scalable infrastructure for AI model training and inference Lead technical decisions and architecture choices What We Look For Core Technical Expertise GPU Fundamentals: Deep understanding of GPU architectures, CUDA programming, and parallel computing patterns. Deep Learning Frameworks: Proficiency in PyTorch, TensorFlow, or JAX, particularly for GPU-accelerated workloads. LLM/AI Knowledge: Strong grounding in large language models (training, fine-tuning, prompting, evaluation). Systems Engineering: Proficiency in C++, Python, and possibly Rust/Go for building tooling around CUDA. Ideal Background Publications or open-source contributions in inference GPU computing or ML/AI for code are a plus. Hands-on experience with large-scale experiments, benchmarking, and performance tuning.
Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.
Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.
Questions and experiences
Nobody has asked anything about Wafer yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.
Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.