Member Of Technical Staff (Winter Intern)

at Wafer — AI that makes AI fast

San Francisco, CA, USInternship$6K - $10K / monthlyYC-S25

What the role involves

Join our team to build the future of inference, GPU optimization and AI infrastructure. You'll work directly with the team to define our technical direction and build the core systems that power our GPU optimization platform.

What You'll Do

Build scalable infrastructure for AI model training and inference

Lead technical decisions and architecture choices

What We Look For

Core Technical Expertise

GPU Fundamentals: Deep understanding of GPU architectures, CUDA programming, and parallel computing patterns.

Deep Learning Frameworks: Proficiency in PyTorch, TensorFlow, or JAX, particularly for GPU-accelerated workloads.

LLM/AI Knowledge: Strong grounding in large language models (training, fine-tuning, prompting, evaluation).

Systems Engineering: Proficiency in C++, Python, and possibly Rust/Go for building tooling around CUDA.

Ideal Background

Publications or open-source contributions in inference GPU computing or ML/AI for code are a plus.

Hands-on experience with large-scale experiments, benchmarking, and performance tuning.

What they ask for

C++PythonTypeScriptTorch/PyTorchCUDA

About Wafer

Wafer builds AI agents that work as autonomous performance engineers, optimizing GPU kernels for AI inference. Our product is serverless and dedicated inference for the world’s fastest open source LLMs, achieved by Wafer's autonomous performance engineers.

Full Wafer profile

Other roles at Wafer

Similar Engineering, Machine learning elsewhere

jobo is a browser extension. Open this on a computer to install it.