Kernel Engineer — Scientific Computing (SPU)

at Vorticity — The Fastest Scientific Computing Platform on the Planet

Redwood CityFull-timeAny (new grads ok)$120K - $170K0.25% - 0.50% equityYC-S19

What the role involves

Vorticity is building the world’s first Scientific Processing Unit (SPU), a new class of silicon purpose-built to accelerate scientific computing beyond the limits of GPUs. We are designing tightly coupled software–hardware systems around applied mathematics workloads to deliver order-of-magnitude performance gains. Unlocking its full potential requires early, deep engagement from applied mathematics–driven software engineers who can translate real-world scientific workloads into executable models, kernels, libraries, and applications that inform both architecture and tooling decisions.

As a Kernel Engineer, you will work at the intersection of applied mathematics, scientific computing, parallel programming, and low-level performance engineering. You will help shape how numerical kernels are implemented, optimized, and eventually mapped onto the SPU. Your work may include building early numerical kernels and libraries, developing prototype applications, and writing Python-based workload models and simulators, all to support and inform the evolving hardware and compiler stack.

This requires both strong applied math fundamentals and deep low-level implementation ability. You should be comfortable moving from mathematical formulations to efficient kernels, reasoning about accuracy, stability, data movement, memory hierarchy, parallel execution, and compiler behavior along the way. This position is ideal for someone who combines strong scientific computing instincts with the low-level habits of a performance engineer.

Responsibilities

Prototyping and implementing core kernels and low-level numerical primitives for the SPU.

Translating mathematical formulations into executable, performance-relevant kernel implementations.

Analyzing and optimizing memory-access patterns, including coalescing, locality, shared memory usage, cache behavior, register pressure, and host-device data movement.

Collaborating closely with hardware architects to evaluate algorithm–architecture tradeoffs around memory hierarchy, synchronization, vector/SIMT execution, instruction behavior, and parallel scheduling.

Working with compiler and runtime teams to ensure kernels map cleanly to the SPU programming model.

Designing microbenchmarks, correctness tests, numerical accuracy tests, and performance models, then iteratively refining kernels based on hardware evolution, compiler behavior, profiler output, and measured performance.

Core Skills

  • Strong applied mathematics and scientific computing judgment, with the ability to understand numerical workloads deeply enough to implement them correctly and efficiently.
  • Strong proficiency in C++ and CUDA, HIP, SYCL, or an equivalent accelerator programming model.
  • Experience writing custom kernels, not just using existing frameworks or vendor libraries.
  • Ability to translate mathematical formulations into low-level implementations while balancing accuracy, stability, precision, data movement, and performance.
  • Deep understanding of GPU execution and memory hierarchy, including global memory, shared memory, registers, caches, coalescing, atomics, reductions, scans, warp-level execution, and occupancy.
  • Experience using profiling and performance tools to identify bottlenecks, test hypotheses, and validate improvements.
  • Ability to reason from profiler output to concrete code changes, rather than treating performance debugging as guesswork.
  • Solid concurrency fundamentals, including race conditions, atomicity, synchronization, and thread/process execution behavior.

Nice to Have Skills

  • Familiarity with performance analysis tools or modeling techniques (profilers, roofline models)
  • Exposure to compilers, runtimes, or code generation frameworks
  • Experience in applied scientific domains such as physics, geophysics, CFD, climate, materials, fusion, or finance.
  • Experience with low-level GPU assembly or intermediate representations.
  • Familiarity with low-level system software or drivers.

Non-Technical Qualities

Excellent written

What they ask for

C++PythonCUDA

About Vorticity

Many pressing problems facing humanity can be solved with faster scientific compute. Currently, it takes billions of dollars to design and develop a new fusion reactor, a hyper-sonic airplane or a new cancer treatment. This is because scientists and engineers have to build very expensive things before they even know if their ideas work. But what if we can simulate the workings of these ideas (and more) on computers first and have high confidence that it works before we build really expensive things? The work of humanity's greatest scientists over the years has given us the physics and math to do this. The problem is that we simply do not have the computing power to solve these big problems. To address this, the Vorticity Inc is re-imagining scientific computing all the way down to processor architecture and system design. Our proven technology is already helping the energy, life sciences and aerospace industries.

Full Vorticity profile

Other roles at Vorticity

Similar Engineering, Full stack elsewhere

jobo is a browser extension. Open this on a computer to install it.