About the role
You make each unit of compute do more useful work, optimising the kernels and compilation paths that matter to real workloads rather than to a paper. Mastery of every backend is not expected; depth in one and the judgement to find the bottleneck is.
The work
Profile operators, memory traffic and scheduling; implement and validate kernels on selected backends. Work with numerical and model researchers on precision choices. Maintain portable abstractions where useful while documenting backend-specific behavior and limits.
What good looks like
In your first 90 days, deliver one measured end-to-end improvement with correctness tests, hardware details and a reproducible comparison.
Evidence we look for
Bring GPU programming, compiler or numerical computing expertise. Relevant experience may include CUDA, Metal, LLVM, MLIR, Triton or other accelerator toolchains; mastery of every backend is unnecessary.
What we need to see
- GPU programming, compiler, or numerical computing expertise with real results behind it
- Depth in at least one accelerator toolchain such as CUDA, Metal, LLVM, MLIR, or Triton
- You profile to find the real bottleneck rather than optimising what is familiar
- You can prove a speedup that survives being used in production
Nice to have
- Compiler internals as well as kernel writing
- Numerical stability under low precision
- Open-source kernel or compiler contributions
The exercise
Show how a faster isolated kernel could still make the complete application slower, and design the correct benchmark.
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.