About the role
You make frontier intelligence run fast on hardware the user already owns. This is the technical heart of the whole promise: if the model has to leave the device, the privacy argument is rhetoric. You will work against real memory, thermal and battery limits on Apple and NVIDIA silicon, and the wins you find are the product.
What we need to see
- Systems and performance engineering on Apple or NVIDIA silicon
- On-device inference in practice: quantisation, memory layout, and working inside thermal limits
- You profile before you optimise and can prove the win with numbers
- Comfort at the boundary between model and machine, rather than treating either as somebody else's layer
Nice to have
- Core ML, Metal, CUDA, or TensorRT specifically
- Kernel or compiler work
- You have published benchmarks somebody else could reproduce
What winning looks like
- On-device tokens/sec and memory footprint within target
- Share of workloads served on-device vs cloud
- Battery and thermal headroom on reference hardware
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.