← All open roles

H31 Β· AI infrastructure

Performance and Energy Benchmark Engineer

Make performance claims useful to the person buying the computer. You will measure what the system accomplishes, what it costs and how reliably it works.

Pilot expansionOffice-first9 cities

About the role

You make performance claims useful to the person buying the computer, by measuring what the system accomplishes, what it costs, and how reliably it does it. This role exists partly to keep us honest: you are the person who will identify our own cherry-picked workload and say so.

The work

Build representative task suites covering inference, retrieval, training and integrations. Record model version, quality, context, precision, memory and device conditions. Measure latency distributions, wall energy and completed-task cost alongside raw throughput.

What good looks like

In your first 90 days, publish a repeatable baseline and a comparison that another engineer can reproduce on the same configuration.

Evidence we look for

Bring experimental design, systems measurement and clear technical writing. Be able to identify cherry-picked workloads and distinguish measurement noise from a meaningful change.

What we need to see

  • Experimental design, systems measurement, and clear technical writing in combination
  • You can identify a cherry-picked workload and distinguish noise from a meaningful change
  • You publish methodology alongside results, so somebody else can reproduce them
  • The independence to report a result the company will not enjoy

Nice to have

  • Energy measurement as well as performance
  • You have built a benchmark suite others adopted
  • Statistics depth

The exercise

Design a fair comparison between two devices when their fastest available model configurations have different quality.

Where and how we work

In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.