About the role
You make personal computing reachable through the ways people naturally communicate: speech, vision, and whatever else somebody brings to it, under the constraints of a device rather than a data centre. The recurring trap in this area is fluency, which is now easy to produce and easy to mistake for understanding. Telling those apart, and evaluating on the input people really give rather than the input that demos well, is the work.
The work
Develop or adapt models for speech, documents, images and grounded interaction. Evaluate diverse accents, environments and accessibility needs with appropriate participant consent. Measure local performance and make capture, retention and activation behavior understandable.
What good looks like
In your first 90 days, deliver a bounded multimodal capability with evaluation across realistic conditions and a clear record of its limitations.
Evidence we look for
Bring research or advanced engineering in speech, vision or multimodal learning. Show how you distinguish apparent fluency from correct understanding.
What we need to see
- Research or advanced engineering in speech, vision, or multimodal learning
- You can show how you distinguish apparent fluency from correct understanding
- You work within real device constraints rather than assuming a server
- You evaluate on realistic input, including accents, noise, and bad lighting
Nice to have
- On-device speech recognition or synthesis
- Accessibility-driven design
- Low-resource languages
The exercise
Design an evaluation for a voice-controlled task in a shared room where only one person has authorized the action.
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.