About the role
You study planning under uncertainty and limited authority: how an agent finishes meaningful work, and how it recognises when it needs a person. The second half is the harder and more important one, because an agent that never asks is an agent that eventually does something irreversible.
The work
Investigate task decomposition, tool selection, error recovery, uncertainty and learning from feedback. Use realistic environments with hidden failures and changing state. Keep authority outside model-generated text and work with product researchers on when asking a question improves the outcome.
What good looks like
In your first 90 days, establish a task benchmark and demonstrate a planning improvement against a strong baseline with failure analysis.
Evidence we look for
Bring research in reasoning, reinforcement learning, planning or decision systems. You should be comfortable explaining a method's assumptions and why a success metric may be misleading.
What we need to see
- Research in reasoning, reinforcement learning, planning, or decision systems
- You can explain a method's assumptions and where they stop holding
- You can say why a success metric may be misleading, from experience of one that was
- You take limited authority seriously as a design constraint rather than a wrapper
Nice to have
- Human-in-the-loop or interactive decision systems
- Formal reasoning about safety or corrigibility
- You have shipped a planner into a product
The exercise
Analyze an agent that appears successful because it declares completion before the external service confirms the action.
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.