← All open roles

H34 Β· AI research

Agent Planning and Reasoning Scientist

Help an agent finish meaningful work while recognizing when it needs a person. You will study planning under uncertainty and limited authority.

Research programOffice-first9 cities

About the role

You study planning under uncertainty and limited authority: how an agent finishes meaningful work, and how it recognises when it needs a person. The second half is the harder and more important one, because an agent that never asks is an agent that eventually does something irreversible.

The work

Investigate task decomposition, tool selection, error recovery, uncertainty and learning from feedback. Use realistic environments with hidden failures and changing state. Keep authority outside model-generated text and work with product researchers on when asking a question improves the outcome.

What good looks like

In your first 90 days, establish a task benchmark and demonstrate a planning improvement against a strong baseline with failure analysis.

Evidence we look for

Bring research in reasoning, reinforcement learning, planning or decision systems. You should be comfortable explaining a method's assumptions and why a success metric may be misleading.

What we need to see

  • Research in reasoning, reinforcement learning, planning, or decision systems
  • You can explain a method's assumptions and where they stop holding
  • You can say why a success metric may be misleading, from experience of one that was
  • You take limited authority seriously as a design constraint rather than a wrapper

Nice to have

  • Human-in-the-loop or interactive decision systems
  • Formal reasoning about safety or corrigibility
  • You have shipped a planner into a product

The exercise

Analyze an agent that appears successful because it declares completion before the external service confirms the action.

Where and how we work

In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.