About the role
You own the agent runtime: tools, memory, planning, and the evaluation suite that keeps all of it honest. This is the part of the system where impressive demos and reliable products diverge, and the difference is almost always the evals. You will build for an agent that often runs on the user's own hardware, under a consent boundary, which makes cost and latency real constraints rather than nice-to-haves.
What we need to see
- Strong engineering plus hands-on LLM and agent work: tool use, retrieval, memory, planning
- You have shipped an agent or AI feature to real users, not only to a demo audience
- Genuine rigor about evaluation and regression, including held-out suites you trust
- You care about cost per successful task and can move it
Nice to have
- You have built an eval harness others adopted
- On-device or quantised inference experience
- Open-source agent tooling anyone can read
What winning looks like
- Task success rate and quality on a held-out eval suite
- First-token and end-to-end latency budgets met
- Cost per successful task trending down
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.