About the role
You show when a private agent deserves a user's trust, building the evidence about task completion, failure, and human oversight. Two failure modes govern this role: benchmark contamination, and quietly substituting a model judge for ground truth because it is cheaper.
The work
Design evaluations for real workflows, adversarial inputs and distribution shifts. Use held-out tasks, blinded review where appropriate and calibrated human evaluation. Measure false completion, unauthorized actions, appropriate escalation and sustained usefulness over time.
What good looks like
In your first 90 days, deliver a versioned evaluation suite and a release report that makes strengths and unresolved failures easy to inspect.
Evidence we look for
Bring experimental design, statistics and hands-on AI evaluation. Show how you prevent benchmark contamination and avoid substituting a model judge for ground truth.
What we need to see
- Experimental design, statistics, and hands-on AI evaluation
- You prevent benchmark contamination deliberately, with a method you can describe
- You do not substitute a model judge for ground truth, and can explain when a judge is admissible
- You report failure rates as plainly as success rates
Nice to have
- Red-teaming or adversarial evaluation
- Human evaluation study design
- You have built an eval suite that changed a shipping decision
The exercise
Design an evaluation that detects an agent becoming more persuasive without becoming more accurate or reliable.
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.