About the role
You let people adapt models and run appropriately sized training jobs on resources they control, and you make those jobs reproducible, interruptible and accountable. Interruptible matters more than it sounds: this is somebody's own machine, and they get to want it back.
The work
Build training pipelines, checkpoints, optimizer-state handling, resource scheduling and distributed execution where the network supports it. Coordinate local inference with background training. Enforce dataset permissions and explicit limits before a job moves to remote compute.
What good looks like
In your first 90 days, deliver a recoverable local fine-tuning workflow and a documented boundary between supported local jobs and remote workloads.
Evidence we look for
Bring practical ML training systems experience and distributed-computing fundamentals. Understand memory accounting, numerical stability and the difference between fine-tuning and large-scale pretraining.
What we need to see
- Practical ML training systems experience plus distributed-computing fundamentals
- Memory accounting and numerical stability in practice
- You understand the difference between fine-tuning and large-scale training, and design for the first honestly
- You make runs reproducible, including the parts people usually leave out
Nice to have
- Parameter-efficient fine-tuning methods
- Checkpointing and preemption handling
- Federated or on-device training
The exercise
Plan a training run that must survive a power interruption while preserving the ability to compare results with the original baseline.
Where and how we work
In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.