← All open roles

H27 Β· AI infrastructure

AI Inference Systems Engineer

Make useful intelligence run responsively on a computer the user owns. You will turn models into a dependable local service.

First teamOffice-first9 cities

About the role

You turn models into a dependable local service, running responsively on a computer the person owns. Local inference is where the sovereignty claim is either true or marketing, and the constraint is unforgiving: fixed memory, no autoscaling, and a user who notices every pause.

The work

Build model serving, memory management, quantization evaluation, batching and scheduling. Support the chosen accelerator backends and measure startup, time to first token, sustained throughput and quality. Keep interactive tasks responsive while background work uses spare capacity.

What good looks like

In your first 90 days, ship a local inference service with a reproducible performance and quality report for the reference device.

Evidence we look for

Bring strong performance engineering and practical experience serving machine-learning models. Understand memory capacity, bandwidth, context length and why theoretical compute figures do not predict user experience.

What we need to see

  • Strong performance engineering plus practical experience serving machine-learning models
  • You reason about memory capacity, bandwidth and context length together rather than in isolation
  • You know why theoretical compute figures do not predict real throughput, and measure instead
  • You optimise against a latency budget somebody set for a user, not for a benchmark

Nice to have

  • Quantisation, speculative decoding, or KV-cache optimisation
  • Apple silicon or consumer GPU targets
  • You have shipped an inference stack somebody else ran

The exercise

Diagnose an inference slowdown as context grows and propose an improvement without quietly reducing output quality.

Where and how we work

In the office together five days a week, in any of these cities. Remote-friendly around your family, arranged one person at a time.