About
Training models at scale creates a lot of infrastructure problems. Researchers need reliable access to GPUs, experiments need to be reproducible, and failures need to be easy to understand. You'll own the platform that makes that possible.
The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.
You'll work across GPU infrastructure, distributed execution, storage and observability. The aim is straightforward: make compute productive and give researchers simple tools to run, inspect and debug their work.
What you'll do
- Build and operate GPU clusters for model training and inference
- Manage job scheduling, networking and storage for distributed ML workloads
- Build systems that can run large numbers of training environments concurrently
- Improve GPU utilisation, startup times and overall platform reliability
- Build reliable pipelines for datasets, model weights and checkpoints
- Improve monitoring, debugging and failure recovery
- Give researchers simple, reproducible ways to run experiments
What you'll need
- Strong Linux and distributed systems fundamentals
- Experience with Docker, Kubernetes and/or Slurm
- Strong Python or Go
- Experience building reliable production infrastructure
- Good understanding of concurrency, networking and failure recovery
- Familiarity with GPU networking and NCCL
- Familiarity with Ray, Terraform, Prometheus, Grafana or OpenTelemetry
Shortlisted candidates will be contacted within 48 hours.