About
This role is about making models better after their initial training. You'll work on reinforcement learning and other post-training methods, while also improving the systems needed to run those experiments efficiently at scale.
The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.
You'll work across both the learning algorithms and the infrastructure underneath them. That means going from an RL experiment to a GPU profiler trace, finding what's limiting performance, and making sure systems improvements don't change the way the model learns.
What you'll do
- Build and improve supervised fine-tuning, preference optimisation and RL methods
- Work with approaches including PPO, GRPO and SDPO
- Own training loops from rollout generation through to policy updates and checkpointing
- Improve training throughput, GPU utilisation and memory efficiency
- Profile and fix bottlenecks across distributed training
- Investigate instability and differences between training and inference
- Use real model failures to improve rewards, training data and overall performance
What you'll need
- Strong programming and quantitative problem-solving skills
- Hands-on experience with PyTorch and model training
- Understanding of reinforcement learning or LLM post-training
- Experience with distributed training and GPU systems
- Strong experimental judgement
- Ability to work comfortably across ML research and systems engineering
Optional
- vLLM, SGLang, Ray, FSDP or Megatron-LM
- CUDA, Triton, CuTE or GPU performance optimisation
Shortlisted candidates will be contacted within 48 hours.