Free cookie consent management tool by TermsFeed Generator MTS Environments / Evals — Frontier AI | London / Paris | Axioma Search
Image
Image
Bg

Member of Technical Staff — Environments / Evals

About

AI models need useful problems to practise on if they're going to improve. This role is about creating those problems automatically — and then building evaluations that tell you whether the model has genuinely learned something.

The company is building AI systems that learn how to carry out complex work inside large organisations. They recreate real-world workflows as interactive training environments, then use those environments to train models through practice and feedback — so the models get better at completing long, multi-step tasks reliably, rather than simply generating answers.

You'll build systems that generate training tasks, adjust their difficulty as models improve, and simulate parts of real-world deployments. You'll also make sure improvements are real rather than models exploiting shortcuts in the training environment.

What you'll do

  • Build pipelines that automatically generate tasks and training environments
  • Create systems that filter tasks for validity, diversity and difficulty
  • Adjust training curricula based on where models succeed and fail
  • Train simulators using data from real model deployments
  • Test how closely simulated environments match real interactions
  • Build reproducible evaluation frameworks and held-out tests
  • Identify reward hacking, shortcuts, contamination and failures to generalise

What you'll need

  • Strong engineering skills and good experimental judgement
  • Experience with PyTorch and modern ML systems
  • Experience in synthetic data, agents, world models, curriculum learning or evaluation
  • Ability to design rigorous model evaluations
  • Strong understanding of model behaviour and failure analysis
  • Comfort working across open-ended research and engineering problems

Optional 

  • Ray, Docker, OpenEnv or Gymnasium
  • vLLM, SGLang, Playwright, Inspect AI or DeepEval

Shortlisted candidates will be contacted within 48 hours.

Back to job listings
  • Location London, Paris
  • Salary / Compensation Up to £160k + equity package
  • Sectors Agentic, Frontier AI / Foundation Models, GenAI
  • Skills PyTorch, Synthetic Data, Model Evaluation, Ray, vLLM, Curriculum Learning
Image

Role Contact

Alex Jouatte

Bg

Didn't find the right role?

Send us your CV.