Responsibilities
- Build and deploy state-of-the-art reinforcement learning pipelines at scale.
- Post-train large (multi-modal) models to align them with human intent and enable general capabilities such as reasoning, pushing the boundaries of model performance, safety, and efficiency
- Always keep the entire lifecycle of research and production in mind: from idea conception, theoretical modeling, prototyping, ablation studies, all the way to production deployment
- Build and foster external collaborations with academic and industrial partners
- Follow scientific and technical standards for experimentation, reproducibility, and model evaluation
- Collaborate deeply with Engineering, ML Platform, and HPC teams to deliver robust and reliable model updates to users
Requirements
- Deep technical background
- Strong leadership skills
- Proven track record of driving research in reinforcement learning or large-scale model alignment to production
- Strong practical background
- Creative mindset
- Passion for solving hard problems with real-world impact
- Solid mathematical background
- Enjoy solving challenging problems
- Masters degree, diploma, PhD, or equivalent industry experience in mathematics, physics, computer science, or a related field
- Deep practical experience in Python
- Deep practical experience in at least one modern machine learning framework such as PyTorch, TensorFlow, or JAX
- Track record of leading self-directed research projects that go well beyond academic exercises and deliver tangible results
Nice to Have
- Experience working with large compute clusters and ML infrastructure
- Expertise in deep reinforcement learning (RLHF/RLAIF/RLVR)
- Hands-on experience scaling and deploying LLMs or other foundation models in real-world systems
Benefits
- Diverse and internationally distributed team with people of more than 90 nationalities
- Open communication, regular feedback
- Hybrid work, flexible hours
- Monthly full-day hacking sessions (Hack Fridays)
- 30 days of annual leave (excluding public holidays)
- Access to mental health resources
- Competitive benefits tailored to location
Work Arrangement
Hybrid — UK, Germany, Netherlands, Poland, US, Japan
Additional Information
- Flexible working hours
- Hybrid schedule: office three times per week
- Mental health resources
- Global team collaboration across time zones