London Hybrid Full-time

DeepL is hiring a Senior Research Scientist FMTA

Responsibilities

  • Build and deploy state-of-the-art reinforcement learning pipelines at scale.
  • Post-train large (multi-modal) models to align them with human intent and enable general capabilities such as reasoning, pushing the boundaries of model performance, safety, and efficiency
  • Always keep the entire lifecycle of research and production in mind: from idea conception, theoretical modeling, prototyping, ablation studies, all the way to production deployment
  • Build and foster external collaborations with academic and industrial partners
  • Follow scientific and technical standards for experimentation, reproducibility, and model evaluation
  • Collaborate deeply with Engineering, ML Platform, and HPC teams to deliver robust and reliable model updates to users

Requirements

  • Deep technical background
  • Strong leadership skills
  • Proven track record of driving research in reinforcement learning or large-scale model alignment to production
  • Strong practical background
  • Creative mindset
  • Passion for solving hard problems with real-world impact
  • Solid mathematical background
  • Enjoy solving challenging problems
  • Masters degree, diploma, PhD, or equivalent industry experience in mathematics, physics, computer science, or a related field
  • Deep practical experience in Python
  • Deep practical experience in at least one modern machine learning framework such as PyTorch, TensorFlow, or JAX
  • Track record of leading self-directed research projects that go well beyond academic exercises and deliver tangible results

Nice to Have

  • Experience working with large compute clusters and ML infrastructure
  • Expertise in deep reinforcement learning (RLHF/RLAIF/RLVR)
  • Hands-on experience scaling and deploying LLMs or other foundation models in real-world systems

Benefits

  • Diverse and internationally distributed team with people of more than 90 nationalities
  • Open communication, regular feedback
  • Hybrid work, flexible hours
  • Monthly full-day hacking sessions (Hack Fridays)
  • 30 days of annual leave (excluding public holidays)
  • Access to mental health resources
  • Competitive benefits tailored to location

Work Arrangement

Hybrid — UK, Germany, Netherlands, Poland, US, Japan

Additional Information

  • Flexible working hours
  • Hybrid schedule: office three times per week
  • Mental health resources
  • Global team collaboration across time zones
Required Skills
MathematicsPythonDeep Reinforcement Learning
About company
DeepL
A global communications platform powered by Language AI that provides translations and intelligent writing suggestions for over 100,000 businesses worldwide.
All jobs at DeepL Visit website
Job Details
Department Research
Category Data & ML
Posted 2 days ago