Responsibilities

Architect and optimize core reinforcement learning infrastructure, from clean training abstractions to distributed experiment management across GPU clusters. Help scale our systems to handle increasingly complex research workflows.
Design, implement, and test novel training environments, evaluations, and methodologies for reinforcement learning agents which push the state of the art for the next generation of models.
Drive performance improvements across our stack through profiling, optimization, and benchmarking. Implement efficient caching solutions and debug distributed systems to accelerate both training and evaluation workflows.
Collaborate across research and engineering teams to develop automated testing frameworks, design clean APIs, and build scalable infrastructure that accelerates AI research.

Team

Structure: The Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.5 and Opus 4.5. Our work spans several key areas:

Anthropic is hiring a Research Engineer, Machine Learning (Reinforcement Learning)

Responsibilities

Team

Invoice multiple clients effortlessly

Similar Jobs

Care Data Insights Manager

AI Implementation Engineer - Moveworks

Junior Data Scientist

Associate Principal Engineer, Big Data

Business Intelligence Lead

Data Scientist