💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Anthropic

Research Engineer, Code RL (Reinforcement Learning)

EngineeringFull-TimeMid-Level
Location
Worldwide
Job Type
Full-Time
Experience
Mid-Level
Apply Now

Job Description

ABOUT THE ROLE Anthropic is a rapidly growing organization dedicated to developing safe and beneficial AI systems. Our mission is to create reliable, interpretable, and steerable AI systems that positively impact users and society. As a Research Engineer on the Code RL team, you will play a critical role in advancing our AI systems, particularly in the areas of code generation, reinforcement learning, and large language models. WHAT YOU'LL DO As a Research Engineer, you will design and implement reinforcement learning environments and coding tasks, build reward signals and verifiers, and run training experiments on frontier models. You will also diagnose and improve the speed and reliability of pipelines that enable models to write, edit, test, debug, and ship real software. Your work will span several focus areas, including agentic coding behaviors, code correctness, long-horizon autonomous engineering, and high-performance code for accelerators. WHAT YOU'LL NEED To succeed in this role, you should have strong software-engineering skills, including deep expertise in Python, async, and concurrent programming. You should be comfortable owning systems end-to-end and debugging across the stack. Additionally, you should be able to balance research exploration with engineering implementation and engage in rigorous experimental design and result interpretation. You should also have a passion for code quality, testing, and performance, as well as a commitment to developing safe and beneficial AI systems. Strong candidates may also have experience with reinforcement learning, RLHF, post-training, or LLM finetuning, and a background in program analysis, testing, verification, compilers, or formal methods. WHY REMOTE As a remote employee, you will have the flexibility to work from anywhere and be part of a collaborative and dynamic team. You will have the opportunity to work on cutting-edge research and engineering projects, and contribute to the development of safe and beneficial AI systems. BENEFITS Anthropic offers a competitive annual salary range of $500,000 to $850,000 USD, as well as a comprehensive benefits package, including a health plan, parental leave, and education stipend. You will also have access to cutting-edge equipment and tools, and the opportunity to work with a talented and dedicated team.