🏢
Anthropic
Research Engineer, Model Evaluations
Job Description
ABOUT THE ROLE
We're seeking a talented Research Engineer to join Anthropic's team. As a Research Engineer, you will play a critical role in building evaluations that help us understand and measure the capabilities of our AI system, Claude. Your work will focus on designing and implementing evaluations across various aspects of Claude's capabilities and personality, and building the infrastructure that runs them reliably at scale. You will partner closely with researchers throughout the lifecycle of a new capability, from defining what to measure to interpreting the results.
WHAT YOU'LL DO
As a Research Engineer at Anthropic, you will be responsible for:
Designing and running new evaluations of Claude's capabilities, including reasoning, agentic behavior, knowledge, and safety properties
Building and hardening the distributed eval execution platform to ensure hundreds of evals run reliably against checkpoints throughout production RL training runs
Developing dashboards that researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss
Debugging anomalous eval results mid-training-run, determining the cause, and communicating the answer clearly under time pressure
Improving the tooling, libraries, and workflows researchers use to implement and iterate on evaluations
Partnering with research teams across the full lifecycle of a new capability, from defining what to measure to interpreting results as training progresses
Running experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks
Communicating evaluations and their results to internal stakeholders and, where appropriate, external audiences
WHAT YOU'LL NEED
To be successful in this role, you will need:
Strong Python programming skills, including production or research infrastructure
Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale
Clear written and verbal communication, especially when explaining technical results to non-specialists
Comfort operating in an on-call or production-support capacity when training runs are live
Care about the societal impacts of your work and an interest in steering powerful AI to be safe and beneficial
WHY REMOTE
As a remote employee at Anthropic, you will have the flexibility to work from anywhere and enjoy a range of benefits, including a competitive salary, comprehensive benefits package, and opportunities for professional growth and development.
BENEFITS
Anthropic offers a competitive salary and comprehensive benefits package, including:
Competitive salary
Comprehensive benefits package
Opportunities for professional growth and development