💼Jobs 📝Blog 🧮Salary Calc 🌍Cost of Living 📋Tax Guide
👤Sign In Post a Job — from $99
Anthropic

Staff+ Site Reliability Engineer, Safeguards ML Infra

DevOpsFull-TimeLead
Location
Worldwide
Job Type
Full-Time
Experience
Lead
Apply Now

Job Description

ABOUT THE ROLE We are seeking a highly experienced Staff+ Site Reliability Engineer to join our Safeguards ML Infra team at Anthropic. As a key member of our team, you will play a critical role in designing, building, and operating the production infrastructure that powers Claude's safety systems. Your expertise will be essential in ensuring the safe and reliable deployment of safeguards across various platforms, including 1P, AWS Bedrock, GCP Vertex, and others. WHAT YOU'LL DO As a Staff+ Site Reliability Engineer, your primary responsibilities will include: - Launch captain model releases, ensuring that safeguards are properly configured and verified for every new model. - Owning the off-cycle deployment of new safety classifiers, including canarying rollouts, running post-deploy validations, and investigating discrepancies. - Verifying that the right safeguards are provably live on the right models across every deployment platform, detecting and eliminating configuration drift between them. - Automating last quarter's work by turning launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Building and maintaining a safeguards registry with full provenance, including what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participating in on-call and operational-duty rotations, covering service incidents, model provisioning, and time-sensitive research and safety launches. WHAT YOU'LL NEED To be successful in this role, you will need: - Deep experience in production change management at scale, including deploy pipelines, config management systems, canary analysis, and strong opinions about what "verified" means. - A track record of running high-stakes releases, serving as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences. - Meaningful on-call experience for production systems, including incident response and postmortem-driven improvements. - Hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale. - Proficiency in Python; experience with Rust is a plus but not required. WHY REMOTE As a remote employee, you will have the flexibility to work from anywhere and be part of a team that values collaboration, innovation, and safety. Our remote work setup allows for seamless communication and collaboration across different time zones and locations. BENEFITS - Annual salary: $405,000 - $485,000 USD - Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience - Required field of study: A field relevant to the role as demonstrated through coursework, training, or experience.