🏢
Anthropic
Data Center Operations Lead - Partner Site Operations
Job Description
ABOUT THE ROLE
Anthropic is building AI systems that are reliable, interpretable, and steerable, and our Data Center Operations team is the backbone that keeps the compute fleet available and secure. As the Data Center Operations Lead for partner‑operated sites, you will own the end‑to‑end operational performance of multiple data halls, ensuring that deployment velocity, uptime, and incident response meet the highest standards. Your role bridges Anthropic’s engineering objectives with the day‑to‑day work of external vendors, setting priorities, defining processes, and driving continuous improvement across the fleet. You will be the primary point of contact for vendor teams, translating technical requirements into actionable plans while maintaining rigorous oversight of service level agreements and operational metrics.
WHAT YOU'LL DO
You will define and enforce the operational playbook for deployment, change management, security, and environmental, health, and safety compliance, and you will own the metrics that validate vendor performance. You will set daily and weekly priorities for on‑site teams, lead stand‑ups and business reviews, and provide tactical direction without directly managing staff. You will participate in the incident escalation on‑call rotation, serve as Incident Commander for site‑specific events, and own communications and post‑mortem processes. You will analyze failure patterns, identify root causes, and collaborate with engineering and vendor owners to implement corrective actions. You will develop and maintain scorecards and performance dashboards that track availability, repair turnaround, and deployment milestones, using independent data sources rather than vendor self‑reporting. Finally, you will guide the onboarding of new data halls, ensuring readiness of spares, security, and operational processes before first‑compute‑online.
WHAT YOU'LL NEED
You bring at least eight years of experience in data center operations, with a proven track record of managing vendors, MSPs, or contract workforces to measurable outcomes. You possess hands‑on technical depth in server, network, and rack‑level infrastructure, enabling you to audit vendor claims and verify quality. You have built or substantially improved operational processes, not merely run them, and you have led incident command or lead‑responder roles, communicating clearly under ambiguity. You can support non‑standard hours, including on‑call rotations and deployment surges, and you hold a bachelor’s degree or equivalent practical experience in a relevant field. Strong candidates will also have experience with third‑party colocation providers, commissioning new sites, GPU or high‑density liquid‑cooled infrastructure, multi‑vendor environments, and incident management frameworks or EHS programs.
WHY REMOTE
Anthropic embraces a distributed, asynchronous culture that empowers teams to work from anywhere while maintaining high collaboration standards. Our remote model offers flexible schedules that accommodate global time zones and the demands of a 24/7 data center environment. You will have the autonomy