🏢
Anthropic
Pre-training Distributed Systems Tech Lead / Manager
Job Description
ABOUT THE ROLE
Anthropic is building reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. The Pre‑training Distributed Systems Tech Lead / Manager will head the Evals Infrastructure team, creating the systems that measure what our models can actually do. This role sits at the intersection of inference, research, and infrastructure engineering, managing large‑scale distributed systems that orchestrate evaluations for frontier models, building and scaling the harnesses researchers use, ensuring results are reproducible and interpretable, and making evaluation signals available to decision makers. Your work will directly influence what the company builds and releases.
WHAT YOU’LL DO
Lead the team that designs, implements, and operates the distributed systems that schedule, orchestrate, and execute evaluations for frontier model training. Own evaluation throughput and cost, managing compute allocation across suites, queueing against constrained accelerator pools, and caching and reusing evaluation work. Build and scale the harnesses researchers use to define, run, and iterate on evaluations. Ensure evaluation results are trustworthy through determinism, reproducibility, and honest uncertainty quantification. Deliver evaluation signals to dashboards and review processes that inform launch decisions. Contribute as an engineer while managing and growing the team, prioritizing work, and coaching reports.
WHAT YOU’LL NEED
You have led technical projects end‑to‑end on large‑scale distributed systems and have at least one year of experience managing engineers or serving as a tech lead with reports. You are strong in Python and Rust and have built high‑throughput, fault‑tolerant systems on cloud or on‑prem accelerator fleets. You care deeply about measurement quality, not just pipeline uptime, and can detect when a metric shifts for the wrong reason. You communicate effectively with researchers and translate research needs into infrastructure solutions. You are deeply interested in the transformative effects of advanced AI and committed to safe development.
Strong candidates may also have experience with LLM inference or training infrastructure, evaluation or benchmarking systems (especially agentic evaluations requiring sandboxed execution), working statistical literacy (variance, confidence intervals, sample‑size sufficiency for noisy metrics), and observability and regression detection over time‑series metrics.
WHY REMOTE
This position is fully remote, allowing you to work from anywhere while collaborating with a globally distributed team of researchers, engineers, policy experts, and business leaders.
BENEFITS
Annual salary range: $500,000 – $850,000 USD. The role includes health coverage, parental leave, education stipends, life insurance, equipment stipend, and a flexible schedule. Anthropic sponsors visas for qualified candidates and encourages applicants from all backgrounds to apply.