🏢
Scale AI
Machine Learning Engineer, Global Public Sector
Job Description
ABOUT THE ROLE
Scale is a leading developer of reliable AI systems for the world's most important decisions. We are hiring ML Research Engineers to bridge the gap between emerging AI capabilities and mission-critical, real-world impact in our Global Public Sector (GPS) division. As a Research Engineer, you will lead the research into Agent Design, Reliability, and AI Safety, developing novel system architectures that power high-stakes government applications.
WHAT YOU'LL DO
As an ML Research Engineer, you will be responsible for designing and building agentic systems, driving reliability and safety, synthesizing deep research, optimizing models for niche domains, building evaluation frontiers, and consulting as a technical authority. Your key responsibilities will include:
- Architecting agentic systems, including designing agent architectures, harnesses, tool-use protocols, and logic flows that allow large language models (LLMs) to function as reliable, autonomous agents in complex workflows.
- Researching and implementing robust evaluation frameworks, including red-teaming for sovereign AI requirements and developing strategies to mitigate hallucinations in regulated data environments.
- Building agents capable of autonomous information synthesis and long-horizon reasoning, enabling users to analyze massive datasets and extract actionable insights.
- Evaluating and adapting models for specialized use cases, such as LLM reasoning for low-resource languages, complex OCR tasks, or working in GPU-constrained environments.
- Creating new, automated benchmarks that define what success looks like for AI in the public sector, ensuring our systems meet the highest standards of accuracy and sovereignty.
- Acting as a subject matter expert for public sector leaders, advising on the practical limits, safety requirements, and performance trade-offs of emerging AI technologies.
WHAT YOU'LL NEED
To be successful in this role, you will need to have exceptional proficiency in Python and experience building agentic harnesses or AI infrastructure. You should have a track record of taking theoretical AI concepts and turning them into functional prototypes or products, as well as experience in LLM benchmarking, red-teaming, or building evaluations that go beyond standard academic datasets. A Master’s or PhD in Computer Science, Mathematics, or a related field (with a focus on ML) is preferred, but we value demonstrated impact and engineering excellence.
WHY REMOTE
Scale is a remote-first company, and this role will be based entirely remotely. As a remote employee, you will have the flexibility to work from anywhere and will be part of a global team that is passionate about developing reliable AI systems for the world's most important decisions.
BENEFITS
Scale offers a competitive salary and benefits package, including health insurance, retirement savings, and paid time off. We are an equal opportunity employer and are committed to creating an inclusive and diverse work environment.