🏢
Anthropic
Safeguards Enforcement Analyst, Violence & Extremism
Job Description
ABOUT THE ROLE
As a Safeguards Enforcement Analyst focused on Violence & Extremism at Anthropic, you will play a critical role in ensuring the safety and reliability of our AI systems. Our mission is to create beneficial AI that aligns with societal values and prevents real-world harm. In this position, you will design and execute operational workflows to detect and mitigate misuse attempts, drive enforcement decisions, and develop evaluations across a range of policy areas.
WHAT YOU'LL DO
As a key member of our Safeguards team, your responsibilities will include:
- Designing and architecting automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
- Developing and maintaining evaluations that measure model performance on policy areas, surface regressions, and inform policy and model improvements
- Partnering with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
- Reviewing flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts and emerging extremist movements, ideologies, and mobilization tactics
- Supporting the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios
- Developing and maintaining enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content
- Staying up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices, and applying these to inform our workflows and evaluations
- Identifying and escalating emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity
WHAT YOU'LL NEED
To succeed in this role, you will need:
- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
- Experience standing up and scaling policy enforcement or content review workflows
- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
- Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
- Experience working with generative AI products, including writing effective prompts for content review and enforcement
- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space
WHY REMOTE
As a remote employee, you will have the flexibility to work from anywhere and maintain a healthy work-life balance. Our remote work environment allows you to collaborate with our team of experts from around the world and contribute to the development of beneficial AI systems.
BENEFITS
- Annual salary range: $285,000 - $330,000 USD
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the level of expertise and qualifications required for the role