The board
Job
Business
Safeguards Enforcement Analyst, Violence & Extremism
Anthropic · Safeguards (Trust & Safety)
- Pay
- $285,000–$330,000/yr
- Where
- San Francisco, CA | New York City, NY | Washington, DC
- Posted
- Jul 15
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas.
What you'd do
- Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
- Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
- Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
- Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging extremist movements, ideologies, and mobilization tactics
- Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios
- Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content
- Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals
- Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity
What they want
- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
- Experience standing up and scaling policy enforcement or content review workflows
- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
- Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
- Experience working with generative AI products, including writing effective prompts for content review and enforcement
- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space