Safeguards Enforcement Analyst, Integrity & Authenticity
Anthropic · Safeguards (Trust & Safety)
- Pay
- $285,000–$330,000/yr
- Where
- San Francisco, CA | New York City, NY | Washington, DC
- Posted
- Jul 10
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
As a Safeguards Analyst focusing on Integrity & Authenticity, you will be responsible for building and executing enforcement workflows for our products and services, with a focus on detecting and mitigating attempts to misuse Anthropic's AI systems for coordinated inauthentic behavior, election manipulation, and targeting, tracking, and surveillance of individuals.
Your work will span a broad and interconnected set of harm areas: AI-enabled influence operations and disinformation campaigns, the abuse of AI to interfere with electoral processes, and the use of AI systems…
What you'd do
- Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
- Partner with Engineering and Data Science teams to optimize detection models for policy violations and automated enforcement systems
- Review flagged content to drive enforcement and policy improvements
- Enforce usage policies with a focus on detecting and mitigating AI-enabled influence operations, coordinated inauthentic behavior, election interference, and targeting, tracking, or surveillance of individuals and groups
- Support the Safeguards policy design team by providing detailed feedback on policy gaps based on real enforcement scenarios
- Keep up to date with emerging AI policy enforcement best practices, evolving threat actor tactics, and the regulatory landscape around elections, privacy, and surveillance, using these to inform our decision-making and workflows
What they want
- Experience in trust & safety, policy enforcement, threat intelligence, or a closely related field with a focus on one or more of: influence operations, disinformation, coordinated inauthentic behavior, election integrity, or privacy and surveillance harms
- Experience standing up and scaling policy enforcement or content review workflows
- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets
- Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
- Experience working with generative AI products, including writing effective prompts for content review and enforcement
- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space