Safeguards Enforcement Analyst, User Well-being
Anthropic · Safeguards (Trust & Safety)
- Pay
- $245,000–$285,000/yr
- Where
- San Francisco, CA | New York City, NY | Washington, DC
- Posted
- Aug 3
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
As a Safeguards Analyst on the User Well-being team, you will be focused on supporting the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones. Interventions range from steering how Claude responds in the conversation itself to in-product features that connect users to other resources.