Full-Time

Safeguards Enforcement Analyst

User Well-being

Updated on 9/3/2026

Anthropic

Anthropic

5,001-10,000 employees

Develops reliable, interpretable AI systems

Compensation Overview

$245k - $285k/yr

H1B Sponsorship Available

Remote in USA + 3 more

More locations: Washington, DC, USA | San Francisco, CA, USA | New York, NY, USA

Hybrid

Staff must work from one of the listed offices at least 25% of the time.

Bachelor's

Category
Software Engineering
Required Skills
LLM
SQL
Quality Assurance (QA)
Data Analysis

Get referred to Anthropic

See people who can refer or advise you

Requirements
  • Experience in trust and safety, product policy, content moderation, or a related field, with direct exposure to mental health, suicide and self-harm, or related well-being harm areas.
  • Experience designing or running experiments, evaluations, or measurement studies to determine whether an intervention worked.
  • Experience translating policy definitions into measurable form, including rubrics, review guidelines, or classification criteria for human reviewers or automated systems.
  • Experience managing or coordinating content review operations, including quality assurance and workflow management.
  • Proficiency in SQL and/or other data analysis tools to measure intervention efficacy, monitor workflow health, and surface policy gaps.
  • Experience working with generative artificial intelligence products, including writing effective prompts for content review, classification, or evaluation.
  • Experience turning open questions and data into concise and insightful analysis.
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
  • Understanding of the challenges involved in implementing product policies at scale in the content moderation space.
  • Sound judgment in ambiguous, high-consequence cases, with the ability to make a decision and escalate appropriately when available signal is incomplete.
Responsibilities
  • Support the design and execution of interventions, define key metrics, and curate evaluation datasets.
  • Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems, including threshold-setting and precision and recall tradeoffs.
  • Monitor how interventions and detection systems perform over time.
  • Review flagged content to drive enforcement and policy improvements.
  • Support the development of in-product features that connect users to crisis resources, working with Product, Legal, and external partners on referral pathways and user-facing content.
  • Support the Safeguards Policy Design team by providing detailed feedback on policy gaps based on real scenarios.
  • Keep up to date with emerging AI policy and external research on AI's relationship to mental health, and use these to inform decision-making and workflows.
Desired Qualifications
  • Subject matter expertise in mental health, whether developed in academia, clinical practice, crisis intervention, trust and safety, or other related settings.
  • Experience building or evaluating large language model-based classification systems.
  • Experience using agentic tools, such as Claude Code, to scale analysis or automate recurring work.
  • Experience working within crisis support.

Anthropic focuses on AI research to build reliable, interpretable, and steerable AI systems. Its main product, Claude, is an AI assistant designed to handle tasks at any scale for clients across industries, delivered through deployment and licensing along with specialized AI R&D services. Claude works by combining natural language processing, human feedback, reinforcement learning, and interpretability techniques to produce a capable, controllable AI assistant that can assist with a wide range of tasks. The company differentiates itself from competitors by prioritizing safety, transparency, and controllability—emphasizing reliability, interpretability of model behavior, and user-controlled steerability in its AI systems. Anthropic’s goal is to make AI systems that people can trust and efficiently use to improve operations and decision-making across sectors.

Company Size

5,001-10,000

Company Stage

Debt Financing

Total Funding

$182.8B

Headquarters

San Francisco, California

Founded

2021

Get referred to Anthropic

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Project Glasswing found over 10,000 critical vulnerabilities by June 2, 2026.
  • Anthropic disclosed annualized revenue above $65 billion in July 2026, signaling explosive demand.
  • September 1, 2026 pricing cuts for cache reads boost agentic API adoption and retention.

What critics are saying

  • Anthropic’s $1.5 billion copyright settlement, approved July 20, 2026, invites more suits.
  • Late-September 2026 IPO pressure exposes weak multiples if growth decelerates after listing.
  • Heavy compute commitments and chip-lease debt create existential financing risk if demand softens.

What makes Anthropic unique

  • Claude Security and Project Glasswing anchor Anthropic’s enterprise security moat in 2026.
  • Anthropic pairs frontier-model capability with explicit safety branding, unlike OpenAI’s consumer-first posture.
  • Multi-cloud distribution across AWS, Google, Microsoft, Lambda, and Nscale reduces single-vendor dependence.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Paid Vacation

Parental Leave

Hybrid Work Options

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

7%

1 year growth

4%

2 year growth

2%
Yahoo Finance
Sep 9th, 2026
Broadcom eyes $40B Anthropic opportunity as Google chip risks weigh on stock

Broadcom could capture a $40 billion opportunity from Anthropic's growing compute needs, according to Macquarie analyst Arthur Lai. This comes as concerns mount over Google developing more chips internally, potentially threatening Broadcom's custom silicon business. The stock has fallen approximately 24% from its all-time high. However, Lai suggests many concerns may already be priced in, creating an attractive entry point. In April, Anthropic partnered with Google and Broadcom to secure next-generation TPU capacity for training its Claude AI models. Broadcom's AI semiconductor revenue surged 221% year-over-year to $16.7 billion in fiscal Q3 2026. Management projects AI semiconductor revenue could reach $115 billion in fiscal 2027 and potentially $230 billion in fiscal 2028. The Anthropic partnership could provide crucial revenue visibility whilst strengthening Broadcom's position in custom AI silicon and networking.

PR Newswire
Sep 8th, 2026
Black Duck joins Anthropic's Project Glasswing to secure critical software with AI

Black Duck has joined Anthropic's Project Glasswing, an industry initiative aimed at securing critical software infrastructure using advanced AI for defensive cybersecurity. The application security company will apply Mythos, Anthropic's AI system, across its full security portfolio. This will combine AI-accelerated vulnerability discovery with remediation workflows, risk-based prioritisation, and compliance-driven governance. "AI is transforming the economics and speed of vulnerability discovery and exploit development," said Dipto Chakravarty, Black Duck's Chief Product & Technology Officer. He explained that pairing Mythos with Black Duck's existing capabilities will enable faster risk reduction whilst maintaining the transparency and auditability required by enterprise security teams. Black Duck specialises in application security, combining deterministic analysis with AI reasoning to identify and fix security issues in code written by developers, generated by AI, or assembled from open source.

Yahoo Finance
Sep 8th, 2026
Goldman Sachs and Morgan Stanley push for OpenAI and Anthropic investment-grade ratings despite $20.9B losses

Goldman Sachs and Morgan Stanley have asked major credit rating agencies to grant investment-grade status to OpenAI and Anthropic upon going public, despite neither company turning a profit, the Financial Times reported. OpenAI posted a $20.9 billion operating loss on $13.1 billion revenue in 2025. Anthropic doesn't expect to break even until 2028, with OpenAI targeting 2030. The investment-grade designation would allow pension funds and insurers to buy their bonds. It would also terminate Nvidia's guarantee of up to $105 billion in lease obligations for OpenAI's Ohio campus. Rating analysts currently describe both labs as speculative-grade and loss-making. When SpaceX received investment-grade ratings after its June IPO, its bonds traded near junk pricing within days. Anthropic could list in late September, whilst OpenAI targets 2027.

Yahoo Finance
Sep 8th, 2026
Interactive Brokers earns interest on $182B of clients' idle cash — will Anthropic's IPO drain it?

Interactive Brokers held $182.4 billion in uninvested client cash at the end of June, up 27% year over year, and this figure grew to $185.6 billion by August. The automated global broker earns interest on this cash by investing it in short-term US government securities whilst paying clients a rate half a percentage point below the federal funds rate. Net interest income rose 23% year over year to $1.06 billion in the second quarter, representing more than half of total net revenues of $1.9 billion. The growth came from larger balances rather than margins, which actually narrowed to 1.93% from 2.07%. Anthropic's potential IPO, rumoured to arrive soon with a possible $2 trillion valuation, could provide clients with an opportunity to deploy some of this cash.

Yahoo Finance
Sep 7th, 2026
Anthropic signs $35B cloud deal with Lambda at Nvidia-leased Texas data centre

Anthropic has reportedly secured a $35 billion cloud deal with Lambda for 350 MW of capacity at Hut 8's Beacon Point campus in Texas, marking its ninth major compute corridor. The arrangement highlights Nvidia's dual role as both GPU supplier and data centre landlord, allowing it to extract value at multiple levels. The deal supports Anthropic's $65 billion annualised revenue run rate but deepens its reliance on Nvidia's ecosystem. Anthropic has diversified across nine corridors, including commitments to AWS (5GW), Google/Broadcom (5GW), Microsoft/Nvidia ($30 billion), Fluidstack ($50 billion), Nscale ($45 billion), Volta ($10 billion), AMD ($5 billion), and SpaceX (300MW). This infrastructure strategy reflects a shift where GPU suppliers increasingly control both hardware and physical environments, positioning themselves as compute landlords.