Full-Time

Staff Software Engineer

AI Reliability Engineering

Updated on 8/22/2026

Anthropic

Anthropic

5,001-10,000 employees

Develops reliable, interpretable AI systems

Compensation Overview

£325k - £390k/yr

London, UK

Hybrid

At least 25% of work time must be in an office; some roles may require more.

Bachelor's

Category
DevOps & Infrastructure (1)
Software Engineering (1)
Required Skills
LLM
Graphics Processing Unit (GPU)
Incident Response
Distributed Systems
Observability

Get referred to Anthropic

See people who can refer or advise you

Requirements
  • Strong background in distributed systems, infrastructure, or reliability engineering.
  • Ability to work comfortably in unfamiliar systems during incidents and help drive resolution.
  • Ability to think holistically about how systems compose and where system seams are located.
  • Ability to build lasting relationships and collaborate effectively across teams.
  • Experience building product stacks, scaling databases, or operating large-scale distributed systems.
  • A bachelor's degree or equivalent combination of education, training, and experience.
  • A field relevant to the role, demonstrated through coursework, training, or professional experience.
  • Years of experience correlating with the internal job level requirements for the position.
Responsibilities
  • Develop Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving.
Desired Qualifications
  • Experience as a Site Reliability Engineer, Production Engineer, or in a similar reliability-focused role on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure exceeding 1,000 GPUs.
  • Experience with machine learning hardware accelerators such as GPUs, TPUs, or Trainium.
  • Understanding of machine-learning-specific networking optimizations such as RDMA and InfiniBand.
  • Expertise in AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contributions to open-source infrastructure or machine learning tooling.

Anthropic focuses on AI research to build reliable, interpretable, and steerable AI systems. Its main product, Claude, is an AI assistant designed to handle tasks at any scale for clients across industries, delivered through deployment and licensing along with specialized AI R&D services. Claude works by combining natural language processing, human feedback, reinforcement learning, and interpretability techniques to produce a capable, controllable AI assistant that can assist with a wide range of tasks. The company differentiates itself from competitors by prioritizing safety, transparency, and controllability—emphasizing reliability, interpretability of model behavior, and user-controlled steerability in its AI systems. Anthropic’s goal is to make AI systems that people can trust and efficiently use to improve operations and decision-making across sectors.

Company Size

5,001-10,000

Company Stage

Debt Financing

Total Funding

$162.8B

Headquarters

San Francisco, California

Founded

2021

Get referred to Anthropic

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Reuters said July 2026 revenue run rate reached $65 billion.
  • Claude Code auto mode caught 89% of harmful actions in Anthropic testing.
  • Project Glasswing launched April 7, 2026 with $100 million in usage credits.

What critics are saying

  • Round Hill sued Anthropic on August 17, 2026 over 500 song lyrics.
  • Anthropic still faces UMG and BMG training lawsuits, risking billion-dollar damages.
  • An IPO tied to $190 billion to $200 billion 2028 revenue demands flawless execution.

What makes Anthropic unique

  • Claude Code auto mode became default on August 14, 2026.
  • Project Glasswing ties Anthropic to AWS, Apple, Google, Microsoft, and JPMorganChase.
  • Public Benefit Corporation structure and Long-Term Benefit Trust reinforce safety-first governance.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Paid Vacation

Parental Leave

Hybrid Work Options

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

9%

1 year growth

6%

2 year growth

5%
PR Newswire
Aug 20th, 2026
MIT and Harvard students unveil AI tools to boost economic mobility for disadvantaged families

MIT and Harvard graduate students have developed artificial intelligence tools designed to improve economic outcomes for disadvantaged families through the inaugural AI for Social Impact Fellowship. The programme, supported by NextLadder Ventures, The Bike Shop @ MIT, and Anthropic, created "Navigation Technology" solutions that provide personalised support during high-stakes financial moments. Four frontline organisations are scaling the tools immediately. Money Management International is deploying call intelligence software for debt counsellors. Center for Employment Opportunities is integrating an AI voice coach for job interview preparation. Climb Together is adding features for networking conversation feedback. Neighborhood Trust Financial Partners is implementing technology to analyse bank statements during coaching calls. The fellowship aims to expand the pipeline of developers building Navigation Technology, which at maturity could help millions of Americans navigate economic challenges. NextLadder Ventures has over $1 billion in capital backing the initiative.

PR Newswire
Aug 20th, 2026
Cyberhill's Cerebro delivers context layer for Claude Enterprise, enabling deployment in days

Cyberhill Partners has launched Cerebro, a context and semantic layer for Anthropic's Claude Enterprise that enables AI models to understand business context rather than merely retrieve information. The system provides Claude Enterprise with structured understanding of an organisation's data, relationships, and business logic whilst keeping data in place. Cerebro allows enterprises to deploy context-aware AI in days instead of months. The platform offers traceability by showing reasoning paths, reduces token costs through logical mapping, and minimises hallucinations by aligning models with specific enterprise data. Originally developed for the US Intelligence Community, Cerebro works across multiple AI models. The Austin-based firm plans to extend Cerebro's capabilities to enterprise versions of ChatGPT, Gemini, and Grok in coming weeks. Cyberhill has completed over 1,000 enterprise software implementations and brings a decade of AI experience from US defence and intelligence work.

TechCrunch
Aug 19th, 2026
OpenAI previews Private Safety Processing to monitor AI misuse without retaining customer data

OpenAI has introduced Private Safety Processing, a new automated system that monitors for AI misuse whilst retaining no customer data. The service analyses multiple conversations for signs of abuse without human review, targeting bad actors who spread malicious requests across sessions to avoid detection. The move appears designed to compete with Anthropic, whose 30-day data retention policy for certain models has concerned enterprises handling sensitive information. OpenAI's system expands on Zero Data Retention protocols, which both companies largely follow. When triggered, the system sends a narrowly defined signal to OpenAI about specific suspicious activity. The company can then contact customers to discuss potential enforcement actions. Both AI firms are competing intensely for market share. Anthropic's annualised revenue reportedly reached $65 billion, whilst investors suggest it could IPO at $2 trillion valuation. OpenAI is also preparing for an IPO.

Bloomberg
Aug 19th, 2026
Anthropic-linked Texas AI data centre secures $1.3B private credit loan from Eagle Point

Anthropic-linked data centre secures $1.3 billion private credit loan from Eagle Point Credit Management to finance an AI facility in Texas. The deal represents one of the latest major financings supporting the artificial intelligence industry's rapid expansion. Eagle Point Credit Management is providing the loan for the sprawling data centre, marking a significant role for the investment firm in AI infrastructure development. The transaction reflects growing investor appetite for backing physical infrastructure needed to support AI operations as the sector continues its boom.

Yahoo Finance
Aug 19th, 2026
OpenAI Q2 revenue hits $6.7B but trails Anthropic's $11.6B as operating loss widens to $12.3B

OpenAI's revenue grew 18% to $6.7 billion in Q2, falling short of investor expectations as rival Anthropic surged ahead, The Wall Street Journal reports. Anthropic's revenue more than doubled to $11.6 billion, overtaking OpenAI for the first time whilst posting a small operating profit. OpenAI's operating loss widened to $12.3 billion from $9.3 billion in Q1, moving the company further from profitability. The contrasting results suggest a potential shift in momentum, with Anthropic's Claude Code gaining traction amongst developers whilst ChatGPT growth slows. The results increase pressure on OpenAI to refine its strategy ahead of a possible IPO. The company has made substantial computing commitments based on expectations of eventually generating hundreds of billions in annual revenue.