Full-Time

Jailbreaking Lead

Red Team

Updated on 9/10/2026

FAR AI

FAR AI

51-200 employees

Non-profit AI safety research institute

Compensation Overview

$170k - $250k/yr

H1B Sponsorship Available

Remote in USA

Remote

Remote globally, with up to one trip per month for convenings, government meetings, or team gatherings.

Category
Cybersecurity (1)
Required Skills
LLM
Cybersecurity

Get referred to FAR AI

See people who can refer or advise you

Requirements
  • A track record of finding non-obvious, high-severity vulnerabilities in frontier AI systems, including universal or near-universal jailbreaks in heavily defended risk domains.
  • Deep, hands-on jailbreaking experience with modern frontier models and layered defenses, including chaining multiple attack techniques through defense-in-depth stacks.
  • Experience with black-box optimization methods, multimodal attacks, and/or agentic red-teaming.
  • Deep understanding of large language model architectures, training processes, and failure modes under adversarial conditions.
  • A strong technical track record in AI, adversarial machine learning, security, computer science, cybersecurity, mathematics, physics, or another highly technical subject.
  • Ability to operate effectively in rapidly evolving environments where techniques become obsolete quickly.
  • Demonstrated drive for mission impact and persistence in achieving ambitious goals.
Responsibilities
  • Personally develop universal and near-universal jailbreaks against frontier closed- and open-weight models across CBRNE, cyber, agentic security, extreme persuasion, and emerging risk domains.
  • Systematically dismantle defense-in-depth stacks, including input filters, model-level refusal and safe completion, reasoning monitors, output filters, and account-level moderation.
  • Escalate initial vulnerabilities to expose their most severe form and maximize universality, success rate, and elicited capability.
  • Own the technical bar for vulnerability severity and generality on major engagements.
  • Invent new attack classes when existing techniques fail and monitor and incorporate state-of-the-art methods from the literature.
  • Shape the jailbreaking research agenda and stress-test new affordances such as agents, tool use, long context, multimodal systems, and reasoning.
  • Set standards for rigor, creativity, and precision in jailbreaking across the red team.
  • Mentor individual contributors through pairing sessions, post-engagement retrospectives, and internal write-ups.
  • Review major red-teaming deliverables for technical quality, severity judgment, and clarity.
  • Work directly with frontier labs and government agencies so findings lead to mitigations.
  • Contribute to public reports, benchmarks, and the FAR.AI safety leaderboard.
  • Make calibrated technical judgments about universality, reliability, and what a capable threat actor could do with a finding.
Desired Qualifications
  • Prior collaboration with AI labs, security teams, or government safety institutes.
  • A track record in top CTF teams, offensive security research, or adversarial machine learning research.
  • Published work in AI safety, security, or robustness.
  • Ability to communicate technical findings and recommended mitigations to technical and non-technical audiences, including frontier lab safety teams and senior policymakers.
  • Prior experience mentoring technical individual contributors or leading a small technical team.

FAR AI is a non-profit research institute focused on making advanced artificial intelligence safer and more beneficial for society. It conducts in-house research on AI safety topics such as model evaluation, interpretability, and robustness, and it also supports safety-driven research through collaborations and targeted grants. A key activity is red-teaming frontier AI systems to identify vulnerabilities and inform safety standards for developers and governments, while it also builds community through FAR.Labs in Berkeley and global dialogues like the International Dialogue on AI Safety. Its funding comes mainly from philanthropy, with profits from for-profit AI work capped at 10% to preserve independence, and its goal is to steer powerful AI toward minimizing risk and maximizing societal benefit through research, governance discussions, and broad collaboration.

Company Size

51-200

Company Stage

N/A

Total Funding

N/A

Headquarters

Berkeley, California

Founded

2022

Get referred to FAR AI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • July 2026 Singapore office expands Asia-Pacific partnerships with IMDA, CSA, and NUS.
  • Over $30 million in 2025 commitments funds growth from 15 to 30-plus researchers.
  • Technical governance, fellowships, and workshops broaden hiring, reputation, and future grantmaking reach.

What critics are saying

  • Budget concentration around Coefficient Giving and peers creates donor dependence if priorities shift in 2027.
  • Leaderboard attacks can alienate frontier labs, shrinking paid red-teaming work and data access quickly.
  • If a major safety evaluation misses harmful capability leakage, FAR.AI loses credibility with regulators.

What makes FAR AI unique

  • Founded July 2022, FAR.AI combines frontier-model red-teaming, governance, and field-building from Berkeley.
  • January 2026 European Commission AI Office selection gives FAR.AI direct regulatory influence on CBRN safety.
  • Its AI Security Leaderboard benchmarks real jailbreak costs across major models, not abstract safety theory.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Remote Work Options

Hybrid Work Options

Paid Vacation

Paid Holidays

Sabbatical Leave

Flexible Work Hours

Wellness Program

Mental Health Support

Gym Membership

Phone/Internet Stipend

Home Office Stipend

Professional Development Budget

Conference Attendance Budget

Training Programs

Tuition Reimbursement

Professional Certification Support

Mentorship Program

Stock Options

Company Equity

Relocation Assistance

Adoption Assistance

Childcare Support

Elder Care Support

Parental Leave

Fertility Treatment Support

Family Planning Benefits

Employee Referral Bonus

Meal Benefits

Commuter Benefits

Legal Services

Employee Discounts

Company Social Events

Growth & Insights and Company News

Headcount

6 month growth

1%

1 year growth

1%

2 year growth

1%
PR Newswire
Jul 30th, 2026
FAR.AI opens first international office in Singapore to advance AI safety research

FAR.AI, a nonprofit AI safety research organisation, has opened its first international office in Singapore. The Singapore office will serve as FAR.AI's Asia Centre of Excellence, deepening collaboration on AI safety research across the Asia-Pacific region. FAR.AI has existing partnerships in Singapore with the Infocomm Media Development Authority (IMDA), the Cyber Security Agency of Singapore (CSA), and NUS. The organisation is currently working with IMDA to examine risks posed by advanced AI systems, including technical assessments of chatbot persuasive behaviour. Previously, FAR.AI collaborated with CSA on agentic AI security research, presenting the discussion paper "Securing Agentic AI" at Singapore International Cyber Week last year. The organisation has also built research ties with the NUS AI Institute. FAR.AI was founded in 2022 and has been selected by the European Commission's AI Office to lead technical safety research.

PR Newswire
Jul 29th, 2026
FAR.AI leaderboard reveals hundredfold gap in AI security: Grok, Gemini broken for under $300

AI security nonprofit FAR.AI has launched a leaderboard revealing a hundredfold gap in safeguards across frontier AI models. Testing found hundreds of "universal jailbreaks" in Grok 4.5 and Gemini 3.1 Pro, costing roughly $58 and $278 respectively to exploit. Claude Fable 5 and GPT-5.6 Sol showed no such vulnerabilities under identical testing, with exploitation costs exceeding $14,200. FAR.AI tested over 60 publicly documented jailbreak techniques across chemical, biological, radiological, nuclear, explosive, and cybersecurity threats. Grok 4.5 yielded 448 distinct universal jailbreaks whilst Gemini 3.1 Pro produced 249. The stronger models employed multiple independent protection layers. The organisation published a Minimal Standard for Safeguards alongside the leaderboard at leaderboard.far.ai, defining baseline attacks frontier models should withstand. FAR.AI shared findings confidentially with evaluated companies before publication.