Technical team members are expected in the Berkeley office 3–5 days per week; in-person work is strongly preferred.
METR is a nonprofit research institute that tests frontier AI models for capabilities that could pose catastrophic risks before they are released, by partnering with leading AI companies to gain early access to models and conduct evaluations. Its work uses red-teaming and capability assessments to probe long-horizon, agentic tasks such as autonomous replication, rapid research and development, and cyberattacks, while not accepting compensation for its testing. The findings inform risk assessment methods and safety policies for developers and policymakers, and METR contributes to governance efforts by supporting frameworks like OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy. Its goal is to provide independent technical evaluation and threat research to improve AI safety and reduce the chance that dangerous capabilities enter real-world use.
Company Size
51-200
Company Stage
Grant
Total Funding
$71M
Headquarters
Berkeley, California
Founded
2022
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Flexible Work Hours
Hybrid Work Options
401(k) Retirement Plan
401(k) Company Match
Wellness Program
Mental Health Support
Conference Attendance Budget
Professional Development Budget
Stock Options
Company Equity
Phone/Internet Stipend
Home Office Stipend
Parental Leave
Family Planning Benefits
Fertility Treatment Support
Adoption Assistance
Childcare Support
Paid Vacation
Paid Sick Leave
Paid Holidays
Remote Work Options
Health Insurance
Dental Insurance
Vision Insurance
Life Insurance
Disability Insurance
Tuition Reimbursement
Professional Certification Support
Mentorship Program
Employee Discounts
Employee Referral Bonus
Relocation Assistance
Meal Benefits
Legal Services
Gym Membership
Commuter Benefits
Sabbatical Leave
Performance Bonus
Profit Sharing
Employee Stock Purchase Plan
Adoption Assistance
Amazon's $190 billion Anthropic stake linked to Sam Bankman-Fried's AI Ventures. Likely Real 42 votes Updated 4 hours ago Amazon is sitting on a $190 billion stake in Anthropic. And buried inside that number is a connection most people haven't looked at closely enough. The investment breaks down into roughly $98 billion in convertible notes and $92 billion in non-voting preferred stock. Big money, clearly. But the story gets messier when you start pulling on the threads connecting Anthropic, Amazon, and a network of effective altruism figures - some of whom trace back, directly or indirectly, to Sam Bankman-Fried, the disgraced founder of FTX. Bankman-Fried's Fingerprints on AI Safety. Before FTX collapsed, Bankman-Fried wasn't just running a crypto exchange. He was funding things. One of those things was the Model Evaluation and Threat Research organization, known as METR - an AI safety evaluator that now works with some of the biggest names in the industry. METR got $1.25 million from the FTX Foundation, which was Bankman-Fried's effective altruism charity. After FTX went bankrupt, METR gave the money back. They cited moral concerns. Fair enough. But the original funding link is still there. Digital Currencies METR didn't come out of nowhere. It grew out of the Alignment Research Center, which was led by Paul Christiano. Christiano wasn't a stranger to Anthropic's orbit - he was a former associate of Dario Amodei, Anthropic's CEO. The two reportedly shared a residence at one point. Christiano also recruited METR's CEO, Beth Barnes. Several METR members have ties to the University of Oxford, which has long been a hub for effective altruism thinking. Discover more Finance News Financial Markets News So you've got an AI safety evaluator that was seeded by Bankman-Fried money, led by someone who came up alongside Anthropic's CEO, and staffed partly by people from the effective altruism world. That's a pretty tangled web for an organization meant to provide independent safety checks. Amazon brought METR in for a pilot review. Amazon ran a 2025 pilot review of its in-house AI model using METR. That's not a small thing. METR's whole job is to flag AI systems that might be developing dangerous capabilities - specifically what they call "recursive" self-coding abilities, where a model could theoretically rewrite and improve itself without human oversight. If that sounds like science fiction, it's not really. It's a concern serious AI researchers take seriously, and it's why safety evaluators exist. But the independence question keeps coming up. Daniela Amodei, Anthropic's president, is deeply tied to the effective altruism movement. She's married to Holden Karnofsky, who is both an Anthropic staffer and a co-founder of effective altruism initiatives including GiveWell and Coefficient Giving. Coefficient Giving, along with Good Ventures Foundation, also provides financial backing to METR. So the money flowing into METR comes partly from people who are also running Anthropic. That's a conflict of interest worth naming plainly. Discover more Finance News Merchant Services & Payment Systems David Sacks, a former AI advisor, has raised concerns about METR's independence. U.S. officials more broadly have expressed skepticism about whether effective altruism's principles align with American national interests. That friction probably won't go away quietly. The network behind the numbers. It's worth slowing down on the personal connections here because they're unusually dense. Dario Amodei and Paul Christiano lived together. Christiano built METR and brought in its leadership. Daniela Amodei runs Anthropic's operations and is married to a man who co-founded the foundations now funding METR. The evaluator checking Anthropic's safety is financially supported by people inside Anthropic's own leadership circle. None of that is necessarily corrupt. People who care about AI safety tend to know each other - it's a small world. But when the stakes are $190 billion and the evaluations are supposed to be independent, the closeness matters. Amazon's commitment to Anthropic keeps growing as Anthropic's valuation climbs. That's not surprising - the company has been aggressive in its AI infrastructure bets. But the due diligence question lingers. How much does Amazon know about the METR relationship? How much does it care? The pilot review happened. The investment kept growing. METR returned Bankman-Fried's $1.25 million. Good Ventures Foundation and Coefficient Giving stepped in. The evaluations continue. Frequently asked questions. What exactly is Amazon's stake in Anthropic worth? How is Sam Bankman-Fried connected to Anthropic's AI safety evaluator? Why it matters. This investment by Amazon highlights the intertwining of major tech firms with figures and ideologies linked to effective altruism, a movement that has faced scrutiny following the collapse of FTX and Sam Bankman-Fried's legal troubles. The connection raises questions about the future of ethical investing in the tech sector and the potential risks associated with backing entities that are tied to controversial figures. As the cryptocurrency and AI landscapes continue to evolve, this relationship could influence investor sentiment and regulatory scrutiny in both arenas. Digital Currencies Community Trust Index High Confidence Real79% 21%Fake 42 community signals Post Views: 36
Anthropic researcher quits over AI doomsday fears - then two founders reveal a bold solution to stop rogue AI agents in enterprises. One day after Anthropic researcher Jacob Coxon resigned over fears that AI might end human life within the decade, I sat down with Rune Kvist and Rajiv Dattani, founders and brothers-in-law. They believe they've found a way to protect Ai Finder Guru - or at minimum, to keep AI agents from turning rogue within corporate environments. "AI is getting smarter at an increasingly rapid rate. The surprising thing about AI is that it becomes harder to adopt and harder to control as AI gets smarter, not easier," said Kvist, who joined Anthropic early on and is married to Dattani's sister. Dattani previously served as COO of METR, an AI safety research organization. Together they've launched the Artificial Intelligence Underwriting Company (AIUC), aiming to bring AI safety measures to enterprises and the companies developing AI models and agents. The startup lists Cursor, Lovable, Harvey, and ElevenLabs among its customers. On Tuesday, AIUC revealed a $40 million Series A round led by Ribbit Capital, with First Harmonic also participating. The company had previously secured $15 million in seed funding from Nat Friedman through his NFDG fund, alongside Emergence, Terrain, and Anthropic co-founder Ben Mann, among others. This brings total funding to $55 million. What drew this impressive group of investors was AIUC's approach: applying a well-established cybersecurity framework to a fresh set of AI risks. The company has created a third-party audit and certification process for AI agents. "Banks, hospitals, governments and militaries no longer decline to deploy AI because a model isn't smart enough," Kvist explained. "They decline because they've made commitments to their own customers about what a system will and won't do, and nobody can currently guarantee that." Taking inspiration from SOC 2, a widely adopted cybersecurity standard, AIUC developed its own standard called AIUC-1, along with a testing service to validate agents against it. To create the standard, AIUC brought together a consortium of roughly 250 security and risk leaders - the people who purchase agents. "These are the people who we meet with on a monthly basis, and the question we ask them is: When you're buying agents from someone, what would you look for. What are the questions you'd want to ask, and what would you want to see addressed." That input directly informs the tests. The startup then puts an agent through approximately 5,000 tests to observe its behavior in scenarios involving jailbreaks, hallucinations, and data leaks. The outcome is a report of roughly 100 pages detailing where an agent performs safely and reliably - and where it falls short. Notably, AIUC relies on AI agents to conduct the tests and AI to analyze the results, though humans verify the final audit, according to Kvist. If this sounds somewhat familiar, that's because it is. Dattani's former employer METR - where he served as COO from 2024 to 2025 and still sits on the board - conducts similar testing for frontier labs, though its work until recently centered mainly on performance, specifically whether agents can reliably complete tasks. METR was among the independent research organizations OpenAI enlisted to investigate its Hugging Face incident. Anthropic CEO Dario Amodei has also recently urged the AI industry to slow frontier development, pointing to a sharp rise in bad-behavior incidents. In his post, Amodei suggested requiring frontier labs to use embedded third-party evaluators to observe and verify safety, mentioning METR as one option. While AIUC doesn't propose embedding itself at customer sites, the underlying concept is comparable: providing enterprises with an independent assessment of how safe their AI agents truly are. "Here's where it passes and where you can trust it. And here's where there's concerns. You should be aware of those references before you make the decision to buy," Dattani said.
Two AI researchers leave Anthropic and Google over safety: 'There are no adults in the room' Jared Perlo Thu, September 10, 2026 at 3:33 PM PDT Citing fears that AI systems may soon spiral out of human control and potentially kill all humans, two more researchers from Anthropic and Google DeepMind who recently left their coveted positions are sounding the alarm about risks from advanced AI systems. Joe Benton, who used to lead a safety research team at Anthropic, and Josh Engels, who used to work on AI safety research at Google, told NBC News in their first interviews since they left that they see an urgent need to boost transparency about incidents at the cutting edge of AI given the rapid pace of AI development. Advances in AI research "could speed up the pace of progress from merely blistering at the minute to uncontrollable" rates of development, Benton said in an interview with Tom Llamas. "There are no adults in the room," Engels added. "People are trying their best, but there is no one coming to save us." Benton and Engels spoke with NBC News in the wake of a viral social media post from former Anthropic researcher Jacob Coxon, who left Anthropic on Tuesday. In his post on X announcing his departure, Coxon highlighted his extreme concern about the pace of AI development and the risks he believes it poses to the future of humanity. The post has been viewed more than 155 million times, spurring calls from legislators to hold special sessions of Congress to take action and sparking a wave of AI employees to speak out in support of his concerns. Benton and Engels both pointed to the recent cyberattack against AI startup Hugging Face, carried out in July by autonomous AI systems powered by an unreleased OpenAI model, as part of the reason for shifting their work now. "If you look at some of the recent incidents, these were not cases where humans told the models to do something bad," Engels said. Instead, OpenAI's AI systems autonomously decided to hack into Hugging Face's systems, create a sort of illicit message board to exchange information and even expose some of OpenAI's own computing infrastructure to the open internet. "The models decided that the best way to to accomplish their task was to commit really egregious actions, to commit crimes," Engels said in an interview with Christine Romans. OpenAI said that it has since strengthened its safeguards and that newer public models, including its most recent Astra system, more reliably follow human instructions. An Anthropic spokesperson said in a statement Wednesday: "We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry." Benton managed a group at Anthropic dedicated to creating ways for humans and weaker AI systems to supervise more capable AI systems. He said that he is particularly worried that the public does not have significant insight into how AI systems have already exceeded the bounds of human instructions and that the lack of transparency could only get worse as systems become more capable. "At the minute, basically all of the transparency about these risks that is coming from the companies is entirely voluntary," Benton told NBC News. In a blog post released Wednesday, OpenAI's head of global affairs, Chris Lehane, agreed that the status quo is insufficient. "Today, frontier laboratories largely set their own rules for managing frontier risks," Lehane wrote. "Democratically accountable standards, independent verification, and meaningful transparency would replace that fragmented system of private governance." No federal law mandates that the largest AI companies, like OpenAI and Anthropic, share reports when agents or AI systems act beyond humans' control. Benton and Engels are joining METR, one of the world's leading AI safety nonprofit research centers, to work on investigations into incidents or episodes in which AI strays from human directions or intentions. METR aims to create scientific ways to evaluate how AI systems could cause catastrophic risks and empower researchers and the public to sway their development. "I left because I think I can have more positive influence on the development of this technology by helping to foster public transparency from outside these companies and to shed light on the risks," Benton said. The researchers joined growing numbers of AI safety researchers raising awareness about today's AI systems following Coxon's post Tuesday. Marcus Williams, an OpenAI employee who works on monitoring the activity of AI agents, wrote Thursday afternoon on X, "Unless there is AI regulation or a coordinated slowdown between labs, human extinction in the next few years seems very likely." Geoffrey Irving, who was the chief scientist at the United Kingdom's AI Security Institute and served stints at Google and OpenAI, seemed to agree Wednesday on X. "I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years." AI researchers have referenced the possibility that AI systems could potentially begin to improve themselves without human input, creating the potential for AI systems to either purposely or incidentally kill humans. Benton said he was most worried that the pace of technological improvement could lead to this sort of digital or artificial superintelligence, in which AI systems are more capable than humans across most tasks. "All of these companies - and this is something I witnessed firsthand at Anthropic - are pretty directly trying to race towards automating the process of AI R&D itself," he said. Benton suggested that there's a chance AI systems could come to exist as a sort of separate species before long. "We're probably going to go from a world where we have very capable systems now to a world where potentially we are co-inhabiting a world with AI agents that are much, much smarter than humans at some point in the next few years," he said. Engels said many people outside the AI industry might not fully appreciate just how intelligent today's AI systems already are - and how capable they might soon become. "We're building these systems that are generally intelligent," he said. "They can generally do what people can do, and soon they might be able to generally do what people can do, but better." Both researchers said they were excited to shift their attention to more public efforts to highlight the latest AI progress, so people can make more informed decisions. "There are many reasons people are excited and racing forward, Engels said, citing huge potential upsides and benefits from AI. But he added that "we should progress as society aware of the risks and okay with where they're at." "I am worried that stuff might end up progressing too fast for us to get our act together in time," Benton said, "unless we worry about it now."
METR AI API key theft leads to $600K credit fraud. Attackers stole METR AI API keys via a fail-open bug, causing $600K in fraudulent AI credit consumption, revealing critical API security risks. The nonprofit AI evaluator METR recently disclosed a major security incident involving the theft of an API key that enabled attackers to illicitly consume approximately $600,000 in AI service credits over a three-week period. This breach, stemming from a fail-open authentication vulnerability in METR's public cloud infrastructure, highlights critical gaps in API key management, monitoring, and cloud security within AI-focused organizations. API vulnerabilities drive significant fraudulent credit consumption. According to multiple reports, attackers capitalized on a misconfigured public Amazon EC2 instance running METR's AI model inference API. This fail-open bug disabled authentication checks entirely, exposing the API key publicly. With this key, attackers were able to submit inference requests that consumed free but limited cloud computing credits provided for METR's research, amassing a staggering $600,000 charge over nearly a month before detection. The attackers further escalated their foothold by obtaining persistent access through an SSH key, presenting ongoing risks to METR's infrastructure. The fraudulent activity largely blended into normal usage patterns, aided by METR's error-tolerant API design and the presence of free usage credits, which delayed anomaly detection. Secondary attacks reveal broader supply chain security concerns. Following the API key theft, METR faced another wave of probing attacks including credential stuffing, phishing attempts, OAuth token exploitation, and the exposure of unpublished evaluation data due to a separate SQL query endpoint vulnerability. Although there is no evidence of unauthorized exposure of private or sensitive data, these attempts underscored deficiencies in endpoint security and cloud infrastructure hardening. These incidents emphasize the elevated risks facing AI supply chain evaluators and research organizations, which rely heavily on API-accessible models and cloud services as part of their operational fabric. Mitigation efforts and lessons for cybersecurity teams. In response, METR undertook multiple remediation actions: isolating public-facing applications from internal infrastructure, rotating and revoking compromised credentials, shutting down legacy systems, and expanding security staffing. Enhanced monitoring and spend limit alerts were implemented to better detect anomalous API usage and prevent similar large-scale credit consumption in the future. For cybersecurity leaders and analysts, this case serves as a potent example of how fail-open vulnerabilities and insufficient API governance can lead to high-impact financial damage and persistent threat actor presence. It underscores the need for: * Rigorous API authentication and authorization controls * Comprehensive monitoring including spend and usage anomaly detection * Segmentation of public and internal environments * Proactive credential rotation and secret management As AI research organizations increasingly depend on cloud-based APIs, these lessons are critical to safeguarding intellectual property and controlling operational costs. Conclusion. The METR API key theft and subsequent $600,000 fraudulent credit consumption incident provide a cautionary tale about the cybersecurity risks in AI research infrastructure. Despite METR's security maturity and partnerships with major AI vendors, a single misconfiguration exposed them to prolonged attack and significant financial loss. This breach illustrates how attackers are adept at exploiting overlooked API vulnerabilities and the importance of layered defenses including strict access controls, continuous monitoring, and rapid incident response. Cybersecurity professionals working with AI services must recognize that API security and cloud governance are integral to protecting both data and costly computational resources in this evolving threat environment.
Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks. The model provider gave METR the credits for free. An actual customer would not have been so lucky Published Tue 01 Sept 2026 // 20:45 UTC AI model testing organization METR has disclosed two attacks that happened earlier this year, including one in which an attacker stole an API key and spent three weeks consuming public-model credits worth about $600,000. METR (short for Model Evaluation and Threat Research) found no evidence that the attackers accessed sensitive information in either incident, and the org said it investigated both with security experts. METR researchers worked with OpenAI to investigate how its agents hacked Hugging Face, and on Monday, it disclosed two of its own security snafus. "In March 2026, attackers stole an API key for inference on public models and consumed a substantial amount of credits," the nonprofit disclosed in a Monday report. "In May 2026, we observed attackers systematically probing our publicly accessible infrastructure, including an unsuccessful attempt to access internal data via an inadvertently exposed endpoint." From fail-open bug to model-credit theft. The March incident involved a METR researcher who didn't have access to sensitive information - including model data and credentials, as well as information about model architectures, training, and release dates. The researcher used agents running on a personal EC2 instance that was "intentionally" left publicly accessible behind Google authentication. The instance contained an API key for METR's public models account. According to METR's account, a "vibe-coded app" included a fail-open bug that disabled authentication, and this exposed the system to the public internet for several days. "We suspect that the attacker found the instance by looking through recently-registered websites (e.g. in certificate transparency lists) to find vibe-coded sites with high-signal keywords relating to LLMs or agents, for purposes of harvesting potentially exposed model provider API keys," the AI research org wrote. Once the attacker found the app, they prompted an agent to reveal its model provider API key, then added an SSH key to maintain persistent access, and over the next three weeks used the stolen credentials to consume API credits on public models worth about $600,000. Luckily for METR, the unnamed model developer had given the credits to the nonprofit for free. How do you not notice the 'large illicit usage?' METR does answer the question on everyone's mind in the report: Why its researchers didn't notice the "large illicit usage?" There are several reasons for this. First, the model testing operation regularly runs evaluations that use a lot of tokens, and this means the organization is "very acclimated to getting lots of weird rate limit and API errors." So the high usage didn't look that out of the ordinary. Plus, since the tokens were free, METR didn't accrue a large bill, and at the time there was no way to put a spending limit on keys like the one that was stolen. In response to the March incident, METR says it improved its security infrastructure, protocols, and review process, and will continue to invest in security. To this end, it also hired a security lead, and plans to add more security staff. Crims used agents to try to access frontier models. The second incident happened in early May, when "METR became the target of a sustained external attack campaign." After being "tipped off" that attackers who appeared financially motivated may have been trying to gain illicit access to frontier models, METR watched the intruders probe its publicly accessible infrastructure. They also used agents to find ways to gain initial access, including automated vulnerability discovery, credential stuffing against authentication providers, attempting OAuth token grants, scanning newly deployed services, and phishing attempts. At the same time, METR unintentionally "exposed a read-only SQL query mechanism via our public transcript viewer." While queries were scoped to public data by default, a bug allowed access to unpublished evaluation data, and "some sensitive model data was accidentally included in this database." However, there's no evidence that the attacker found the exploit or accessed any non-public data, according to the model testing body. An independent bug hunter discovered the vulnerability and reported it to METR, which paid the researcher a bounty, and took the API offline. In response, METR says it now uses an isolated production environment for public-facing applications that is separate from its internal infrastructure.(R)