
Work Here?
Preparing a concise, high-school-friendly company summary for BrainTrust based on the provided description.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
51-200
Company Stage
Series B
Total Funding
$121.1M
Headquarters
San Francisco, California
Founded
2023
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$121.1M
Above
Industry Average
Funded Over
3 Rounds
Industry standards
Health Insurance
Dental Insurance
Vision Insurance
Unlimited Paid Time Off
Competitive salary and equity
AI Stipend
4 Levels of AI Agent Maturity: don't Build Slop. Ara Khan outlines a 4-level framework for building mature AI agents, emphasizing state machines, visualization, and cloud-native deployment to avoid "slop" and ensure scalability. In a recent talk titled "Don't Build Slop (4 Levels of AI Agent Maturity)", Ara Khan outlined a framework for developing robust and effective AI agents. The presentation, delivered at AI Engineer Europe and sponsored by Google DeepMind, Braintrust, and WorkOS, emphasized a structured approach to agent development, moving beyond ad-hoc solutions towards more mature and maintainable systems. Understanding AI Agent Maturity. Khan began by addressing a common pitfall in AI agent development: the tendency to create what he termed "slop." This arises from the factor of multi-agent orchestration, where the complexity of coordinating multiple agents can overwhelm even experienced engineers. Khan highlighted that many current AI agent flows from frontier models suffer from two main problems: inference bounds and take isolations.
AI firm Braintrust prompts API key rotation after data breach. Hackers accessed one of the company's AWS accounts and compromised AI provider secrets stored in Braintrust. | May 8, 2026 (7:14 AM ET) AI evaluation and observability platform Braintrust urged customers this week to rotate API keys that may have been compromised after hackers accessed an AWS account. The incident, the company says, was discovered on May 4, after receiving a report of suspicious behavior, and was communicated to customers via email on May 5. The message also included indicators of compromise (IOCs) and remediation steps. Immediately after learning of the incident, Braintrust locked down the compromised account, audited related systems and restricted access to them, rotated internal secrets, and launched an investigation into the matter. The internal AWS account used by its systems, Braintrust says, likely provided the attackers with access to API keys that organizations use to access AI models. "As a precaution, we recommend that all customers rotate any org-level AI provider keys used with Braintrust," the company said in an incident notice. According to the company, at least one customer has been affected by the incident, with three other customers reporting suspicious spikes in AI provider usage. "We have not identified broader customer exposure based on our investigation to date, but as a precaution we informed all org admins with stored AI provider secrets in Braintrust. The investigation is ongoing," the company says. Braintrust recommends that customers access their org-level settings page, delete or revoke the existing secrets, configure new secrets, and confirm that they were rotated by checking their timestamps. The org-level AI provider API keys potentially exposed in the incident were likely stored for AI-forward companies such as Box, Cloudflare, Dropbox, Notion, Ramp, Stripe, and others, Nudge Security CTO Jaime Blasco told SecurityWeek. "The blast radius isn't Braintrust, it's every downstream customer's AI stack, and a single SaaS compromise fans out across dozens of LLM provider accounts. This is the new shape of supply chain risk: every AI eval, observability, and gateway tool a company adopts becomes a credential warehouse, and those warehouses are now a tier-one target," Blasco said. Ionut Arghire is an international correspondent for SecurityWeek.
Braintrust breach exposes customer API keys in AWS incident. Key takeaways. * Braintrust confirmed unauthorized access to an AWS account containing customer API keys for cloud-based AI models * The company is asking every customer to rotate API keys stored with Braintrust, despite claiming only one customer was impacted * Security experts warn of potential downstream implications for AI companies relying on Braintrust's platform What happened. Braintrust, an AI evaluation startup valued at $800 million, has confirmed a security breach affecting customer API keys. The company disclosed that an attacker gained unauthorized access to one of its Amazon Web Services cloud accounts. That account contained API keys customers use to access cloud-based AI models. In an email sent to customers on Monday and seen by TechCrunch, Braintrust acknowledged the incident and urged immediate action. The company is asking every customer to rotate any API keys stored with the platform. "We've communicated with one impacted customer and to date have not found evidence of broader exposure." - Braintrust customer email Braintrust publicly disclosed the incident on its website Tuesday. The company said it has contained the incident, locked down the compromised account, audited and restricted access across related systems, and rotated internal secrets. Mixed messages on severity. Braintrust's public statements contain a notable contradiction. While the company confirmed unauthorized access and is asking all customers to rotate keys, spokesperson Martin Bergman told TechCrunch that there is "no evidence of a breach at this time." He said the company sent the email "out of an abundance of caution." The cause of the breach remains under investigation. Braintrust has not disclosed how the attacker gained access, how long they had access, or what specific data may have been exposed beyond customer API keys. Why this matters for AI companies. Braintrust provides a platform for companies to monitor AI models and products. CEO Ankur Goyal has described it as an "operating system for engineers building AI software." The startup raised $80 million in a Series B funding round in February 2026, reaching an $800 million valuation. The breach has implications beyond Braintrust's direct customers. Jaime Blasco, co-founder of cybersecurity startup Nudge Security, received a breach alert from Braintrust. He warned that the incident could have "downstream implications for affected customers," particularly AI companies that rely on Braintrust's services. API keys are prime targets for attackers. They provide direct access to cloud services, AI models, and sensitive data. Once stolen, attackers can use these keys to access customer systems, run up compute costs, or extract proprietary data. The keys Braintrust stores let customers access cloud-based AI models from providers like OpenAI, Anthropic, and others. Third-Party risk in the AI stack. This breach highlights a growing concern in the AI industry: supply chain risk. Companies building AI products often rely on multiple third-party services for model access, evaluation, monitoring, and deployment. Each service becomes a potential attack vector. Hackers frequently target corporate accounts on cloud services and third-party platforms. These services often store secrets like API keys, making them high-value targets. A single breach can cascade across an entire customer base. What Braintrust customers should do now. * Rotate all API keys stored with Braintrust immediately * Check logs for any unusual API activity during the exposure window * Review access patterns on connected AI model providers like OpenAI or Anthropic * Update keys in all production systems that use Braintrust-stored credentials * Enable additional monitoring on cloud accounts connected to Braintrust Companies should also consider whether to continue storing API keys with third-party platforms. Alternatives include using secrets managers with tighter access controls, or implementing just-in-time credential provisioning. Logicity's take. Frequently asked questions. What data was exposed in the Braintrust breach? Braintrust confirmed that customer API keys for accessing cloud-based AI models were stored in the compromised AWS account. The company has not disclosed the full scope of exposed data. Should I rotate my API keys if I use Braintrust? Yes. Braintrust is asking every customer to rotate any API keys stored with the platform, regardless of whether they've been notified of direct impact. How did attackers access Braintrust's AWS account? Braintrust has not disclosed the attack vector. The company says the cause of the breach is under investigation. Is Braintrust safe to use after the breach? Braintrust says it has contained the incident and locked down the compromised account. However, customers should make their own risk assessment based on their security requirements. Need help implementing this? Huma Shazia Senior AI & Tech Writer
Ameya Bhatawdekar on building AI evaluations at Braintrust. Michael Grinich interviews Ameya Bhatawdekar from Braintrust on AI evaluation, prompt engineering, and building reliable AI products at HumanX 2026. April 15, 2026 At HumanX 2026 in San Francisco, Michael Grinich sat down with Ameya Bhatawdekar from Braintrust to talk about one of the hardest problems in shipping AI products: knowing whether they actually work. The evaluation problem. Building AI features is the easy part relative to validating them. Knowing whether those features are reliable - across thousands of edge cases, user inputs, and real-world conditions - is where teams get stuck. Ameya Bhatawdekar leads work at Braintrust, a platform focused on AI evaluation and observability. The core challenge he sees: teams ship AI-powered features without a rigorous way to measure whether they're improving or regressing with each change. Traditional software testing doesn't map cleanly onto AI systems. When your output is probabilistic, deterministic unit tests alone aren't sufficient - the same input can produce different valid outputs. You need evaluation frameworks that account for the inherent variability of model outputs while still giving you confidence that your system is getting better over time. What braintrust is building. Braintrust provides tools for teams to evaluate, iterate on, and monitor their AI applications. The platform lets developers define evaluation criteria, run experiments against datasets, and track how changes to prompts, models, or retrieval pipelines affect output quality. Ameya's key point: evaluation isn't a one-time gate before deployment. It's a continuous process that runs alongside development. Teams that treat eval as an afterthought end up with AI features that degrade silently in production. Braintrust's approach includes: * Eval datasets - curated sets of inputs and expected outputs that represent real usage patterns * Scoring functions - automated and human-in-the-loop scoring to measure output quality * Experiment tracking - side-by-side comparison of how changes affect results across your entire eval suite Prompt engineering is engineering. One theme that came up in the conversation: prompt engineering is real engineering work, not a hack or a temporary workaround. Ameya emphasized that the teams getting the best results from AI treat prompt development with the same rigor they'd apply to any other part of their codebase. That means version control, systematic testing, and data-driven iteration - not ad hoc tweaking without measurement. This aligns with what Warrant see across the industry. The gap between a demo that works and a production system that's reliable is largely filled by evaluation infrastructure. Teams that invest in that infrastructure ship with more confidence because they have data on how changes affect output quality before those changes reach users. Shipping AI with confidence. Many teams still lack structured evaluation workflows. They ship changes to prompts or swap models without a systematic way to know whether things got better or worse. Braintrust's bet is that evaluation tooling becomes as fundamental to AI development as CI/CD is to traditional software. Just as you wouldn't merge code without running tests, you shouldn't deploy prompt changes without running evals. For teams building AI features today: invest in your evaluation pipeline early. The cost of building it upfront is concrete and bounded; the cost of debugging production regressions without evaluation data compounds over time. This interview was recorded at HumanX 2026 in San Francisco.
Braintrust raises $80M Series B to power AI observability. Published on Feb 20, 2026 Key takeaways. * Braintrust Data Inc. secured $80 million in Series B funding, led by ICONIQ Capital, with participation from Andreessen Horowitz, Greylock, Basecase Capital, and Elad Gil. The round values the company at $800 million, reflecting strong investor confidence in AI infrastructure platforms. * As enterprises integrate AI agents and large language models into mission critical workflows, structured evaluation frameworks have become essential. Braintrust delivers infrastructure that measures model performance, identifies hallucinations, detects data drift, and flags regressions before they affect end users. * The platform is already embedded within leading AI driven enterprises such as Notion, Replit, Cloudflare, Ramp, Dropbox, Vercel, Navan, and BILL. This adoption indicates increasing demand for continuous AI observability and production grade monitoring tools. * The newly raised capital will be allocated toward expanding engineering capabilities, strengthening go to market operations, establishing additional office locations, launching enhanced observability features, and entering new geographic markets. Quick recap. San Francisco-based Braintrust Data Inc. has officially announced the close of an $80 million Series B funding round, led by ICONIQ Capital at an $800 million post-money valuation. The round included returning backers Andreessen Horowitz, Greylock, Elad Gil, and Basecase Capital. The announcement was made via the company's official X (formerly Twitter) account, with CEO Ankur Goyal signaling that Braintrust is "building the infrastructure that helps teams measure, evaluate, and improve their AI products". Inside Braintrust's AI observability platform. Braintrust has built an AI-native observability and evaluation platform designed specifically for monitoring the quality of AI models and their outputs in production a fundamentally different challenge than traditional system-health monitoring. The platform integrates several critical workflows: * Exhaustive Tracing: Automatically captures every step of an AI model or agent's reasoning process, including prompts, tool calls, retrieved context, and metadata on latency and cost. * Automated Evaluation: Uses built-in scorers and an LLM-as-a-judge approach to evaluate model outputs for accuracy, relevance, and safety. Teams can run both offline experiments during development and online scoring on live production traffic. * Prompt Playground: A visual interface to test and version-control prompt changes against real production data before deployment. * AI-Powered Assistant: Analyzes millions of traces to suggest better prompts, create new datasets, and identify patterns that cause specific hallucination types. Critically, all of this runs on Brainstore, Braintrust's purpose-built database, which is reportedly 80% faster at querying complex AI traces than alternatives. This performance advantage is essential as enterprise AI deployments scale to millions of daily interactions. What leadership is saying? Matt Jacobson of ICONIQ noted that companies with enduring impact typically demonstrate strong and consistent customer focus. He stated that Ankur and the Braintrust team have embedded this principle into their product strategy from the outset, aligning development closely with evolving user requirements. Competitive landscape. The competitive intensity in AI observability is increasing at a measured but decisive pace. In February 2025, Arize AI secured $70 million in Series C funding to expand its large language model evaluation and monitoring capabilities. The round was positioned as one of the largest investments in the AI observability segment, reflecting growing enterprise demand for structured performance tracking and risk management across AI systems. At the same time, Langfuse, widely adopted within the developer community, was acquired by ClickHouse in January 2026 as part of a $400 million Series D financing at a $15 billion valuation. The transaction highlights how observability is moving beyond a developer focused capability and becoming a core component of enterprise grade AI infrastructure, supporting governance, reliability, and scalable deployment. Strategic analysis. Braintrust leads in developer experience and UI-driven evaluation workflows, making it the strongest choice for product and engineering teams that want an integrated, non-code-heavy approach to AI observability. Arize AI, with $131M in total funding and deep roots in traditional ML observability, holds the edge for large enterprises with complex, multi-model production environments. Langfuse, now backed by ClickHouse's $15 billion infrastructure, offers the most compelling option for teams that prioritize open-source flexibility and self-hosting. Bayelsa Watch's takeaway. I think this is a big deal, $800 million valuation at the Series B stage for an observability focused company indicates a structural shift in the AI ecosystem. The industry is moving beyond rapid model deployment toward ensuring models operate reliably, consistently, and within defined performance standards. Across the AI infrastructure landscape, capital is increasingly being allocated to accountability rather than experimentation. While many startups previously raised funding based on model capability claims, Braintrust's positioning centers on evaluation, transparency, and measurable outcomes. Add Bayelsa Watch as a Preferred Source on Google for instant updates! Sources. Pramod Pawar Pramod Pawar is the Founder of Bayelsa Watch and a digital entrepreneur behind multiple technology focused ventures. With 10+ years of experience in SEO and content strategy, he is known for converting complex research into clear statistics and practical insights. He holds a Bachelor of Engineering in Information Technology from Shivaji University, and his work is centered on AI, machine learning, big data analytics, and other emerging technologies. Coverage is frequently focused on fast moving areas such as AR, VR, robotics, cybersecurity, and next generation digital platforms, where trends are best understood through data. A strong focus is placed on accuracy, source checking, and simple explanations that support both general readers and business decision makers. Outside of work, cricket and reading across multiple genres are enjoyed, which helps new ideas and continuous learning remain part of his writing process. Statistics
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
51-200
Company Stage
Series B
Total Funding
$121.1M
Headquarters
San Francisco, California
Founded
2023
Find jobs on Simplify and start your career today