Simplify Logo
Artificial Analysis

Artificial Analysis

Independent AI analysis and insights firm

Member of Technical Staff

Full-Time
No salary listed
Mid
San Francisco, CA, USA
In Person

San Francisco is preferred; other listed locations include Sydney, Melbourne, and Brisbane.

About the job

Requirements
  • At least 3 years of relevant professional experience.
  • Strong analytical and critical thinking skills.
  • Proficiency in Python and data analysis.
  • Genuine, demonstrable interest and knowledge of frontier artificial intelligence.
  • For strategy consulting backgrounds, the candidate should be able to structure ambiguous problems, build analytical frameworks, and communicate findings to senior stakeholders, with the ability to code and a strong technical interest in artificial intelligence.
  • For artificial intelligence and machine learning backgrounds, the candidate should have hands-on experience with modern artificial intelligence systems and understand how models work at a technical level.
  • For technical product management backgrounds, the candidate should have built and shipped artificial intelligence products in a fast-moving environment across research, engineering, product, and commercial work.
Responsibilities
  • Structure, design, and execute projects to evaluate artificial intelligence systems and technologies, including developing new artificial intelligence evaluation methodologies and datasets.
  • Develop reports and data visualizations to communicate complex artificial intelligence concepts to enterprises shaping their artificial intelligence strategy.
  • Collaborate with leading artificial intelligence companies to benchmark their technologies, including agentic artificial intelligence applications, models, and hardware.
  • Identify opportunities to enhance the artificial intelligence benchmarking platform and work with developers to implement them.
  • Use cutting-edge artificial intelligence tools in an artificial-intelligence-native workflow.
  • Contribute to company strategy and drive large initiatives supporting Artificial Analysis's goal of being the leading artificial intelligence benchmarking company.
  • Manage relationships with leading artificial intelligence labs and enterprise customers.
  • Evaluate major models across major capabilities as they are released and provide analysis that influences artificial intelligence development priorities.

About the company

Artificial Analysis provides ongoing, independent analysis of the artificial intelligence field. It produces reports and briefings about AI technologies, industry trends, company strategies, and policy/regulatory implications, using evidence-based research and data to inform decision-makers. Its product is a steady stream of analyses that help clients understand what is happening in AI and what it might mean for business, research, and society. What sets it apart is its independence from vendors or lobby groups and its backing by well-known AI leaders, including Nat Friedman, Daniel Gross, and Andrew Ng, which supports rigorous, credible insights. The company’s goal is to provide clear, trustworthy analysis that helps people navigate the fast-changing AI landscape and make better-informed decisions.

Company Size

11-50

Company Stage

Seed

Total Funding

$250K

Headquarters

Newark, Delaware

Founded

2024

Get referred to Artificial Analysis

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Optima launched August 13, 2026, turning proprietary customer workflows into paid benchmark demand.
  • The Search Index drew immediate attention because buyers lacked a neutral AI search API comparison.
  • IBM-backed ITBench-AA expands Artificial Analysis into enterprise operations, opening larger B2B budgets.

What critics are saying

  • Parallel, Exa, and Firecrawl sat within two points on August 18, 2026, crushing pricing power.
  • OpenAI, Anthropic, and Google can copy benchmark formats and bury Artificial Analysis features.
  • If model vendors and cloud labs internalize evaluations, Artificial Analysis becomes a commoditized leaderboard publisher.

What makes Artificial Analysis unique

  • Artificial Analysis fixes models and swaps providers, isolating retrieval quality from answer-model variance.
  • On August 13, 2026, Optima benchmarked companies on their own data and workflows.
  • Its IBM partnership on ITBench-AA targets agentic enterprise tasks, not generic chatbot leaderboards.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

29%

1 year growth

29%

2 year growth

46%
serp.fast
Aug 29th, 2026
Artificial Analysis benchmarked AI search APIs: what the Search Index does and doesn't tell you.

Artificial Analysis benchmarked AI search APIs: what the Search Index does and doesn't tell you. Artificial Analysis launched a neutral AI search API benchmark on August 18, 2026. What its Search Index measures, what it misses, and how to read it. Artificial Analysis launched its Search Index on August 18, 2026, the first neutral third-party AI search API benchmark. It scores 11 provider configurations across seven providers on a 0-100 scale by holding one answer model fixed and swapping only the search provider across three research and fact-finding tasks. For anyone building a RAG pipeline or an agent, this is the number that was missing. Until now, the only comparisons of Exa, Tavily, Parallel, Firecrawl, and their peers came from the vendors themselves, each drawn to a test where its own product wins. A group that spent two years turning model benchmarks into a shared reference point has now aimed the same method at search. That is worth understanding, and it is worth reading with some care, because the headline ranking is the least useful thing in it. What artificial Analysis actually shipped. The Search Index announcement covers 11 provider results across seven search providers, among them Parallel, Exa, Firecrawl, You.com, Tavily, and Keenable. Several providers appear more than once because they ship distinct modes, and a mode is really a separate product: Parallel's basic and advanced tiers, or Exa's auto configuration, behave differently enough that averaging them would hide the thing you are choosing between. Each provider gets a single 0-100 score, which is the equal-weighted mean of three component benchmarks. That score is the part people will screenshot. It is also the part that rewards the least attention, for reasons the rest of this piece works through. If you want the qualitative version of the same layer, its head-to-head of the AI search APIs walks the providers one at a time; the Index adds a measured axis on top of that. How the provider-swap design works (and why it's the smart part). The benchmark methodology is a provider-swap. Every task is answered by the same fixed model, GPT-5.6 Luna at medium settings, temperature 0.6, medium reasoning effort. The only thing that changes between runs is which search provider feeds that model. So the score is not "how good is this provider's own answer." It is "how much did this provider's results improve a fixed model's answer against a fixed task." That isolates the search contribution from model variance, which is the most common way vendor comparisons mislead: swap the model, watch the answer improve, then credit the search tool. The three components test different things. DeepSearchQA is 900 broad research questions that need many searches and resolve to list-style answers, graded by F1 so both misses and spurious additions cost you. BrowseComp is a 200-task subset of hard-to-find facts that need multi-hop browsing, scored on exact-answer accuracy. AA-Omniscience is 600 private factual questions balanced across six domains, and it lets the model abstain, so a confident wrong answer is penalized rather than rewarded. Artificial Analysis also says it filtered known contamination sources out of the search results before grading, which matters because a benchmark answer sitting in the retrieved page is not retrieval, it is a leak. None of this makes the Index the last word. It makes it an honest measurement of one clearly defined thing, which is more than the layer had a week before it shipped. What the first Index says. At the top: Parallel Search in advanced mode scores 75, Exa Search on auto scores 74, Firecrawl Search scores 73, and Parallel Search in basic mode also scores 73. That is the leaderboard. Look at the spread. The top four sit inside two points. On a 0-100 scale, across three benchmarks totaling 1,700 tasks, the difference between first and fourth is smaller than the noise most teams would see re-running the suite on a different day. Read literally, the Index does not say Parallel is better than Firecrawl. It says the leading providers are, for this workload, roughly interchangeable on quality. That is the finding, and it is easy to miss because a ranked table invites you to treat rank one as the answer. A two-point gap is not a verdict. It is a tie with an ordering imposed on it. A tie on quality throws the decision onto the axes the score compresses away, which is where similar-looking providers stop being similar. Two providers can post near-identical scores while running on completely different underlying indexes, which changes freshness, coverage, and what breaks when your traffic shifts. The three numbers under the score most builders skip. Cost is two lines, not one. Artificial Analysis reports cost per 1,000 tasks as two figures: the search provider's charge and the answer model's charge. Parallel advanced runs $47.93 in search plus $35.58 in model cost. Exa auto runs $65.57 plus $61.58. Firecrawl runs $30.48 plus $44.94. The model line is the one people forget, and it is not fixed across providers even though the model is. A provider that returns longer or noisier results makes the model read more and search more often, so it drives up the model bill it does not send you. The provider with the cheaper per-call search can end up more expensive per answer. You cannot see that from the search price on a pricing page, which is why the split matters. Latency is per task, not per call. The Index reports latency per task, not per search call, and the gap is large: Parallel advanced averages 35.9 seconds per task while Keenable in realtime mode averages 15.1 seconds. Artificial Analysis states the reason plainly: "A provider can be fast per call and still contribute more total time, if the model searches more often against it." Per-call latency is the number vendors publish. Per-task latency is the number your user feels, because an agent may issue five or ten searches to answer one question. A fast call attached to a strategy that calls repeatedly is not a fast experience. If you are building something interactive, this axis can outrank quality outright. The benchmark's workload is not your workload. DeepSearchQA, BrowseComp, and AA-Omniscience are research and hard-fact tasks: many-hop questions, obscure facts, broad-domain trivia. If your product does deep research, that is a fair proxy. If it looks up current prices, monitors news, or answers narrow questions inside one vertical, the ranking can invert, because the skills that win a multi-hop trivia hunt are not the skills that win a freshness race. The native-search-versus-API evals showed the same thing, with headline scores that were low and tightly bound to the task set. A benchmark measures its own workload faithfully and yours only by luck. How to use the Index without over-fitting. Use it as a filter, not a decision. The Index is strong evidence that the top cluster is competent and that anything far down the table has something to prove. That alone is worth having, and it is enough to build a shortlist from. Then stop reading the rank. Weight the two axes the single score hides against your own shape: if you serve interactive traffic, latency per task leads; if you run high volume, put both cost lines in a spreadsheet with your real call pattern, not the vendor's headline price. A two-point quality spread will not survive contact with a 2x swing in cost or latency, so let those decide among the leaders. Then run your own evaluation. The providers at the top are close enough that the tiebreaker has to come from your queries, not this one. Running the same comparison on your own traffic is the step the Index makes easier, not the step it replaces. It tells you where to point a small eval, so you test three plausible providers on your questions instead of seven on someone else's. The measured read. A neutral scoreboard for AI search APIs is genuinely new, and it beats a market where every comparison came from a seller. The provider-swap design is honest, the contamination filtering is the right instinct, and the component benchmarks are well chosen for what they cover. Treat the Index as a credible filter and a sanity check on your own results, not as a ranking to buy off the top row. It measures one workload carefully. Whether that workload is yours is the question the leaderboard cannot answer, and the one worth spending your evaluation budget on. * #benchmarks * #ai-search-apis * #vendor-evaluation * #api-selection

MarkTechPost
Aug 25th, 2026
Liquid AI open-sources pipette: A reproducible benchmarking suite that measures on-device models, quantization, runtime and hardware together.

Liquid AI open-sources pipette: A reproducible benchmarking suite that measures on-device models, quantization, runtime and hardware together. August 25, 2026 Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week, Liquid AI released Pipette. It is an open-source platform for benchmarking foundation models on edge devices, built in partnership with Artificial Analysis as an independent methodology validator. Pipette treats on-device behavior as a property of the deployed system, not the model in isolation. Its unit of measurement is a full configuration: model + quantization + runtime + device. The launch dataset covers five on-device performance metrics across more than 1,000 model x quantization x runtime x device x context configurations, spanning 30+ models, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens. Initial verified results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra. The practical claim is testable: two 350M models at the same quantization on the same phone retain 78.4% and 33.8% of decode throughput at 4,096 tokens. Is it deployable? Yes, Pipette ships as Apache 2.0 infrastructure (pipette-mgmt, pipette-clients, pipette-scores), a public results dataset, a hosted dashboard, and native iOS and Android benchmark apps. Nothing is waitlisted. Publication of community-submitted results is still in beta. * Which companies: Any team shipping a model onto hardware it does not own. Solo developers and seed-stage startups can use the dashboard and apps without infrastructure. Mid-market product teams can run the clients across an internal device fleet. Large OEMs, chip vendors and enterprises can operate the whole pipeline behind their own firewall. * Industries: Consumer electronics and smartphone OEMs, automotive, industrial and robotics, healthcare devices, financial services, defense - anywhere latency, privacy or connectivity forces inference onto the device. * Applications: Model and quantization selection before a sprint commits; SoC and hardware procurement validation; regression testing when a runtime, OS or driver updates; context-length capacity planning; independent verification of vendor performance claims. What Liquid AI shipped. Liquid AI released Pipette in partnership with Artificial Analysis, an independent validator that reviewed and verified the methodology. The premise is narrow and useful: on-device behavior is a property of the deployed system, not of the model in isolation. The launch dataset covers five on-device performance metrics across more than 1,000 model x quantization x runtime x device x context configurations. It spans 30+ models, multiple quantization formats, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens. Initial published results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra, with AMD Ryzen AI Max+ 395 and Radeon 8060S results listed as coming soon. In Pipette, the unit of measurement is a deployment configuration: model + quantization + runtime + device. A benchmark then defines the metric and token shape, producing a latency, throughput or memory result. Quality is tracked separately on IFBench, GPQA Diamond and MATH-500. Those quality scores currently come from llama.cpp evaluation runs on NVIDIA H100 80GB reference systems, then get matched to on-device runs sharing the same model and quantization - a quality number shown next to phone throughput was not produced on the phone. Why the deployment context changes the answer. Four published comparisons show how far a configuration can move a decision: * Context scaling can diverge at identical parameter counts. At Q4_K_M on Galaxy S26 Ultra, Granite-4.0-H-350M retains 78.4% of its decode throughput from 256 to 4,096 input tokens, while Granite-4.0-350M retains only 33.8%. * Sparse activation buys speed, not memory. At 2,048 input tokens on the same phone, LFM2.5-8B-A1B decodes 2.4x faster than Qwen3.5-4B and 2.6x faster than Ministral-3-3B-Instruct-2512. It activates 1.5B of 8.5B parameters per token, yet still peaks at 5.29 GiB because all expert weights occupy memory. * Speed and quality do not co-locate. On iPhone 17 Pro at Q4_K_M, MiniCPM5-1B completes a 2,048-in / 256-out workload in 3.47 seconds versus 4.12 seconds for LFM2.5-1.2B-Instruct, a 15.8% reduction in elapsed time. On the same artifacts, LFM scores 9.0 points higher on MATH-500. * Near-identical system profiles can hide task-level reversals. At Q4_K_M and 2,048 input tokens on M5 Max, Granite-4.1-8B and Ministral-3-8B-Instruct-2512 differ by 2.4% in decode throughput and 1.2% in peak RAM. Granite leads IFBench by 7.3 points; Ministral leads GPQA Diamond by 14.0 points. How the measurements are produced. Performance runs follow a published methodology: fixed token shapes, greedy decoding, a discarded warm-up, five measured repetitions and readiness gating. Before each timed repetition, a platform-specific check verifies thermal and load conditions; failing runs are not published. Evaluations use a separate protocol with deterministic, model-blind scoring, and pipette-scores never sees generation provenance. Every submission records benchmark version, token shape, model artifact, quantization, runtime version and settings, and device hardware and OS. Interactive explainer. Key takeaways. * Pipette benchmarks configurations, not models: model + quantization + runtime + device. * Apache 2.0 stack, 1,000+ configurations, 30+ models, three verified devices at launch. * Quality evals run on H100 references and are matched to on-device performance, not measured on-device. * Identical parameter counts can differ 78.4% vs 33.8% in context-scaling retention.

Auraboros
Aug 18th, 2026
The Agentic intelligence report: what happened in AI agents On August 18, 2026.

The Agentic intelligence report: what happened in AI agents On August 18, 2026. Inside the August 18, 2026 report: An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Docum..., followed by the wider AI signals worth carrying forward. Published August 19, 2026 · Neutral source-linked reporting Executive summary. On August 18, 2026, the clearest AI pattern was practical validation. Across arXiv cs.AI, The Decoder AI, the cycle kept returning to the same operator question: which claims are strong enough to change how teams build, buy, or govern AI systems right now. The dominant themes were evaluation and reliability, agent workflows, tooling and developer workflows. The source material was more detailed than usual, which made the cycle easier to read through an operator lens. For serious operators, the right response is disciplined narrowing: treat launches as hypotheses, use benchmarks as filters rather than verdicts, and only move quickly when capability, workflow fit, and operating constraints all point in the same direction. An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive document layouts: A plant science use case. Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Nicolas Turenne [view email] [v1] Fri, 19 Jun 2026 11:50:27 UTC (2,002 KB) Full-text links: Access Paper: View a PDF of the paper titled An Agentic Framework Using Rules and LLMs for... Why this matters now: Research and evaluation stories matter because they reset the standard for what counts as credible model evidence. If the claim holds up, it will influence how teams benchmark, buy, and govern AI systems. What still needs proof: The main uncertainty is transferability. Strong benchmark or research results do not automatically mean better performance in messy production settings with long context, tools, and human oversight in the loop. Practical read: Treat this as a scoring signal, not a verdict. Fold it into your eval suite and decision rubric before you let it change procurement or deployment choices. New benchmark ranks search APIs for AI agents on quality, cost, and speed. Artificial Analysis has released the "Search Index," a benchmark that rates search API providers for AI agents on quality, cost, and speed. Of seven providers tested with GPT-5.6 Luna, Parallel, Exa, and Firecrawl scored highest. Artificial Analysis has released the "Search Index," a benchmark that measures how well search API providers work for AI agents across quality, cost, and speed. Why this matters now: Research and evaluation stories matter because they reset the standard for what counts as credible model evidence. If the claim holds up, it will influence how teams benchmark, buy, and govern AI systems. What still needs proof: The main uncertainty is transferability. Strong benchmark or research results do not automatically mean better performance in messy production settings with long context, tools, and human oversight in the loop. Practical read: Treat this as a scoring signal, not a verdict. Fold it into your eval suite and decision rubric before you let it change procurement or deployment choices. Do LLM Agents Negotiate Rationally? A mechanism-design Framework for verifiable multi-agent interaction over A2A/MCP. Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individually rational, or strategy-proof outcomes. Focus to learn more arXiv-issued DOI via DataCite Submission history From: Wael Albayaydh [view email] [v1] Fri, 10 Jul 2026 07:32:19 UTC (529 KB) Full-text links: Access Paper: View a PDF of the paper titled Do LLM Agents Negotiate Rationally? Why this matters now: Launch stories matter because they force immediate stack decisions. The key question is whether the capability survives real prompts, latency targets, and budget constraints or remains mostly release framing. What still needs proof: Headline momentum is clear, but the important questions are still practical: pricing, rollout scope, reliability under load, and whether the capability improvement shows up in everyday workflows. Practical read: Do not upgrade on launch energy alone. Put the claim through your own prompts, latency checks, and budget constraints before you touch a production default. Crosscurrents to watch. The deeper pattern in this cycle is evaluation pressure. The individual stories are also getting more concrete: vendor blogs, research notes, and media coverage are all pointing at operational detail rather than abstract possibility. The names will change tomorrow, but the operating pressure is stable: teams are being forced to make faster calls on evaluation and reliability, agent workflows, tooling and developer workflows while still carrying the burden of reliability, cost discipline, and governance. * evaluation and reliability: More of the cycle is being decided by whether outputs are verifiable, benchmarked, and resilient under real usage conditions. * agent workflows: The strongest stories are increasingly about whether agents can handle real multi-step work, not just produce impressive demos. * tooling and developer workflows: Practical tooling is becoming a bigger source of advantage because it changes build speed, iteration quality, and failure handling. * infrastructure economics: Cost, latency, and serving constraints still determine whether strong capability can survive contact with production. Benchmark context. Benchmark leaders still matter, but only when paired with deployment fit and real workflow validation. * GPT-5 (OpenAI, overall 98) * Claude Opus 4.1 (Anthropic, overall 97) * Gemini 2.5 Pro (Google, overall 96) Operator note: Benchmark leadership is useful for orientation, not for skipping reliability, integration, or cost validation. Operator bottom line. Today's winners will not be the teams that react fastest to every AI headline. They will be the teams that separate genuine operating leverage from launch theater, test the important claims quickly, and move only when the evidence is good enough. References. AI transparency. This report and its hero image were produced with AI systems and AI agents under human direction. Publishing workflow and controls are documented at Journey. Want This Daily? If this report was useful, get the next one by email with fresh sources, tools, and benchmark movement. Fresh report every day Operator-grade summaries No spam No spam. Unsubscribe anytime.

Renascence
Aug 17th, 2026
Optima lets enterprises benchmark AI models against their own data.

Optima lets enterprises benchmark AI models against their own data. Artificial Analysis has launched Optima, letting organisations test AI models against their own data and workflows, measuring cost and time per task rather than relying on generic public leaderboards. Renascence Newsdesk What happened. Artificial Analysis has launched Optima, a benchmarking platform that lets organisations test AI models against their own data and workflows rather than relying solely on generic, public leaderboards. The tool allows users to build custom evaluations using their own tasks, then compares how different models perform not just on output quality but on cost and time taken to complete each task. According to The Decoder, this shifts the comparison away from standardised, one-size-fits-all benchmarks toward metrics that reflect how a model actually behaves inside a specific organisation's use case. For agent-based applications in particular, where a model might make multiple calls, chain reasoning steps or interact with tools, the report notes that cost and time per completed task are described as more informative than headline token pricing alone. Why it matters. Public AI benchmarks have long been criticised for measuring performance on generic tasks that may bear little resemblance to how a business actually deploys a model. Optima's approach - evaluating models against an organisation's real data and workflows - points to a broader shift in how enterprises are expected to select and manage AI systems: less by trusting aggregate leaderboard rankings, and more by running their own fit-for-purpose trials before committing. For teams building AI-powered agents or automations, this matters because token-based pricing alone can obscure the true operating cost of a workflow. A model that looks cheaper per token may still take longer or require more steps to complete a task, driving up total cost and latency. Tools that surface cost and time per task give technology and operations leaders a more grounded basis for model selection, budgeting and vendor comparison as agentic AI moves from pilot to production. The Renascence take. The interesting part of this launch isn't the benchmarking tool itself - it's the admission behind it: that generic AI leaderboards have been quietly misleading the people making real deployment decisions. Most organisations still choose AI models the way shoppers choose a phone by spec sheet alone - headline scores, not lived performance. But a model's real value shows up in the friction it creates or removes for actual users completing actual tasks, at actual cost and speed. Any team deploying AI at scale should treat vendor-published benchmarks as a shortlist tool at best, and insist on testing against their own workflows before committing - because the gap between "benchmark-good" and "operationally good" is exactly where customer and employee experience gets quietly eroded. This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage. Questions Renascence get on this topic. Optima is a benchmarking platform launched by Artificial Analysis that lets organisations test AI models against their own data and workflows, rather than relying solely on generic public leaderboards. More in AI Stay ahead of CX Get the signal, not the noise. The stories shaping customer experience - plus the Journal and Experience Loom - in your inbox.

Hugging Face
May 27th, 2026
ITBench-AA: frontier models score below 50% on the first benchmark for agentic enterprise IT tasks - by artificial Analysis and IBM.

ITBench-AA: frontier models score below 50% on the first benchmark for agentic enterprise IT tasks - by artificial Analysis and IBM. Artificial Analysis and IBM Software Innovation Lab are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA's SRE tasks benchmark model performance on Kubernetes incident response, where models and agents must diagnose live systems by reading logs, tracing dependencies, and identifying root-cause entities across complex infrastructure. The underlying ITBench dataset has been developed by IBM, leveraging deep expertise in enterprise IT operations. Artificial Analysis has worked closely with IBM over the last 6 months to develop an implementation of the dataset for frontier AI evaluation, beginning with Site Reliability Engineering (SRE) and expanding to Financial Operations (FinOps) and Chief Information Security Officer (CISO) tasks over time. Key findings: * Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42%. * All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks in its suite. For context, frontier models score considerably higher on Terminal-Bench. * Turn counts vary nearly 3x and longer trajectories do not translate to higher accuracy. GPT-5.5 (xhigh) averages 31 turns per task at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Models that over-investigate tend to surface upstream fault-injection mechanisms or co-occurring symptoms as false positives. * GLM-5.1 (Reasoning) leads open weights models at 40%, effectively tied with Gemini 3.5 Flash (high). DeepSeek V4 Pro (Reasoning, Max Effort) follows at 38%, with Gemma 4 31B (Reasoning) at 37%, ahead of Gemini 3.1 Pro Preview at 30%. ITBench-AA SRE overview: * 59 SRE tasks in total: 40 public tasks and 19 brand new, held-out tasks * Each task provides a Kubernetes incident snapshot containing alerts, events, traces, metrics, logs, and application topology. The model must identify the minimal set of independent root-cause Kubernetes entities responsible for the incident. * Faults span typical SRE failure modes including infrastructure, service, application, and chaos-injected incidents, such as resource quota exhaustion, rollout failures, connection pool exhaustion, and network partitions. Methodology details: * Agentic harness: each task is solved by the model running in its open-source Stirrup reference harness, with shell access to a sandboxed file system containing the relevant logs and snapshots. 100-turn cap per task, 3 repeats per task. * Models and agents submit a list of root-cause entities (Kubernetes Deployments, Services, Pods, etc.) they believe caused the incident. Each submission is compared against a ground-truth set of root causes provided by IBM. * Scoring uses average precision at full recall: if a model misses any of the ground-truth root causes, it scores 0.0 for that repeat. If it identifies all of them, it is awarded a score equal to its precision - the share of its submitted entities that are actual root causes, i.e. true positives / (true positives + false positives). The headline score is the average across 59 tasks x 3 repeats. * The harness (Stirrup) is held constant across all evaluated models, allowing an apples-to-apples comparison between models. Highlights. * Tasks require agents to investigate Kubernetes incident snapshots through shell commands and submit a structured JSON diagnosis identifying the responsible root-cause entities. In one public SRE task, the agent sees user-facing failures in the frontend path. It uses shell commands to inspect the offline snapshot: reviewing alerts shows the incident window, then traces/logs narrow the failure to frontend traffic. Topology pins down the affected services, and Kubernetes manifests reveal a network policy blocking the frontend. The successful diagnosis identifies the responsible root-cause entity: otel-demo/NetworkPolicy/frontend-block-all-ports. * More turns do not mean better answers. Models that submit additional contributing entities beyond the true root cause get penalized: identifying the correct root cause but adding upstream mechanisms (e.g., a chaos-mesh controller) or co-occurring symptoms counts as a false positive under recall-gated precision. This is why some models with long trajectories underperform terser ones: Gemini 3.1 Pro Preview averages 83 turns and scores 30%, while Gemma 4 31B (Reasoning) averages 58 turns and scores 37%. * Open weights models sit on the cost frontier of ITBench-AA SRE. Gemma 4 31B (Reasoning) scores 37% at $0.14 per task, outperforming Gemini 3.1 Pro Preview ($2.23 per task, 30%) on both score and cost. GLM-5.1 (Reasoning) scores 40% at $1.23 per task, matching Gemini 3.5 Flash (high) ($1.70) on score at lower cost. Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads the leaderboard at 47% but is the most expensive at $5.38 per task. ITBench-AA is built in partnership with @IBM based on their ITBench benchmark.