Full-Time

Product Manager

Media Generation, Image & Video

Artificial Analysis

Artificial Analysis

11-50 employees

Independent AI analysis and insights firm

No salary listed

San Francisco, CA, USA

In Person

Category
Product (1)
Required Skills
Python
Data Visualization
TensorFlow
Product Management
PyTorch
Matplotlib
Pandas
NumPy
Video Editing
Data Analysis

Get referred to Artificial Analysis

See people who can refer or advise you

Requirements
  • At least 3 years of professional experience.
  • Strong analytical and critical thinking skills.
  • Excitement about artificial intelligence and eagerness to work at the forefront of technological innovation.
  • Demonstrated excellence academically and professionally.
  • A genuine interest in artificial intelligence image and video generation.
Responsibilities
  • Contribute to image and video product strategy and help execute the roadmap for arenas and leaderboards, including video editing, reference-to-image, and reference-to-video capabilities.
  • Build prompt libraries, categories, and evaluation techniques that reflect how creators and developers use image and video models.
  • Keep coverage up to date by benchmarking new models and providers as they launch, managing generation pipelines and human preference studies.
  • Understand how creators and developers use these models and translate those insights into evaluation frameworks that help users understand capability differences.
  • Help amplify the impact of media benchmarks by collaborating with the team to maximize the impact of content across platforms and audiences.
  • Work with leading image and video companies to benchmark their models, understand their capabilities, and help shape industry standards for evaluating media generation.
  • Use cutting-edge artificial intelligence tools in an artificial-intelligence-native workflow to create leverage in a fast-changing industry and maintain competitiveness in artificial intelligence benchmarking.
Desired Qualifications
  • Experience at an artificial intelligence image or video generation company or a media generation team at a larger lab, especially in product, operations, or research roles.
  • Experience building products leveraging image or video generation models.
  • Familiarity with human preference evaluation methodologies or arena-style ranking systems.
  • Experience in strategy consulting.
  • Strong writing skills, especially in relation to explaining technical topics to both technical and non-technical audiences.
  • Visual communication skills, such as making slides and visualizing data.
  • Proficiency in Python and data analysis libraries such as pandas, NumPy, and Matplotlib; experience with artificial intelligence and machine learning frameworks such as PyTorch and TensorFlow.

Artificial Analysis provides ongoing, independent analysis of the artificial intelligence field. It produces reports and briefings about AI technologies, industry trends, company strategies, and policy/regulatory implications, using evidence-based research and data to inform decision-makers. Its product is a steady stream of analyses that help clients understand what is happening in AI and what it might mean for business, research, and society. What sets it apart is its independence from vendors or lobby groups and its backing by well-known AI leaders, including Nat Friedman, Daniel Gross, and Andrew Ng, which supports rigorous, credible insights. The company’s goal is to provide clear, trustworthy analysis that helps people navigate the fast-changing AI landscape and make better-informed decisions.

Company Size

11-50

Company Stage

Seed

Total Funding

$250K

Headquarters

Newark, Delaware

Founded

2024

Get referred to Artificial Analysis

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Search Index launched August 18, 2026, positioning Artificial Analysis in agentic search infrastructure.
  • Optima, updated August 23, 2026, directly targets enterprise model-selection workflows.
  • ITBench-AA revealed frontier models below 50%, validating demand for harder benchmarks now.

What critics are saying

  • OpenAI, Anthropic, and Google can replicate benchmarking features inside their own platforms.
  • Optima depends on enterprise adoption; slow sales would cap revenue before 2027.
  • If model vendors standardize proprietary evals, Artificial Analysis loses pricing power and relevance.

What makes Artificial Analysis unique

  • August 2026 launches: Optima, Search Index, and ITBench-AA span models, search, and agents.
  • Artificial Analysis tests real workloads, comparing quality, cost, and time per task.
  • IBM partnership on ITBench-AA gives enterprise SRE benchmarks credibility competitors lack.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

23%

1 year growth

23%

2 year growth

40%
Kekulai, Inc
Aug 18th, 2026
The Agentic intelligence report: what happened in AI agents On August 18, 2026.

The Agentic intelligence report: what happened in AI agents On August 18, 2026. Inside the August 18, 2026 report: An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Docum..., followed by the wider AI signals worth carrying forward. Published August 19, 2026 · Neutral source-linked reporting Executive summary. On August 18, 2026, the clearest AI pattern was practical validation. Across arXiv cs.AI, The Decoder AI, the cycle kept returning to the same operator question: which claims are strong enough to change how teams build, buy, or govern AI systems right now. The dominant themes were evaluation and reliability, agent workflows, tooling and developer workflows. The source material was more detailed than usual, which made the cycle easier to read through an operator lens. For serious operators, the right response is disciplined narrowing: treat launches as hypotheses, use benchmarks as filters rather than verdicts, and only move quickly when capability, workflow fit, and operating constraints all point in the same direction. An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive document layouts: A plant science use case. Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Nicolas Turenne [view email] [v1] Fri, 19 Jun 2026 11:50:27 UTC (2,002 KB) Full-text links: Access Paper: View a PDF of the paper titled An Agentic Framework Using Rules and LLMs for... Why this matters now: Research and evaluation stories matter because they reset the standard for what counts as credible model evidence. If the claim holds up, it will influence how teams benchmark, buy, and govern AI systems. What still needs proof: The main uncertainty is transferability. Strong benchmark or research results do not automatically mean better performance in messy production settings with long context, tools, and human oversight in the loop. Practical read: Treat this as a scoring signal, not a verdict. Fold it into your eval suite and decision rubric before you let it change procurement or deployment choices. New benchmark ranks search APIs for AI agents on quality, cost, and speed. Artificial Analysis has released the "Search Index," a benchmark that rates search API providers for AI agents on quality, cost, and speed. Of seven providers tested with GPT-5.6 Luna, Parallel, Exa, and Firecrawl scored highest. Artificial Analysis has released the "Search Index," a benchmark that measures how well search API providers work for AI agents across quality, cost, and speed. Why this matters now: Research and evaluation stories matter because they reset the standard for what counts as credible model evidence. If the claim holds up, it will influence how teams benchmark, buy, and govern AI systems. What still needs proof: The main uncertainty is transferability. Strong benchmark or research results do not automatically mean better performance in messy production settings with long context, tools, and human oversight in the loop. Practical read: Treat this as a scoring signal, not a verdict. Fold it into your eval suite and decision rubric before you let it change procurement or deployment choices. Do LLM Agents Negotiate Rationally? A mechanism-design Framework for verifiable multi-agent interaction over A2A/MCP. Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individually rational, or strategy-proof outcomes. Focus to learn more arXiv-issued DOI via DataCite Submission history From: Wael Albayaydh [view email] [v1] Fri, 10 Jul 2026 07:32:19 UTC (529 KB) Full-text links: Access Paper: View a PDF of the paper titled Do LLM Agents Negotiate Rationally? Why this matters now: Launch stories matter because they force immediate stack decisions. The key question is whether the capability survives real prompts, latency targets, and budget constraints or remains mostly release framing. What still needs proof: Headline momentum is clear, but the important questions are still practical: pricing, rollout scope, reliability under load, and whether the capability improvement shows up in everyday workflows. Practical read: Do not upgrade on launch energy alone. Put the claim through your own prompts, latency checks, and budget constraints before you touch a production default. Crosscurrents to watch. The deeper pattern in this cycle is evaluation pressure. The individual stories are also getting more concrete: vendor blogs, research notes, and media coverage are all pointing at operational detail rather than abstract possibility. The names will change tomorrow, but the operating pressure is stable: teams are being forced to make faster calls on evaluation and reliability, agent workflows, tooling and developer workflows while still carrying the burden of reliability, cost discipline, and governance. * evaluation and reliability: More of the cycle is being decided by whether outputs are verifiable, benchmarked, and resilient under real usage conditions. * agent workflows: The strongest stories are increasingly about whether agents can handle real multi-step work, not just produce impressive demos. * tooling and developer workflows: Practical tooling is becoming a bigger source of advantage because it changes build speed, iteration quality, and failure handling. * infrastructure economics: Cost, latency, and serving constraints still determine whether strong capability can survive contact with production. Benchmark context. Benchmark leaders still matter, but only when paired with deployment fit and real workflow validation. * GPT-5 (OpenAI, overall 98) * Claude Opus 4.1 (Anthropic, overall 97) * Gemini 2.5 Pro (Google, overall 96) Operator note: Benchmark leadership is useful for orientation, not for skipping reliability, integration, or cost validation. Operator bottom line. Today's winners will not be the teams that react fastest to every AI headline. They will be the teams that separate genuine operating leverage from launch theater, test the important claims quickly, and move only when the evidence is good enough. References. AI transparency. This report and its hero image were produced with AI systems and AI agents under human direction. Publishing workflow and controls are documented at Journey. Want This Daily? If this report was useful, get the next one by email with fresh sources, tools, and benchmark movement. Fresh report every day Operator-grade summaries No spam No spam. Unsubscribe anytime.

Renascence
Aug 17th, 2026
Optima lets enterprises benchmark AI models against their own data.

Optima lets enterprises benchmark AI models against their own data. Artificial Analysis has launched Optima, letting organisations test AI models against their own data and workflows, measuring cost and time per task rather than relying on generic public leaderboards. Renascence Newsdesk What happened. Artificial Analysis has launched Optima, a benchmarking platform that lets organisations test AI models against their own data and workflows rather than relying solely on generic, public leaderboards. The tool allows users to build custom evaluations using their own tasks, then compares how different models perform not just on output quality but on cost and time taken to complete each task. According to The Decoder, this shifts the comparison away from standardised, one-size-fits-all benchmarks toward metrics that reflect how a model actually behaves inside a specific organisation's use case. For agent-based applications in particular, where a model might make multiple calls, chain reasoning steps or interact with tools, the report notes that cost and time per completed task are described as more informative than headline token pricing alone. Why it matters. Public AI benchmarks have long been criticised for measuring performance on generic tasks that may bear little resemblance to how a business actually deploys a model. Optima's approach - evaluating models against an organisation's real data and workflows - points to a broader shift in how enterprises are expected to select and manage AI systems: less by trusting aggregate leaderboard rankings, and more by running their own fit-for-purpose trials before committing. For teams building AI-powered agents or automations, this matters because token-based pricing alone can obscure the true operating cost of a workflow. A model that looks cheaper per token may still take longer or require more steps to complete a task, driving up total cost and latency. Tools that surface cost and time per task give technology and operations leaders a more grounded basis for model selection, budgeting and vendor comparison as agentic AI moves from pilot to production. The Renascence take. The interesting part of this launch isn't the benchmarking tool itself - it's the admission behind it: that generic AI leaderboards have been quietly misleading the people making real deployment decisions. Most organisations still choose AI models the way shoppers choose a phone by spec sheet alone - headline scores, not lived performance. But a model's real value shows up in the friction it creates or removes for actual users completing actual tasks, at actual cost and speed. Any team deploying AI at scale should treat vendor-published benchmarks as a shortlist tool at best, and insist on testing against their own workflows before committing - because the gap between "benchmark-good" and "operationally good" is exactly where customer and employee experience gets quietly eroded. This briefing was written by the Renascence newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage. Questions Renascence get on this topic. Optima is a benchmarking platform launched by Artificial Analysis that lets organisations test AI models against their own data and workflows, rather than relying solely on generic public leaderboards. More in AI Stay ahead of CX Get the signal, not the noise. The stories shaping customer experience - plus the Journal and Experience Loom - in your inbox.

Hugging Face
May 27th, 2026
ITBench-AA: frontier models score below 50% on the first benchmark for agentic enterprise IT tasks - by artificial Analysis and IBM.

ITBench-AA: frontier models score below 50% on the first benchmark for agentic enterprise IT tasks - by artificial Analysis and IBM. Artificial Analysis and IBM Software Innovation Lab are launching ITBench-AA, the first in a new series of benchmarks evaluating models on agentic enterprise IT tasks, starting with Site Reliability Engineering tasks where frontier models score below 50% ITBench-AA's SRE tasks benchmark model performance on Kubernetes incident response, where models and agents must diagnose live systems by reading logs, tracing dependencies, and identifying root-cause entities across complex infrastructure. The underlying ITBench dataset has been developed by IBM, leveraging deep expertise in enterprise IT operations. Artificial Analysis has worked closely with IBM over the last 6 months to develop an implementation of the dataset for frontier AI evaluation, beginning with Site Reliability Engineering (SRE) and expanding to Financial Operations (FinOps) and Chief Information Security Officer (CISO) tasks over time. Key findings: * Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42%. * All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks in its suite. For context, frontier models score considerably higher on Terminal-Bench. * Turn counts vary nearly 3x and longer trajectories do not translate to higher accuracy. GPT-5.5 (xhigh) averages 31 turns per task at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Models that over-investigate tend to surface upstream fault-injection mechanisms or co-occurring symptoms as false positives. * GLM-5.1 (Reasoning) leads open weights models at 40%, effectively tied with Gemini 3.5 Flash (high). DeepSeek V4 Pro (Reasoning, Max Effort) follows at 38%, with Gemma 4 31B (Reasoning) at 37%, ahead of Gemini 3.1 Pro Preview at 30%. ITBench-AA SRE overview: * 59 SRE tasks in total: 40 public tasks and 19 brand new, held-out tasks * Each task provides a Kubernetes incident snapshot containing alerts, events, traces, metrics, logs, and application topology. The model must identify the minimal set of independent root-cause Kubernetes entities responsible for the incident. * Faults span typical SRE failure modes including infrastructure, service, application, and chaos-injected incidents, such as resource quota exhaustion, rollout failures, connection pool exhaustion, and network partitions. Methodology details: * Agentic harness: each task is solved by the model running in its open-source Stirrup reference harness, with shell access to a sandboxed file system containing the relevant logs and snapshots. 100-turn cap per task, 3 repeats per task. * Models and agents submit a list of root-cause entities (Kubernetes Deployments, Services, Pods, etc.) they believe caused the incident. Each submission is compared against a ground-truth set of root causes provided by IBM. * Scoring uses average precision at full recall: if a model misses any of the ground-truth root causes, it scores 0.0 for that repeat. If it identifies all of them, it is awarded a score equal to its precision - the share of its submitted entities that are actual root causes, i.e. true positives / (true positives + false positives). The headline score is the average across 59 tasks x 3 repeats. * The harness (Stirrup) is held constant across all evaluated models, allowing an apples-to-apples comparison between models. Highlights. * Tasks require agents to investigate Kubernetes incident snapshots through shell commands and submit a structured JSON diagnosis identifying the responsible root-cause entities. In one public SRE task, the agent sees user-facing failures in the frontend path. It uses shell commands to inspect the offline snapshot: reviewing alerts shows the incident window, then traces/logs narrow the failure to frontend traffic. Topology pins down the affected services, and Kubernetes manifests reveal a network policy blocking the frontend. The successful diagnosis identifies the responsible root-cause entity: otel-demo/NetworkPolicy/frontend-block-all-ports. * More turns do not mean better answers. Models that submit additional contributing entities beyond the true root cause get penalized: identifying the correct root cause but adding upstream mechanisms (e.g., a chaos-mesh controller) or co-occurring symptoms counts as a false positive under recall-gated precision. This is why some models with long trajectories underperform terser ones: Gemini 3.1 Pro Preview averages 83 turns and scores 30%, while Gemma 4 31B (Reasoning) averages 58 turns and scores 37%. * Open weights models sit on the cost frontier of ITBench-AA SRE. Gemma 4 31B (Reasoning) scores 37% at $0.14 per task, outperforming Gemini 3.1 Pro Preview ($2.23 per task, 30%) on both score and cost. GLM-5.1 (Reasoning) scores 40% at $1.23 per task, matching Gemini 3.5 Flash (high) ($1.70) on score at lower cost. Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads the leaderboard at 47% but is the most expensive at $5.38 per task. ITBench-AA is built in partnership with @IBM based on their ITBench benchmark.

VentureBeat
Jun 16th, 2025
Groq Just Made Hugging Face Way Faster — And It’S Coming For Aws And Google

Join the event trusted by enterprise leaders for nearly two decades. VB Transform brings together the people building real enterprise AI strategy. Learn more. Groq, the artificial intelligence inference startup, is making an aggressive play to challenge established cloud providers like Amazon Web Services and Google with two major announcements that could reshape how developers access high-performance AI models.The company announced Monday that it now supports Alibaba’s Qwen3 32B language model with its full 131,000-token context window — a technical capability it claims no other fast inference provider can match. Simultaneously, Groq became an official inference provider on Hugging Face’s platform, potentially exposing its technology to millions of developers worldwide.The move is Groq’s boldest attempt yet to carve out market share in the rapidly expanding AI inference market, where companies like AWS Bedrock, Google Vertex AI, and Microsoft Azure have dominated by offering convenient access to leading language models.“The Hugging Face integration extends the Groq ecosystem providing developers choice and further reduces barriers to entry in adopting Groq’s fast and efficient AI inference,” a Groq spokesperson told VentureBeat. “Groq is the only inference provider to enable the full 131K context window, allowing developers to build applications at scale.”How Groq’s 131k context window claims stack up against AI inference competitorsGroq’s assertion about context windows — the amount of text an AI model can process at once — strikes at a core limitation that has plagued practical AI applications. Most inference providers struggle to maintain speed and cost-effectiveness when handling large context windows, which are essential for tasks like analyzing entire documents or maintaining long conversations.Independent benchmarking firm Artificial Analysis measured Groq’s Qwen3 32B deployment running at approximately 535 tokens per second, a speed that would allow real-time processing of lengthy documents or complex reasoning tasks

Decrypt
Jun 1st, 2025
Best Short-Form Ai Video Generator? Kling 2.1 Vs Google Veo 3

In brief Kling 2.1 launched to compete directly with Google's Veo 3 in the AI video generation market.Testing reveals Kling 2.1 excels at image-to-video conversion while Veo 3 dominates with integrated audio generation capabilities .Both models deliver cinema-quality results, but require different workflows and budget considerations.Decrypt’s Art, Fashion, and Entertainment Hub. Discover SCENEAI video generation just got a serious upgrade. Kuaishou’s Kling 2.1 can now produce videos that look genuinely cinematic—the kind of footage that would have required a film crew and expensive equipment just months ago. Characters move naturally, emotions feel authentic, and complex action sequences unfold without the telltale artifacts that usually scream "this was made by AI."Kling is one of the better-known, advanced video-generation platforms, and was launched a year ago by Kuaishou, a Chinese tech company also known for its social media innovations. It’s especially known for its ability to create HD videos up to two minutes long—and for being the model picked by many meme makers to animate their political satire of people like Trump, Elon Musk, and other influential figures.The new technical improvements include faster generation speeds, better prompt adherence, more realism, and less artifacts. The Master tier utilizes advanced 3D spatiotemporal attention mechanisms and proprietary 3D VAE technology for what the company describes as cinema-grade output.The timing couldn't be more pointed