Full-Time

Applied AI Engineer

Posted on 9/11/2026

Arena Intelligence

Arena Intelligence

51-200 employees

Crowdsourced LLM evaluation and comparison platform

Compensation Overview

$150k - $250k/yr

+ Equity

San Francisco Bay Area, CA, USA

Hybrid

Category
Software Engineering (1)
Required Skills
Data Structures & Algorithms
TypeScript
DevOps
Data Analysis

Get referred to Arena Intelligence

See people who can refer or advise you

Requirements
  • At least 4 years of software engineering experience with strong fundamentals in data structures, algorithms, and system design.
  • Production experience with TypeScript and modern web frameworks.
  • Hands-on experience integrating large language model provider application programming interfaces, including streaming responses.
  • Comfort working across frontend and backend systems.
  • Familiarity with cloud infrastructure and continuous integration and continuous delivery workflows.
  • Ability to solve problems and operate effectively in ambiguous situations.
  • Ability to communicate clearly and concisely about complex technical ideas to engineers and non-technical stakeholders.
  • A customer-first orientation and willingness to work closely with multiple stakeholders.
Responsibilities
  • Own 5–15 active model lab relationships as the day-to-day point of contact.
  • Partner directly with researchers and product teams to integrate models, run evaluations, and deliver high-quality results.
  • Drive consistent delivery across onboarding, data and product releases, co-launches, and unexpected issue resolution.
  • Own a key process improvement or automation, such as the model onboarding flow, customer data insights, or delivery tooling.
  • Translate customer feedback and research direction into concrete product ideas and engineering work.
  • Work across the stack to ship pragmatic solutions for customers.
  • Help expand strategic accounts through trust, execution, and technical judgment.

Arena Intelligence provides a crowdsourced, open platform for evaluating large language models (LLMs) by running head-to-head comparisons of two anonymous models on user prompts. Users vote on the better response and, after voting, model identities are revealed; the results generate large-scale preference data used to compute Elo ratings published on a public leaderboard. Originating from LMSYS and UC Berkeley’s Sky Computing Lab as Chatbot Arena, the project evolved into a company and rebranded to Arena, with $100 million seed funding in 2024 to scale the platform. The underlying framework, FastChat, remains open source, including the user interface and serving backend. Arena differentiates itself by providing a transparent, real-world benchmarking method that aggregates community-driven judgments and offers analytics to AI labs and enterprises, aiming to guide model development with observable performance data.

Company Size

51-200

Company Stage

Series A

Total Funding

$250M

Headquarters

San Francisco, California

Founded

2025

Get referred to Arena Intelligence

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Arena reached $100 million annualized revenue in June 2026, eight months after launch.
  • A January 2026 Series A raised $150 million at a $1.7 billion valuation.
  • The March 2026 leaderboard changelog added direct-battle votes, increasing volume and improving ranking stability.

What critics are saying

  • 2025 research exposed provider-specific sampling, score retraction, and private-variant gaming on Arena.
  • If enterprises trust distorted rankings, Arena’s neutrality brand breaks, crippling AI Evaluations sales in 2026.
  • Scale AI, Mercor, and Surge compete directly on evaluation workflows and can undercut pricing.

What makes Arena Intelligence unique

  • Arena’s leaderboard shapes model selection across OpenAI, Google, Anthropic, and DeepSeek.
  • Its 10-million-plus human comparisons create a live preference dataset competitors cannot replicate.
  • Arena’s AI Evaluations monetizes that trust with enterprise-grade analytics and audits.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

-8%

1 year growth

-3%

2 year growth

-9%
Business Insider
Aug 8th, 2026
Arena AI CEO: Enterprises torn between costly frontier models and Chinese open-source AI

Companies are struggling to decide which AI models to trust, according to Arena AI CEO Anastasios Angelopoulos. Speaking on the "20VC" podcast, he said enterprises face a difficult choice between frontier labs and Chinese open-source providers. Frontier companies like OpenAI and Anthropic offer advanced models but charge premium prices and can shut off access without warning or impose safeguards limiting capabilities. Chinese firms like DeepSeek offer cheaper open-weight models that run locally, but face potential security concerns and possible US restrictions. Arena operates a crowdsourced platform where users compare AI models. In June, the company announced it had reached $100 million in annualised run-rate revenue within eight months of launching its enterprise service. Angelopoulos said some model providers have tried using Arena's internal data to get employees to label data, though he declined to name specific companies.

TechCrunch
Jun 29th, 2026
Arena, the AI leaderboard everyone uses, is now a $100M business

Arena, the AI leaderboard provider that began as a UC Berkeley research project in 2023, has reached $100 million in annualised run-rate revenue just eight months after launching its commercial service. The company's post-money valuation stands at $1.7 billion following a $150 million Series A round in January. Arena is known for its crowdsourced AI model performance leaderboard, generated from over 10 million user evaluations. Whilst the public leaderboard remains free, the company monetises through AI Evaluations, a service providing model labs and enterprises with performance analytics. The startup competes with human labelling companies like Mercor, Surge and Scale AI for post-training refinement services. Arena has raised $250 million from investors including Felicis, Andreessen Horowitz and Kleiner Perkins. Co-founders include CEO Anastasios Angelopoulos, CTO Wei-Lin Chiang and Databricks co-founder Ion Stoica.

TechCrunch
Mar 18th, 2026
Arena hits $1.7B valuation ranking AI models with funding from the companies it judges

Arena, formerly LM Arena, has been valued at $1.7 billion just seven months after launching as a UC Berkeley PhD research project. The startup operates the leading public leaderboard for frontier large language models, influencing funding decisions and product launches across the AI industry. Co-founders Anastasios Angelopoulos and Wei-Lin Chiang claim their platform is harder to manipulate than static benchmarks, using what they call "structural neutrality" to evaluate AI models. The company is backed by the very companies it ranks, including OpenAI, Google and Anthropic. Arena currently shows Claude topping expert leaderboards in legal and medical use cases. The startup is expanding beyond chat to benchmark AI agents, coding and real-world tasks through a new enterprise product.

YuJiaComm
Jan 7th, 2026
LMArena achieves $1.7B valuation four months after launching its product

LMArena achieves $1.7B valuation four months after launching its product. LMArena, a startup that originally launched as a UC Berkeley research project in 2023, announced on Tuesday that it raised a $150 million Series A at a post-money valuation of $1.7 billion. The round was led by Felicis and the university's fund, UC Investments. The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months. LMArena is best known for its crowdsourced AI model performance leaderboards. Its consumer website lets a user type a prompt that it sends to two models, with the user then choosing which model did a better job. Those results, which now span more than 5 million monthly users across 150 countries and 60 million conversations a month, the company says, fuel the leaderboards. It ranks various models on a variety of tasks including text, web development, vision, text-to-image, and other criteria. The models it tests include various flavors of OpenAI GPT, Google Gemini, Anthropic Claude, and Grok, as well as ones that are geared toward specialties like image generation, text to image, or reasoning. The company began as Chatbot Arena, an open research project built by UC Berkeley researchers Anastasios Angelopoulos and Wei-Lin Chiang, and was originally funded through grants and donations. LMArena's leaderboards became something of an obsession among model makers. When LMArena started pursuing revenue, it partnered with select model companies such as OpenAI, Google, and Anthropic to make their flagship models available for its community to evaluate. In April, a group of competitors published a paper alleging that this helped those model makers game the startup's benchmarks, an allegation LMArena has vehemently denied. In September, it publicly launched a commercial service, AI Evaluations, in which enterprises, model labs, and developers can hire the company to perform model evaluations through its community. This gave LMArena an annualized "consumption rate" - as the company describes its annual recurring revenue (ARR) - of $30 million as of December, less than four months after launch. Join the disrupt 2026 waitlist. Add yourself to the disrupt 2026 waitlist to be first in line when early bird tickets drop. Past disrupts have brought Google cloud, netflix, microsoft, box, phia, a16z, elevenlabs, wayve, hugging face, elad gil, and vinod khosla to the stages - part of 250+ industry leaders driving 200+ sessions built to fuel your growth and sharpen your edge. Plus, meet the hundreds of startups innovating across every sector.

Dealroom.co
Jan 6th, 2026
LMArena company information, funding & investors

LMArena, open community platform to benchmark and compare ai models through side-by-side evaluations and user voting. Here you'll find information about their funding, investors and team.