Full-Time
Crowdsourced LLM evaluation and ranking platform
No salary listed
San Francisco Bay Area, CA, USA
Hybrid
See people who can refer or advise you
Arena Intelligence runs a crowdsourced platform that evaluates and compares large language models. It shows two anonymous models’ responses to the same prompt side-by-side, collects user votes, then reveals model identities; the resulting preference data feeds Elo ratings published on a public leaderboard. The platform serves researchers, developers, and AI enthusiasts who reference Arena’s rankings to gauge model progress, with the FastChat framework powering the interface and backend as open source. Its goal is to provide transparent, real-world evaluation and analytics to help AI labs and enterprises improve model performance while staying neutral and community-driven.
Company Size
51-200
Company Stage
Series A
Total Funding
$250M
Headquarters
San Francisco, California
Founded
2025
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Dental Insurance
Vision Insurance
Company Equity
Companies are struggling to decide which AI models to trust, according to Arena AI CEO Anastasios Angelopoulos. Speaking on the "20VC" podcast, he said enterprises face a difficult choice between frontier labs and Chinese open-source providers. Frontier companies like OpenAI and Anthropic offer advanced models but charge premium prices and can shut off access without warning or impose safeguards limiting capabilities. Chinese firms like DeepSeek offer cheaper open-weight models that run locally, but face potential security concerns and possible US restrictions. Arena operates a crowdsourced platform where users compare AI models. In June, the company announced it had reached $100 million in annualised run-rate revenue within eight months of launching its enterprise service. Angelopoulos said some model providers have tried using Arena's internal data to get employees to label data, though he declined to name specific companies.
Arena, the AI leaderboard provider that began as a UC Berkeley research project in 2023, has reached $100 million in annualised run-rate revenue just eight months after launching its commercial service. The company's post-money valuation stands at $1.7 billion following a $150 million Series A round in January. Arena is known for its crowdsourced AI model performance leaderboard, generated from over 10 million user evaluations. Whilst the public leaderboard remains free, the company monetises through AI Evaluations, a service providing model labs and enterprises with performance analytics. The startup competes with human labelling companies like Mercor, Surge and Scale AI for post-training refinement services. Arena has raised $250 million from investors including Felicis, Andreessen Horowitz and Kleiner Perkins. Co-founders include CEO Anastasios Angelopoulos, CTO Wei-Lin Chiang and Databricks co-founder Ion Stoica.
Arena, formerly LM Arena, has been valued at $1.7 billion just seven months after launching as a UC Berkeley PhD research project. The startup operates the leading public leaderboard for frontier large language models, influencing funding decisions and product launches across the AI industry. Co-founders Anastasios Angelopoulos and Wei-Lin Chiang claim their platform is harder to manipulate than static benchmarks, using what they call "structural neutrality" to evaluate AI models. The company is backed by the very companies it ranks, including OpenAI, Google and Anthropic. Arena currently shows Claude topping expert leaderboards in legal and medical use cases. The startup is expanding beyond chat to benchmark AI agents, coding and real-world tasks through a new enterprise product.
LMArena achieves $1.7B valuation four months after launching its product. LMArena, a startup that originally launched as a UC Berkeley research project in 2023, announced on Tuesday that it raised a $150 million Series A at a post-money valuation of $1.7 billion. The round was led by Felicis and the university's fund, UC Investments. The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months. LMArena is best known for its crowdsourced AI model performance leaderboards. Its consumer website lets a user type a prompt that it sends to two models, with the user then choosing which model did a better job. Those results, which now span more than 5 million monthly users across 150 countries and 60 million conversations a month, the company says, fuel the leaderboards. It ranks various models on a variety of tasks including text, web development, vision, text-to-image, and other criteria. The models it tests include various flavors of OpenAI GPT, Google Gemini, Anthropic Claude, and Grok, as well as ones that are geared toward specialties like image generation, text to image, or reasoning. The company began as Chatbot Arena, an open research project built by UC Berkeley researchers Anastasios Angelopoulos and Wei-Lin Chiang, and was originally funded through grants and donations. LMArena's leaderboards became something of an obsession among model makers. When LMArena started pursuing revenue, it partnered with select model companies such as OpenAI, Google, and Anthropic to make their flagship models available for its community to evaluate. In April, a group of competitors published a paper alleging that this helped those model makers game the startup's benchmarks, an allegation LMArena has vehemently denied. In September, it publicly launched a commercial service, AI Evaluations, in which enterprises, model labs, and developers can hire the company to perform model evaluations through its community. This gave LMArena an annualized "consumption rate" - as the company describes its annual recurring revenue (ARR) - of $30 million as of December, less than four months after launch. Join the disrupt 2026 waitlist. Add yourself to the disrupt 2026 waitlist to be first in line when early bird tickets drop. Past disrupts have brought Google cloud, netflix, microsoft, box, phia, a16z, elevenlabs, wayve, hugging face, elad gil, and vinod khosla to the stages - part of 250+ industry leaders driving 200+ sessions built to fuel your growth and sharpen your edge. Plus, meet the hundreds of startups innovating across every sector.
LMArena, open community platform to benchmark and compare ai models through side-by-side evaluations and user voting. Here you'll find information about their funding, investors and team.