Full-Time

Backend Engineer

Arena Intelligence

Arena Intelligence

51-200 employees

Crowdsourced LLM evaluation and ranking platform

No salary listed

San Francisco Bay Area, CA, USA

Hybrid

Category
Software Engineering (1)
Required Skills
OpenAI
Postgres
Data Engineering
Role-based Access Control
Go
Redis
REST APIs
LangChain
Data Modeling
Data Analysis

Get referred to Arena Intelligence

See people who can refer or advise you

Requirements
  • At least 5 years of backend engineering experience, including meaningful experience building product-facing application programming interfaces, services, and data systems at scale.
  • Strong proficiency in a modern backend language, with Go preferred, and the judgment to design durable application programming interfaces.
  • Experience modeling, querying, and scaling PostgreSQL, with judgment about when to use Redis, a queue, or a data warehouse.
  • Experience composing services into coherent and reliable products by integrating payments, authentication, analytics, or data pipelines.
  • Ability to think about the developer experience of application programming interfaces, not only their implementation.
  • Comfort working in a startup environment with fluid scope, shifting context, and varied responsibilities.
Responsibilities
  • Design and ship clean, versioned, low-latency, high-reliability application programming interfaces and services for Leaderboards, Evals, and Arena data products.
  • Partner with the research team to turn novel evaluation methods into durable, full-featured products, including scoring pipelines, data models, and application programming interfaces.
  • Build enterprise backend systems for usage metering, cost attribution, billing integration, authentication, role-based access control, multi-tenancy, and audit logging.
  • Unify public and private evaluation data, design schemas that support product growth, and make Arena’s data queryable, consistent, and fast.
  • Contribute to the backend of the Leaderboards and Evals platforms and help unify public and private data architectures.
Desired Qualifications
  • Experience with large language model provider application programming interfaces such as OpenAI, Anthropic, or Google, including streaming, token accounting, rate limits, and model-specific behaviors.
  • Background in machine learning infrastructure, model serving, or evaluation frameworks.
  • Experience building enterprise-ready features such as single sign-on, role-based access control, audit logs, and multi-tenancy.
  • Experience building billing and usage infrastructure with systems such as Stripe, Metronome, and Orb.
  • Familiarity with the modern artificial intelligence stack, including vLLM, LiteLLM, and LangChain.

Arena Intelligence runs a crowdsourced platform that evaluates and compares large language models. It shows two anonymous models’ responses to the same prompt side-by-side, collects user votes, then reveals model identities; the resulting preference data feeds Elo ratings published on a public leaderboard. The platform serves researchers, developers, and AI enthusiasts who reference Arena’s rankings to gauge model progress, with the FastChat framework powering the interface and backend as open source. Its goal is to provide transparent, real-world evaluation and analytics to help AI labs and enterprises improve model performance while staying neutral and community-driven.

Company Size

51-200

Company Stage

Series A

Total Funding

$250M

Headquarters

San Francisco, California

Founded

2025

Get referred to Arena Intelligence

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Arena hit $100 million annualized revenue in June 2026, validating paid evaluations.
  • June 2026 traffic doubled to ten million monthly visitors, expanding benchmark data advantages.
  • Alibaba's August 2026 Qwen3.8-Max launch cited Arena rankings, proving industry dependence.

What critics are saying

  • The Leaderboard Illusion exposed Meta-style variant cherry-picking, weakening Arena's credibility by 2026.
  • If model labs shift scoring to private evals, Arena's influence collapses within months.
  • Crowdsourced votes favor style over truth, making enterprise customers distrust rankings for procurement.

What makes Arena Intelligence unique

  • UC Berkeley-born Arena Intelligence owns the default public LLM leaderboard since 2023.
  • June 4, 2026 Agent Arena extends evaluations from chat preferences to real workflows.
  • January 2026 rebrand to Arena and $1.7 billion valuation reinforce category leadership.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

3%

2 year growth

18%
Business Insider
Aug 8th, 2026
Arena AI CEO: Enterprises torn between costly frontier models and Chinese open-source AI

Companies are struggling to decide which AI models to trust, according to Arena AI CEO Anastasios Angelopoulos. Speaking on the "20VC" podcast, he said enterprises face a difficult choice between frontier labs and Chinese open-source providers. Frontier companies like OpenAI and Anthropic offer advanced models but charge premium prices and can shut off access without warning or impose safeguards limiting capabilities. Chinese firms like DeepSeek offer cheaper open-weight models that run locally, but face potential security concerns and possible US restrictions. Arena operates a crowdsourced platform where users compare AI models. In June, the company announced it had reached $100 million in annualised run-rate revenue within eight months of launching its enterprise service. Angelopoulos said some model providers have tried using Arena's internal data to get employees to label data, though he declined to name specific companies.

TechCrunch
Jun 29th, 2026
Arena, the AI leaderboard everyone uses, is now a $100M business

Arena, the AI leaderboard provider that began as a UC Berkeley research project in 2023, has reached $100 million in annualised run-rate revenue just eight months after launching its commercial service. The company's post-money valuation stands at $1.7 billion following a $150 million Series A round in January. Arena is known for its crowdsourced AI model performance leaderboard, generated from over 10 million user evaluations. Whilst the public leaderboard remains free, the company monetises through AI Evaluations, a service providing model labs and enterprises with performance analytics. The startup competes with human labelling companies like Mercor, Surge and Scale AI for post-training refinement services. Arena has raised $250 million from investors including Felicis, Andreessen Horowitz and Kleiner Perkins. Co-founders include CEO Anastasios Angelopoulos, CTO Wei-Lin Chiang and Databricks co-founder Ion Stoica.

TechCrunch
Mar 18th, 2026
Arena hits $1.7B valuation ranking AI models with funding from the companies it judges

Arena, formerly LM Arena, has been valued at $1.7 billion just seven months after launching as a UC Berkeley PhD research project. The startup operates the leading public leaderboard for frontier large language models, influencing funding decisions and product launches across the AI industry. Co-founders Anastasios Angelopoulos and Wei-Lin Chiang claim their platform is harder to manipulate than static benchmarks, using what they call "structural neutrality" to evaluate AI models. The company is backed by the very companies it ranks, including OpenAI, Google and Anthropic. Arena currently shows Claude topping expert leaderboards in legal and medical use cases. The startup is expanding beyond chat to benchmark AI agents, coding and real-world tasks through a new enterprise product.

YuJiaComm
Jan 7th, 2026
LMArena achieves $1.7B valuation four months after launching its product

LMArena achieves $1.7B valuation four months after launching its product. LMArena, a startup that originally launched as a UC Berkeley research project in 2023, announced on Tuesday that it raised a $150 million Series A at a post-money valuation of $1.7 billion. The round was led by Felicis and the university's fund, UC Investments. The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months. LMArena is best known for its crowdsourced AI model performance leaderboards. Its consumer website lets a user type a prompt that it sends to two models, with the user then choosing which model did a better job. Those results, which now span more than 5 million monthly users across 150 countries and 60 million conversations a month, the company says, fuel the leaderboards. It ranks various models on a variety of tasks including text, web development, vision, text-to-image, and other criteria. The models it tests include various flavors of OpenAI GPT, Google Gemini, Anthropic Claude, and Grok, as well as ones that are geared toward specialties like image generation, text to image, or reasoning. The company began as Chatbot Arena, an open research project built by UC Berkeley researchers Anastasios Angelopoulos and Wei-Lin Chiang, and was originally funded through grants and donations. LMArena's leaderboards became something of an obsession among model makers. When LMArena started pursuing revenue, it partnered with select model companies such as OpenAI, Google, and Anthropic to make their flagship models available for its community to evaluate. In April, a group of competitors published a paper alleging that this helped those model makers game the startup's benchmarks, an allegation LMArena has vehemently denied. In September, it publicly launched a commercial service, AI Evaluations, in which enterprises, model labs, and developers can hire the company to perform model evaluations through its community. This gave LMArena an annualized "consumption rate" - as the company describes its annual recurring revenue (ARR) - of $30 million as of December, less than four months after launch. Join the disrupt 2026 waitlist. Add yourself to the disrupt 2026 waitlist to be first in line when early bird tickets drop. Past disrupts have brought Google cloud, netflix, microsoft, box, phia, a16z, elevenlabs, wayve, hugging face, elad gil, and vinod khosla to the stages - part of 250+ industry leaders driving 200+ sessions built to fuel your growth and sharpen your edge. Plus, meet the hundreds of startups innovating across every sector.

Dealroom.co
Jan 6th, 2026
LMArena company information, funding & investors

LMArena, open community platform to benchmark and compare ai models through side-by-side evaluations and user voting. Here you'll find information about their funding, investors and team.