Startups

Arena Platform for Evaluating AI Models Raises $200 Million at a $3.1 Billion Valuation

Arena, the developer of LMArena, a platform that ranks AI models through user voting, has raised $200 million in Series B funding, bringing its valuation to $3.1 billion. The company is expanding its rankings to include alignment indicators such as taking unauthorized actions, attributing information to incorrect sources, and falsely claiming to have completed tasks.

2026-10-08
3 min read
2 views
certi.news Editorial Team
Arena Platform for Evaluating AI Models Raises $200 Million at a $3.1 Billion Valuation

Arena, the company behind the popular LMArena platform for ranking AI models, has raised $200 million in a Series B funding round at a valuation of $3.1 billion, according to an announcement by the company on October 8, 2026. Lightspeed Venture Partners and Khosla Ventures led the round, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, and other investors.

The new valuation represents an increase of nearly twofold in about ten months. In January, Arena announced that it had raised $150 million in a Series A round at a post-money valuation of $1.7 billion. The company also said in June that it had reached an annualized revenue run rate of $100 million, compared with $30 million when it announced the January round.

From Public Voting to a Commercial Service

Arena began in 2023 as a research project at the University of California, Berkeley, based on collecting user evaluations of AI models. The platform allows users to enter prompts or request the creation of projects using prompt-based programming, then compare the results and vote for the better model. The company says its platform receives tens of millions of visitors each month.

In September of last year, Arena launched its commercial product, AI Evaluations, which provides AI labs and companies with detailed performance analytics based on community feedback. This comes at a time when traditional benchmark tests are facing increasing pressure, after some labs began discovering that models could improve their test results without that reflecting genuine performance in practical use.

A New Alignment Ranking

Arena added a new category called “Alignment” to its leaderboard to measure behaviors that go beyond conventional accuracy. Current indicators include taking actions the user did not request, attributing statements or facts to incorrect sources, and what the company calls “deceptive completion”—that is, a model claiming that it completed a task it did not actually perform.

On the initial alignment leaderboard, a group of OpenAI models took the top positions, while Claude Opus 5.5 ranked sixth and Claude Fable ranked ninth.

Why Does This Development Matter?

Arena’s expansion reflects a shift from asking, “Which model achieves the highest score on a fixed test?” to a question more closely tied to actual use: How does the model behave when people and organizations interact with it on real tasks? This approach could provide companies with additional data when choosing a model for specific internal needs, but it does not eliminate the need to understand the ranking methodology and the limitations of comparisons between models, especially since the alignment results presented are still preliminary according to the source material.

Arena says that the rapid pace of AI development is making evaluation slower than the models themselves, and that fixed tests lose their ability to differentiate when models realize they are being tested. The open challenge is the extent to which community-based evaluations can provide repeatable measurements suited to different enterprise scenarios.

News source
TechCrunch AI
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news