Arena AI
FreemiumThe official crowdsourced leaderboard for comparing frontier AI models across text, code, vision, and agentic tasks.
Arena AI is the official crowdsourced platform for ranking and comparing frontier AI models through real-world human preference voting. Originally launched as Chatbot Arena by LMSYS and UC Berkeley SkyLab in 2023, it evolved into LMArena and was rebranded to Arena in January 2026. The platform lets users pit anonymous models against each other in head-to-head battles, vote on the better response, and reveal identities afterward, with results feeding public Elo-based leaderboards across text, code, vision, document, search, and image-generation arenas.

What is it
A community-driven AI model evaluation platform that ranks frontier large language and multimodal models through live pairwise battles and direct testing.
What it can do
Run anonymous side-by-side model battles, chat directly with specific models, explore specialized leaderboards for text/code/vision/document/search/image tasks, and use Agent Mode for multi-step research, coding, and analysis workflows.
Who is it for
AI researchers, developers, data scientists, and enthusiasts who need unbiased, real-world evidence to choose the right model for a task, plus AI labs and enterprises that purchase custom evaluation services.
Key Features
Battle Mode — Anonymous Model Battles
Submit one prompt and receive two anonymous model responses side by side. After voting for the better output, Arena reveals which models competed and records the result toward a live Elo leaderboard that reflects genuine human preference rather than marketing claims.
Direct & Side-by-Side Modes — Controlled Comparison
Skip anonymity and select specific frontier or open-source models from a ranked dropdown to chat directly or compare them side by side, with modality icons showing whether a model supports text, vision, image generation, or coding.
Specialized Arenas — Domain-Specific Leaderboards
Explore dedicated leaderboards including Text Arena, Code Arena, Vision Arena, Document Arena, Search Arena, Text-to-Image Arena, Image Edit Arena, and Agent Arena, each collecting task-specific human judgments.
Agent Mode — Multi-Step Autonomous Workflows
Switch from chat to a connected workflow where an orchestrator model plans, researches, writes code, generates images, and iterates inside a sandbox/bash environment to complete complex tasks such as building a website or planning a product launch.
Use Cases
AI researchers benchmarking model progress
Collect live human-preference data across multiple capability domains and publish findings backed by millions of real user votes rather than static benchmarks.
Developers choosing a coding model
Use Code Arena and WebDev leaderboards to compare how models build, render, and iterate on real web apps before committing to an API provider.
Product teams selecting a vision or image model
Run Vision Arena and Text-to-Image Arena battles to see which models most effectively handle image understanding, editing, or generation for a specific visual use case.
Data scientists evaluating document-understanding models
Upload PDFs to Document Arena to compare long-context reasoning and structured extraction across frontier models.
Pricing plans
freemium
Visit the website for detailed pricing information.
Frequently Asked Questions
Discussion
No comments yet. Be the first to start the thread.