Skip to main content

Arena AI

Freemium

The official crowdsourced leaderboard for comparing frontier AI models across text, code, vision, and agentic tasks.

Arena AI is the official crowdsourced platform for ranking and comparing frontier AI models through real-world human preference voting. Originally launched as Chatbot Arena by LMSYS and UC Berkeley SkyLab in 2023, it evolved into LMArena and was rebranded to Arena in January 2026. The platform lets users pit anonymous models against each other in head-to-head battles, vote on the better response, and reveal identities afterward, with results feeding public Elo-based leaderboards across text, code, vision, document, search, and image-generation arenas.

arenallm leaderboardmodel comparisonai benchmarkingchatbot arenalmsysmodel evaluationarena scoreAI SearchData AnalysisCode GenerationImage GenerationImage RecognitiongenerationanalysisrecognitionIntelligent SearchData AnalysisCode AssistantImage & Vision
Visit Website
Arena AI preview

What is it

A community-driven AI model evaluation platform that ranks frontier large language and multimodal models through live pairwise battles and direct testing.

What it can do

Run anonymous side-by-side model battles, chat directly with specific models, explore specialized leaderboards for text/code/vision/document/search/image tasks, and use Agent Mode for multi-step research, coding, and analysis workflows.

Who is it for

AI researchers, developers, data scientists, and enthusiasts who need unbiased, real-world evidence to choose the right model for a task, plus AI labs and enterprises that purchase custom evaluation services.

Key Features

Battle Mode — Anonymous Model Battles

Submit one prompt and receive two anonymous model responses side by side. After voting for the better output, Arena reveals which models competed and records the result toward a live Elo leaderboard that reflects genuine human preference rather than marketing claims.

Direct & Side-by-Side Modes — Controlled Comparison

Skip anonymity and select specific frontier or open-source models from a ranked dropdown to chat directly or compare them side by side, with modality icons showing whether a model supports text, vision, image generation, or coding.

Specialized Arenas — Domain-Specific Leaderboards

Explore dedicated leaderboards including Text Arena, Code Arena, Vision Arena, Document Arena, Search Arena, Text-to-Image Arena, Image Edit Arena, and Agent Arena, each collecting task-specific human judgments.

Agent Mode — Multi-Step Autonomous Workflows

Switch from chat to a connected workflow where an orchestrator model plans, researches, writes code, generates images, and iterates inside a sandbox/bash environment to complete complex tasks such as building a website or planning a product launch.

Use Cases

AI researchers benchmarking model progress

Collect live human-preference data across multiple capability domains and publish findings backed by millions of real user votes rather than static benchmarks.

Developers choosing a coding model

Use Code Arena and WebDev leaderboards to compare how models build, render, and iterate on real web apps before committing to an API provider.

Product teams selecting a vision or image model

Run Vision Arena and Text-to-Image Arena battles to see which models most effectively handle image understanding, editing, or generation for a specific visual use case.

Data scientists evaluating document-understanding models

Upload PDFs to Document Arena to compare long-context reasoning and structured extraction across frontier models.

Pricing plans

freemium

Visit the website for detailed pricing information.

Frequently Asked Questions

Discussion

No comments yet. Be the first to start the thread.