Skip to main content
UFOZoo

AI Arena Leaderboard: Elo for Chat and Image Models

AI arena leaderboard online: compare chat and image models by Elo ratings from public blind-test arena votes, with vote counts and confidence intervals.

Updated 2026-09-11

Related Tools

Features

  • Twelve arena leaderboards: Chat, Agent, Web Dev, Image to Web Dev, Text to Image, Image Edit, Text to Video, Image to Video, Video Edit, Vision, Document and Search
  • Elo ratings from blind head-to-head votes: real users compare two anonymous models and pick the better one, the same methodology as the original Chatbot Arena
  • Chat board with 385 models: the largest text arena covering open and proprietary LLMs with Elo, vote counts and confidence intervals
  • Agent board with five signals: task success, praise vs complaint, steerability, bash recovery and tool hallucination, each with its own confidence interval
  • Price column: per-1M-token input and output rates so you can weigh quality against cost
  • Vote count column: more votes means a more reliable Elo rating
  • Confidence interval on every rating: shows how much a score could move with more data
  • Context window column for chat models: compare 128K, 1M and larger windows
  • Dated snapshot: every rating carries the snapshot date at the top of the page

How to Use

  1. 1The page opens on the Chat arena, the largest board, with the current Elo leaders at the top of the table
  2. 2Switch tabs to change arena: Agent, Web Dev, Text to Image, Text to Video, Vision, Document, Search and more
  3. 3Read the Elo column: a 100-point gap means the higher-rated model wins roughly 64% of head-to-head votes
  4. 4Check the vote count: models with very few votes may shift significantly as more votes come in
  5. 5Check the confidence interval: a wide interval means the rating is still moving
  6. 6Use the price column to balance quality and cost: the best model per dollar may not be the highest-ranked
  7. 7On the Agent board, read the five signal columns: task success, praise, steerability, bash recovery and tool hallucination (tool hallucination is lower-is-better)
  8. 8Use the top 3-5 models as a shortlist and test them on your own prompts before committing

Frequently Asked Questions

What is an Elo rating and how is it different from benchmark scores?

Elo comes from chess: models are compared head-to-head in blind tests, winners gain points and losers lose points. A 100-point gap means the higher-rated model wins about 64% of matchups. Unlike benchmark scores (accuracy on test sets), Elo reflects real user preference on open-ended tasks: it answers 'which model feels better to use', not 'which scores higher on a test'.

Is this the same as the Chatbot Arena (LMArena)?

Yes, the data is from arena.ai, the platform formerly known as Chatbot Arena / LMArena (IM-bench). It is the same blind-vote methodology and the same public arena data, reorganized into twelve separate leaderboards in one page.

How often is the data updated?

Votes are collected continuously and the snapshot refreshes regularly; the snapshot date is always shown at the top of the page. Check it before citing any ranking.

Why is my favorite model missing from the ranking?

Models appear only after they receive enough arena votes to be ranked. Brand-new or low-traffic models may have no published Elo yet. The snapshot date also matters: a model released after the snapshot is simply not in this data.

What does the confidence interval (±) mean?

The ± range is the statistical uncertainty of the Elo rating. A wider interval means the rating could move more with additional votes. Compare models by treating intervals that overlap as roughly tied.

What does the vote count tell me?

Vote count is the number of head-to-head comparisons the model received. More votes generally mean a more reliable rating. Models with very few votes (under a few hundred) can move a lot as new votes arrive.

How do I read the Agent board's five signals?

Task success is the share of agentic tasks completed correctly. Praise vs complaint is the sentiment of user feedback. Steerability measures how well the agent follows directions. Bash recovery measures recovery from terminal errors. Tool hallucination is the share of invented or wrong tool calls; lower is better. Each signal has its own confidence interval.

Why is the Web Dev board separate from the Chat board?

Web dev tasks evaluate a different skill: turning prompts (or screenshots, on the Image to Web Dev board) into working web pages. A model that chats well is not automatically good at building websites, so the arenas are ranked separately.

Are the image and video boards the same as the AI Video & Image Rankings tool?

No, the AI Video & Image Rankings tool uses benchmark-based Elo from a different evaluator. This page uses the arena.ai blind-vote data, which measures real user preference instead of task scores. The two answers can differ: a model can rank high on objective tests but lower on user preference.

Can I trust the ranking for a production decision?

Use it as a shortlist, not a verdict. Elo reflects crowd preference on generic prompts; your specific workload may favor a lower-ranked model. Pick the top 3-5, test them on your own tasks, and decide on measured results plus cost.

Which open-source models are ranked on the arena?

Open-weights models compete side by side with closed ones: DeepSeek, Llama, Qwen, Mistral, and Gemma appear in the same boards as OpenAI and Anthropic models. The arena does not split the ranking by license, so to compare open models among themselves, filter by model family and look at their relative positions and confidence intervals.

Where does the ranking data come from?

Elo ratings, vote counts and prices are a dated snapshot from arena.ai, the public arena platform (formerly Chatbot Arena / LMArena) that runs blind human-preference votes. We scrape their published leaderboards and clean the values; the snapshot date is at the top of the page.

What is the best AI model right now?

It depends on the arena. In the current Chat snapshot, Claude Fable 5 leads at about 1509 Elo, with Claude Opus 4.6/4.7 Thinking and Qwen3.8 Max close behind. For agentic work, Claude Opus 5 (High) leads the Agent board. For web dev, Claude Opus 5 (max) leads. 'Best' varies by task, so pick the arena that matches what you are building.

Why is there no dedicated coding arena on this leaderboard?

The arena.ai snapshot this page is built from publishes twelve arenas (Chat, Agent, Web Dev, image, video, vision, document, and search), and none of them is a coding-only board. The Agent board is the closest signal for coding work (it measures tool use and bash recovery), and the Chat board covers general reasoning. For coding-specific rankings, look at benchmark-based leaderboards such as SWE-bench, or test a shortlist from this page on your own prompts.