Skip to main content
UFOZoo

AI Model Leaderboard: Compare LLMs by Intelligence & Speed

AI model leaderboard online: compare LLMs by intelligence, speed, and cost with updated 2026 rankings from public benchmarks.

Updated 2026-09-11

Related Tools

Features

  • Seven ranking dimensions: Intelligence, Coding, Agentic, Cheapest, Fastest, Responsive and Knowledge
  • Parameter-size tiers: Flagship, Mid-size, Compact and Mini, so models are only compared within their own size class
  • Real benchmark scores: dated snapshot of public benchmark aggregates, not vendor claims
  • Cost per task column: average dollars to complete one standard benchmark task, the honest way to compare real usage cost
  • Price column: blended per-1M-token rate (3:1 input-output ratio) plus cache price for every model
  • Context window column: compare 128K, 1M and larger windows side by side
  • Open-weights marker: unlock icon flags models you can self-host
  • Reasoning marker: brain icon flags thinking models that spend tokens before answering
  • est. marker: estimated scores (inferred, not measured) are labelled so you can weigh them accordingly

How to Use

  1. 1The page opens on the Intelligence tab of the Flagship tier: the strongest models by score come first
  2. 2Switch parameter size (Flagship / Mid-size / Compact / Mini): comparing models of different sizes is meaningless, stay within one tier
  3. 3Change tabs to re-rank by another dimension: Coding, Agentic, Cheapest, Fastest, Responsive or Knowledge
  4. 4The Cheapest tab sorts by cost per standard benchmark task ascending: the least expensive per task leads
  5. 5The Fastest tab sorts by output speed in tokens per second descending: for throughput-heavy workloads
  6. 6The Responsive tab sorts by end-to-end response time ascending: for chat and interactive apps
  7. 7The Knowledge tab uses an accuracy-minus-hallucination score that can go negative: a deep negative means heavy hallucination
  8. 8Score bars are normalized within the current list: the highest score fills the bar, relative strength is visible at a glance
  9. 9Hover the est. marker to confirm a score is estimated: estimated scores are less reliable than measured ones

Frequently Asked Questions

How accurate are these leaderboard scores?

Scores are a dated snapshot of public benchmark aggregates, relative indicators, not official results. They work for comparing model tiers, but a 2-point gap between adjacent models is noise. Use them to shortlist, then test the finalists on your own task.

Why is my model missing from the ranking?

261 models are in this snapshot. A model is missing when it has no benchmark score in the source data or when the source does not classify its parameter size (mostly previews and retired models). You can still look it up in the AI Model Database; it just has no score here.

What does the Cheapest tab rank by?

It ranks by cost per task: the average dollars a model spends to complete one standard benchmark task, including input, output and reasoning tokens. Lower is better. Free models cost zero, so they lead. It is the source's own metric, not our calculation.

Is speed the same as latency?

No. Speed is output throughput in tokens per second: how fast the model generates. Responsiveness is end-to-end response time, which also includes time to first token and depends on provider infrastructure. Pick the metric that matches your use case.

Do reasoning models look slow on purpose?

Yes. Reasoning models spend tokens thinking before answering, so raw tok/s is lower than non-reasoning models. That is expected; compare reasoning models against each other, not against non-reasoning ones.

How often is the leaderboard updated?

The snapshot refreshes when the source publishes a new benchmark round, roughly monthly. The snapshot date is always shown at the top of the page; check it before citing any ranking.

Which model is best for coding?

The Coding tab ranks by the coding index: GPT-5.6 Sol (xhigh) leads the current Flagship snapshot at 78.3, ahead of Claude Opus 5. For coding tools like Claude Code and Codex, see the AI Coding Tools leaderboard.

Is this tool different from the AI Model Database?

Yes. The Model Database is for looking up specs and comparing prices: search and filter 300+ models with side-by-side details. This leaderboard answers 'who is best right now?'; it ranks by benchmark scores first.

Why do free models rank first in the Cheapest tab?

The tab sorts by cost per task, and a free model costs zero, so free models naturally lead. Within the same parameter-size tier, free options are the cheapest possible choice; the value question is whether their scores are good enough for your task.

Can I trust the ranking for a production decision?

Use it as a shortlist, not a verdict. Benchmark aggregates average many tasks; your workload may favor a model ranked lower. Pick the top 3-5 for your case, run them on a representative sample, and decide on measured results.

How should I interpret a model's position in the overall ranking?

The overall rank is an average across benchmark dimensions (knowledge, coding, math, instruction following). A model ranked 10th overall can be 2nd in coding; always check the per-dimension rank before choosing. Rankings also change between snapshots, so compare models within one snapshot date rather than across dates.

How do I find open-source (open-weights) models?

Models with open weights carry an unlock icon next to their name: they can be self-hosted. The current snapshot includes DeepSeek V4 Flash, GLM-5.2 and Nemotron 3 Ultra among others. This page does not rank openness separately; it flags open-weights models within each dimension.

What is the cheapest AI model API right now?

By cost per task in the current Flagship snapshot: Command A+ is free, MiMo-V2.5 about $0.01, DeepSeek V4 Flash about $0.03 per task. Outside the Flagship tier, Compact and Mini models are even cheaper. The Price column shows per-1M-token rates; cheapest per token and cheapest per task are different questions.

Which LLM is the fastest?

The Fastest tab ranks by measured output speed: Step 3.7 Flash leads the current Flagship snapshot at about 364 tok/s, followed by Muse Spark 1.1 and Gemini 3.6 Flash. Reasoning models report lower raw tok/s because they generate hidden thinking tokens first.

Where does the ranking and price data come from?

Both scores and prices are a dated snapshot scraped from Artificial Analysis, an independent third-party evaluator. We clean the values and bundle them as static data; the snapshot date is shown at the top of the page. Scores are relative indicators, not vendor results.