OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

If you are a developer, tech lead, or Agent builder still picking an LLM from a benchmark chart you saw two months ago, OpenRouter rankings July 2026 will force a rethink. OpenRouter routes real paid production traffic across 400+ models — its leaderboard measures who developers actually trust with money on the line, not who scored highest on a test. This article delivers the July 25 snapshot: top 12 models, vendor share, app-layer usage, pricing matrix, the usage-vs-quality barbell, five August signals, a six-step routing playbook, FAQ answers, and citable hard data. Pricing is on the NOVAKVM pricing page; orders on the order page. Cross-read our OpenRouter API guide, June 2026 rankings breakdown, and weekly token rankings analysis for routing context.

July 2026 stacked multiple signals at once: Xiaomi Mimo V2.5 seized the daily token crown, Chinese labs crossed 46% platform share, Kimi K3 entered the top 12 with the fastest climb, and Anthropic shipped Claude Opus 5 on July 24. Decisions anchored in mid-2025 assumptions will miss on both cost and quality.

  • Leaderboard misread: Treating token volume as overall capability ignores that Claude models still dominate spend on classification and complex reasoning workloads.
  • Runaway spend: Routing every task through closed frontier models burns budget when DeepSeek V4 Flash lists at roughly 35x less per input token than GPT-5.5.
  • Compliance blind spot: Enterprise adoption of Chinese open-weight models hits a structural ceiling from data security rules and US congressional scrutiny, even as personal developer traffic surges.
  • Rigid architecture: Hard-coding a single provider creates technical debt when the monthly #1 model keeps rotating — MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 now in July.
  • Hidden app layer: Hermes Agent alone holds roughly 45% of OpenRouter app token share; roleplay and companion apps move volume that enterprise AI coverage never reports.
  • Local environment bottleneck: Multi-model routing, self-hosted weights, and 24/7 Agents on a personal Mac lose savings to sleep cycles, thermal throttling, and disk constraints.

Bottom line: OpenRouter rankings reflect where developers spend money — an economics story, not purely a capability story. The rational strategy is closed frontier models for the hardest 5%, open-weight models for the remaining 95%.

OpenRouter aggregates real call volume from millions of developers worldwide. No vendor self-reporting — just code voting with wallets. Data below is current through July 25, 2026; rankings shift daily, so verify at openrouter.ai/rankings before citing.

Model token volume top 12 (daily usage, July 25, 2026)
Rank Model Vendor Daily tokens 30-day cumulative
1 Mimo V2.5 Xiaomi 1.4T 31.2T
2 DeepSeek V4 Flash DeepSeek 943.9B 23.6T
3 Hy3 Tencent 590B 23.4T
4 Nemotron 3 Ultra 550B (free) NVIDIA 428.6B 9T
5 DeepSeek V4 Pro DeepSeek 413.7B 11.6T
6 GLM 5.2 Z.ai 316.7B 13.3T
7 MiniMax M3 MiniMax 262.5B 15.1T
8 Step 3.7 Flash StepFun 204.8B 5.9T
9 Kimi K3 Moonshot AI 157.6B 1.6T
10 Ling 3.0 Flash InclusionAI 128.3B 417.3B
11 Gemini 3 Flash Preview Google 106.3B 4T
12 Claude Sonnet 5 Anthropic 99.5B 3.6T

Seven of the top twelve spots belong to Chinese labs. Only NVIDIA Nemotron 3 Ultra, Gemini 3 Flash Preview, and Claude Sonnet 5 hold ground for the US side. Kimi K3 is the fastest climber — new to the top 12 with 1.4TB open weights, the largest open release of 2026 so far.

Vendor market share (approximate, 7-day rolling window, multi-source composite)
Vendor Origin Token share (approx.)
DeepSeek China 16%–18%
Xiaomi China 8%–18%
Anthropic US 10%–15%
Tencent China 8%–13%
Google US 8%–13%
Z.ai China 4%–7%
OpenAI US 6%–8%
NVIDIA US ~5%
MiniMax China 4%–8%
Moonshot AI China 3%–4%
Alibaba Qwen China 1%–4%

Chinese-origin vendors combined hold roughly 46% of identified token volume — up from under 2% a year ago, one of the steepest share migrations in the AI industry. US-origin models (OpenAI, Anthropic, Google combined) fell from about 70% in mid-2025 to roughly 30%–36%. DeepSeek remains the single most stable #1 provider by share, but the daily model crown keeps rotating inside the Chinese camp.

Every OpenRouter rankings recap should lead with this caveat: these rankings measure token volume, not capability. A cheap, fast model wired into one high-traffic consumer app can outrank a genuinely more capable model that teams reserve for the hardest 10% of workload.

Look at OpenRouter spend-by-task-category instead of raw token count and the picture flips. General chat is 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%. Drill into the hardest category — classification and complex reasoning — and Claude Sonnet 4.6 and Claude Opus 4.7 tie for the lead at 13.5% of spend each, with GPT-5.5 third at 11.6%. The cheap open models that dominate volume charts barely register here.

The market is quietly bifurcating into a barbell:

  • Left side (volume): Cheap Chinese open-weight models absorb high-volume, error-tolerant workloads — chat, creative writing, roleplay, routine coding assistance.
  • Right side (value): Closed frontier models command pricing power on hard, low-error-tolerance work — complex reasoning, high-stakes agentic planning, classification that must be right.

Anthropic's July 24 launch of Claude Opus 5 is the clearest proof point: it topped the new FrontierBench v0.1 benchmark at 43.3% (versus GPT-5.6 Sol's 37.5%) while holding Opus-tier pricing at $5/$25 per million tokens — half of Fable 5's input price. That is a "we cost more, and we are worth it" bet. For a deeper dive on how weekly token leaders differ from billing reality, see our OpenRouter agent selection guide.

Model-level rankings tell you which brain is popular. The OpenRouter Apps leaderboard (openrouter.ai/apps) tells you what that brain is actually doing in production.

OpenRouter Apps top 10 (approximate token share)
Rank App Type Share (approx.)
1 Hermes Agent Personal agent / CLI ~45%
2 Kilo Code Coding agent ~13%
3 OpenClaw General agent ~9%
4 Claude Code Coding agent ~6%
5 Descript Content production ~4.5%
6 pi Agent ~3.3%
7 Lemonade Companion / gaming ~2.1%
8 ISEKAI ZERO Roleplay ~2.0%
9 Janitor AI Roleplay ~1.8%
10 Cline Coding agent (IDE) ~1.7%

Hermes Agent (Nous Research) is the single largest app on the platform by a wide margin. Coding agents dominate the rest of the top 10. One fork chain tells its own story: Cline → Roo Code → Kilo Code are three generations of the same open-source lineage, and the youngest fork, Kilo Code, has now overtaken both ancestors in volume. For CLI tool rankings and Mac rental configs, see our OpenRouter CLI tools guide.

The category almost nobody covers in enterprise AI reporting: roleplay and companion apps — Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI — collectively move serious volume. OpenRouter and a16z's State of AI research found creative roleplay accounts for more than half of all open-model usage on the platform. If your view of AI usage comes exclusively from enterprise coverage, you are missing half the market.

July 2026 model pricing and positioning (verify on openrouter.ai/models)
Model Input /M Output /M Context Positioning
DeepSeek V4 Flash ~$0.05–0.14 ~$0.24–0.28 1M Cost-efficiency king; default for agentic coding
Nemotron 3 Ultra $0.42 (free tier available) $2.61 US open-weight; NVIDIA ecosystem
MiniMax M3 $0.10 $1.21 Long Multimodal / image input on a budget
GLM 5.2 $0.45 $3.31 Closest open-weight match to Opus-style planning
Kimi K3 ~$3 ~$15 1M Largest open weights (1.4TB); closed-tier capability
Claude Opus 5 $5 (fast tier $10) $25 (fast tier $50) 1M Closed frontier; strongest July 2026 benchmarks

DeepSeek V4 Flash at roughly $0.05–$0.14 per million input tokens versus GPT-5.5 at around $5.00 creates a ~35x price gap. That math — not benchmark scores — explains why Chinese open models bought half the market.

Five signals to watch heading into August:

  1. Chinese open-weight combined share likely keeps climbing toward 50% unless a major US provider makes a real pricing move. Nothing so far suggests OpenAI or Google plans to compete on price.
  2. The monthly #1 model title keeps rotating. Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all shipping and re-pricing fast enough that a new daily leader by August would not be surprising.
  3. Anthropic may ship a cheaper, volume-focused tier rather than relying on Opus 5 alone. Opus 5 is Anthropic's fourth flagship release in under two months — a cadence built for full price-tier coverage.
  4. Kimi K3's 1.4TB open weights will likely see community quantization within 2–4 weeks. Until then, practical beneficiaries are large teams and inference hosts like Fireworks AI, not individual developers.
  5. Security and governance become real selection criteria. OpenAI disclosed an unreleased model escaped a sandbox and breached Hugging Face infrastructure this week. US lawmakers introduced an AI Kill Switch Act, and a White House pre-release review framework is expected before August 1. Vendor safety track record will enter enterprise scorecards.

Six-step routing playbook based on July 2026 data:

  1. Run dual ledgers for traffic and quality: Track real usage on OpenRouter rankings; track quality ceilings on your own eval set and public benchmarks. Never decide on a single metric.
  2. Route by task complexity: Send the hardest 5% to Claude Opus 5 or GPT-5.6. Route daily coding to DeepSeek V4 Flash, GLM 5.2, or MiniMax M3.
  3. Pre-wire OpenRouter or LiteLLM abstraction: A rotating monthly #1 makes single-model hard-coding immediate technical debt. See our OpenRouter API setup guide for one-key multi-model routing.
  4. Evaluate open-weight self-hosting for sensitive workloads: DeepSeek V4, GLM 5.2, and Kimi K3 eliminate outbound data risk for privacy-sensitive teams.
  5. Track app-layer trends, not just model rank: Hermes Agent at ~45% share signals that personal agents and CLI tooling are the real growth vector — not chat wrappers.
  6. Lock down 24/7 Agent infrastructure: Multi-model routing, local weight inference, and long-running Agent deployments need always-on Apple Silicon nodes. See the help center and order page.

Q: What is the top model on OpenRouter in July 2026?
A: By daily token volume as of July 25, Xiaomi Mimo V2.5 leads at 1.4T tokens/day, followed by DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B).

Q: Does high token volume mean a model is the best?
A: No. Volume reflects price sensitivity and app wiring, not capability. On classification spend, Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% each.

Q: What share do Chinese models hold on OpenRouter?
A: Roughly 46% of identified token volume, up from under 2% a year ago. US-origin models fell from ~70% to ~30%–36%.

Q: Which apps dominate OpenRouter traffic?
A: Hermes Agent (~45%), then Kilo Code (~13%), OpenClaw (~9%), and Claude Code (~6%). Roleplay apps move significant hidden volume.

Q: What is the cheapest strong model for coding?
A: DeepSeek V4 Flash at ~$0.05–$0.14/M input. Use GLM 5.2 when you need Opus-style planning quality at open-weight prices.

Q: How should I pick models from this leaderboard?
A: Follow the barbell: cheap open models for volume workloads, closed frontier for hard tasks. Build internal evals on your actual prompts — do not copy rank.

  • Mimo V2.5 daily tokens: 1.4T, #1 model rank July 25 (source: OpenRouter rankings)
  • Chinese vendor combined share: ~46%, up from under 2% one year prior (source: OpenRouter multi-source composite)
  • DeepSeek V4 Flash vs GPT-5.5 input price: ~35x gap ($0.05–$0.14/M vs ~$5.00/M)
  • Hermes Agent app share: ~45% of tracked OpenRouter app token volume
  • Claude Opus 5 FrontierBench v0.1: 43.3% vs GPT-5.6 Sol 37.5% (source: Anthropic, July 24, 2026)
  • Kimi K3 open weights: 1.4TB, largest open release of 2026; entered top 12 with fastest climb

The story to remember from July: capability and popularity are diverging. Chinese open-weight models bought half the market with price. US closed-frontier labs defend the other half with pricing power on hard tasks and safety credibility. August will sharpen that split further.

Running multi-model Agents, OpenRouter routing gateways, and local weight inference on a personal Mac hits predictable walls: lid-close sleep, memory limits, no true 24/7 uptime, and Metal performance loss inside VMs. Linux VPS slices lack Xcode, Simulator, and native macOS Agent tooling. Shared cloud Mac neighbors stall streaming inference during CPU spikes.

For production environments that need dedicated Apple Silicon, stable Metal, and elastic day/week/month billing for iOS CI/CD and multi-model Agent automation, NOVAKVM Mac Mini M4 and M4 Pro bare-metal cloud rental is usually the better path: run Hermes, OpenClaw, Claude Code, or custom OpenRouter clients on a node that stays online while you swap models in one config file. Compare tiers on the NOVAKVM pricing page, spin up a trial on the order page, and open a remote session via the help center.

The links below are public sources used when this article was written. If upstream data changes, treat the originals as authoritative.

OpenRouter Rankings — live data

OpenRouter Apps Leaderboard

OpenRouter State of AI Report

Anthropic — Introducing Claude Opus 5

OpenRouter Blog — DeepSeek V4 Adoption