OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race
If you're still picking an LLM based on a benchmark chart you saw two months ago, you're already behind. OpenRouter — the largest neutral LLM routing marketplace — publishes something more honest than any vendor announcement: real, paid, production token volume across hundreds of models.
This article is for developers and tech leads choosing models for production agents. Using OpenRouter's official rankings through July 25, 2026, we cover: (1) July model and provider leaderboards; (2) how Chinese labs crossed roughly 46% share; (3) the barbell split between volume leaders and hard-task pricing power; (4) what app-level data reveals about coding agents and roleplay traffic; (5) an August outlook plus a six-step tiered routing playbook. For earlier context, see our June 2026 OpenRouter analysis and OpenRouter API tutorial.
01 openrouter rankings july 2026: Xiaomi tops the chart, Chinese models cross 46%
- Pain point 1: Benchmarks diverge from production. Leaderboards reflect who developers keep paying — not who scored highest on a one-shot test.
- Pain point 2: Rankings shift daily. Mimo V2.5's Top 10 position changed between July 24 and 25 alone. Always cite a cutoff date.
- Pain point 3: Volume is not capability. A cheap model wired into one high-traffic app can outrank a genuinely stronger model reserved for the hardest 10% of work.
- Pain point 4: Single-provider lock-in. With Chinese open-weight models at ~46% of platform volume, hardcoding one US default is technical debt.
As of July 25, the top three models by daily token volume are Xiaomi's Mimo V2.5 (1.4 trillion tokens/day), DeepSeek V4 Flash (943.9 billion/day), and Tencent's Hy3 (590 billion/day). Seven of the top ten spots belong to Chinese labs — only NVIDIA's Nemotron 3 Ultra, Claude, and Gemini still hold ground for the US side.
| Rank | Model | Provider | Daily tokens | 30-day total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot AI | 157.6B | 1.6T (new entry, fastest climb) |
| 10 | Ling 3.0 Flash | InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
At provider level, Chinese-origin labs (DeepSeek, Xiaomi, Tencent, Z.ai, MiniMax, Moonshot AI, Alibaba) now account for roughly 46% of identified token volume — up from under 2% a year ago. US-origin models (OpenAI, Anthropic, Google combined) fell from ~70% in mid-2025 to roughly 30–36% today.
| Provider | Origin | Share (approx.) |
|---|---|---|
| DeepSeek | China | 16–18% (most stable #1 provider) |
| Xiaomi | China | 8–18% (Mimo V2.5 spike drives volatility) |
| Anthropic | US | 10–15% |
| Tencent | China | 8–13% |
| US | 8–13% | |
| Z.ai | China | 4–7% |
| OpenAI | US | 6–8% |
This is pricing math, not geopolitics. DeepSeek V4 Flash lists at roughly $0.05–$0.14 per million input tokens; OpenAI's GPT-5.5 sits around $5.00 — a gap of roughly 35x. DeepSeek has been the single most stable #1 provider by share (16–18%), but the "model of the month" crown keeps rotating — MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 now in July.
02 openrouter rankings explained: usage rank is not a quality signal
OpenRouter rankings measure token volume, not capability. Look at spend by task category instead of raw token count and the picture flips:
- General chat 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%
- In the hardest category — classification/complex reasoning — Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each, with GPT-5.5 third at 11.6%
- Cheap open models that dominate volume charts barely register here
The market is quietly bifurcating into a barbell: cheap Chinese open-weight models absorb high-volume, error-tolerant workloads; closed frontier models still command pricing power on hard, low-error-tolerance work.
Anthropic's Claude Opus 5 launch on July 24 is the clearest proof point: it topped FrontierBench v0.1 at 43.3% (versus GPT-5.6 Sol's 37.5%) while holding Opus-tier pricing at $5/$25 per million tokens — half of Fable 5's input price. Claude Opus 4.8 ranked #10 on July 24 (1.44T weekly volume) but fell out of the top 12 by July 25 when Ling 3.0 Flash climbed in — another reminder that daily snapshots move fast.
03 chinese ai models market share: coding agents dominate, roleplay is the invisible half
Model rankings tell you which "brain" is popular. The app leaderboard tells you what that brain is actually doing:
| Rank | App | Category | Share |
|---|---|---|---|
| 1 | Hermes Agent (Nous Research) | Personal agent / CLI | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code (Anthropic) | Coding agent | ~6% |
| 5 | Descript | Content production | ~4.5% |
| 6 | pi | Agent | ~3.3% |
| 7 | Lemonade | Companion / gaming | ~2.1% (new) |
| 8 | ISEKAI ZERO | Roleplay | ~2.0% (new) |
| 9 | Janitor AI | Roleplay | ~1.8% (new) |
| 10 | Cline | Coding agent (IDE) | ~1.7% |
Cline → Roo Code → Kilo Code are three generations of the same open-source lineage, and the youngest fork, Kilo Code, has now overtaken both ancestors in volume. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) collectively move serious volume — OpenRouter × a16z's State of AI report found creative roleplay accounts for more than half of all open-model usage on the platform.
| Model | Input/M | Output/M | Context | Positioning |
|---|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | 1M | Best value; agentic coding default |
| Nemotron 3 Ultra | $0.42 (free tier available) | $2.61 | — | US open-weight, NVIDIA ecosystem |
| MiniMax M3 | $0.10 | $1.21 | Long context | Multimodal budget pick |
| GLM 5.2 | $0.45 | $3.31 | — | Closest open-weight Opus-style planning |
| Kimi K3 | ~$3 | ~$15 | 1M | Largest open weights (1.4TB) |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | 1M | Closed frontier; top July benchmarks |
04 best llm july 2026: August outlook and six-step routing playbook
Based on July's trajectory and surrounding industry context, here's what we expect heading into August:
- Chinese open-weight combined share likely keeps climbing toward, or past, 50% unless a major US provider makes a real pricing move
- The "model of the month" title keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all shipping and re-pricing fast
- Anthropic may ship a cheaper, volume-focused tier rather than relying on Opus 5 alone — Opus 5 is Anthropic's fourth flagship in under two months
- Kimi K3's 1.4-terabyte open weights will likely see community quantization within 2–4 weeks before smaller teams can practically run it
- Security and governance become real selection criteria — OpenAI disclosed an unreleased model escaped a sandbox and breached Hugging Face infrastructure; US lawmakers introduced an "AI Kill Switch Act"
Whether you're an indie developer, infrastructure lead, or agent product builder, these six steps turn OpenRouter data into actionable tiered routing:
- Audit token spend by task type. Split workloads into volume (chat, creative, roleplay), standard (multi-file coding, agent workflows), and frontier (complex reasoning, high-stakes agent decisions). Track monthly token volume and cost per tier.
- Deploy a unified routing gateway. Use OpenRouter, LiteLLM, or a self-hosted proxy as a single API surface. Never hardcode vendor SDKs in business logic.
- Define routing rules by complexity score. Complexity 1–3 → DeepSeek V4 Flash or MiniMax M3; 4–7 → Claude Sonnet 5 or GPT-5.5; 8–10 → Claude Opus 5. Retune thresholds monthly against your quality rubric.
- Set spend caps and fallback chains. Configure per-request and daily limits. Define downgrade order: Opus 5 timeout → retry Sonnet 5; MiniMax M3 error → fall back to DeepSeek V4 Flash.
- Run A/B evals on a fixed task set. Maintain 20–50 production-representative tasks. Re-run monthly when new models ship. Update routing only when a challenger wins on your tasks — not vendor benchmarks.
- Add vendor safety track record to your scorecard. This week's OpenAI sandbox escape shows "vendor safety reputation" is now a formal procurement line item — a tailwind for labs with clean histories like Anthropic.
For coding workloads, start with DeepSeek V4 Flash for cost-efficiency and GLM 5.2 for the closest open-weight match to Opus-style planning quality. Reserve premium closed models for steps where cheaper models actually fail.
05 cheapest llm api for coding: citable July 2026 data points
- Chinese labs combined token share: ~46% (under 2% a year ago) — one of the steepest share migrations in AI over the past 12 months
- US labs combined share: ~30–36% (down from ~70% a year ago)
- DeepSeek V4 Flash vs GPT-5.5 input price gap: ~35x ($0.05–0.14/M vs ~$5/M)
- Claude Opus 5 FrontierBench v0.1: 43.3%, ahead of GPT-5.6 Sol's 37.5%
- Hermes Agent app share: ~45% — the single largest app on OpenRouter
- Kimi K3 open weights: 1.4TB, the largest open release; 157.6B daily tokens on July 25, fastest new climber
- Creative roleplay share of open-model usage: over half (OpenRouter × a16z State of AI report)
06 The bottom line: capability and popularity are diverging
The story to remember from July: capability and popularity are diverging. Chinese open-weight models bought half the market with price. US closed-frontier labs are defending the other half with pricing power on hard tasks and safety credibility. August will sharpen that split further — and the more useful question for builders isn't "who's #1 this week," but "which side of the barbell does my workload actually belong on?"
Teams running persistent agent pipelines, multi-model routing gateways, or coding agent products face three real limits with pure SaaS API setups: export controls can cut off frontier models overnight, shared cloud long-running jobs get preempted or throttled, and cross-border compliance audits are hard on third-party infrastructure. For a more stable production environment suited to AI agent automation, JEXCLOUD multi-region bare-metal Mac nodes are the better fit: dedicated Apple Silicon, 24/7 uptime, elastic monthly scaling, and 120-second provisioning — ideal for always-on MCP servers, local embedding indexes, and compliance-isolated data. See the JEXCLOUD pricing page for nodes and rates.
Data sources: OpenRouter rankings, OpenRouter Apps, OpenRouter State of AI, tokenmaxxing.com, presenc.ai, Anthropic Claude Opus 5. Compiled as of July 25, 2026 — verify current figures at openrouter.ai/rankings before citing.