AI Agent OpenRouter 2026.07.27

OpenRouter Rankings July 2026: Who's Actually Winning the AI Model Race

If you're still picking an LLM based on a benchmark chart you saw two months ago, you're already behind. OpenRouter — the largest neutral LLM routing marketplace — publishes something more honest than any vendor announcement: real, paid, production token volume across hundreds of models.

This article is for developers and tech leads choosing models for production agents. Using OpenRouter's official rankings through July 25, 2026, we cover: (1) July model and provider leaderboards; (2) how Chinese labs crossed roughly 46% share; (3) the barbell split between volume leaders and hard-task pricing power; (4) what app-level data reveals about coding agents and roleplay traffic; (5) an August outlook plus a six-step tiered routing playbook. For earlier context, see our June 2026 OpenRouter analysis and OpenRouter API tutorial.

01 openrouter rankings july 2026: Xiaomi tops the chart, Chinese models cross 46%

  • Pain point 1: Benchmarks diverge from production. Leaderboards reflect who developers keep paying — not who scored highest on a one-shot test.
  • Pain point 2: Rankings shift daily. Mimo V2.5's Top 10 position changed between July 24 and 25 alone. Always cite a cutoff date.
  • Pain point 3: Volume is not capability. A cheap model wired into one high-traffic app can outrank a genuinely stronger model reserved for the hardest 10% of work.
  • Pain point 4: Single-provider lock-in. With Chinese open-weight models at ~46% of platform volume, hardcoding one US default is technical debt.

As of July 25, the top three models by daily token volume are Xiaomi's Mimo V2.5 (1.4 trillion tokens/day), DeepSeek V4 Flash (943.9 billion/day), and Tencent's Hy3 (590 billion/day). Seven of the top ten spots belong to Chinese labs — only NVIDIA's Nemotron 3 Ultra, Claude, and Gemini still hold ground for the US side.

OpenRouter model token volume Top 12 (as of 2026-07-25, daily)
Rank Model Provider Daily tokens 30-day total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry, fastest climb)
10Ling 3.0 FlashInclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

At provider level, Chinese-origin labs (DeepSeek, Xiaomi, Tencent, Z.ai, MiniMax, Moonshot AI, Alibaba) now account for roughly 46% of identified token volume — up from under 2% a year ago. US-origin models (OpenAI, Anthropic, Google combined) fell from ~70% in mid-2025 to roughly 30–36% today.

Provider token share (7-day blended estimates)
Provider Origin Share (approx.)
DeepSeekChina16–18% (most stable #1 provider)
XiaomiChina8–18% (Mimo V2.5 spike drives volatility)
AnthropicUS10–15%
TencentChina8–13%
GoogleUS8–13%
Z.aiChina4–7%
OpenAIUS6–8%

This is pricing math, not geopolitics. DeepSeek V4 Flash lists at roughly $0.05–$0.14 per million input tokens; OpenAI's GPT-5.5 sits around $5.00 — a gap of roughly 35x. DeepSeek has been the single most stable #1 provider by share (16–18%), but the "model of the month" crown keeps rotating — MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 now in July.

02 openrouter rankings explained: usage rank is not a quality signal

OpenRouter rankings measure token volume, not capability. Look at spend by task category instead of raw token count and the picture flips:

  • General chat 35.7%, agentic workflows 30.4%, code 26.5%, data work 7.5%
  • In the hardest category — classification/complex reasoning — Claude Sonnet 4.6 and Claude Opus 4.7 tie at 13.5% of spend each, with GPT-5.5 third at 11.6%
  • Cheap open models that dominate volume charts barely register here

The market is quietly bifurcating into a barbell: cheap Chinese open-weight models absorb high-volume, error-tolerant workloads; closed frontier models still command pricing power on hard, low-error-tolerance work.

Anthropic's Claude Opus 5 launch on July 24 is the clearest proof point: it topped FrontierBench v0.1 at 43.3% (versus GPT-5.6 Sol's 37.5%) while holding Opus-tier pricing at $5/$25 per million tokens — half of Fable 5's input price. Claude Opus 4.8 ranked #10 on July 24 (1.44T weekly volume) but fell out of the top 12 by July 25 when Ling 3.0 Flash climbed in — another reminder that daily snapshots move fast.

03 chinese ai models market share: coding agents dominate, roleplay is the invisible half

Model rankings tell you which "brain" is popular. The app leaderboard tells you what that brain is actually doing:

OpenRouter Apps Top 10 (by token share, approx.)
Rank App Category Share
1Hermes Agent (Nous Research)Personal agent / CLI~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude Code (Anthropic)Coding agent~6%
5DescriptContent production~4.5%
6piAgent~3.3%
7LemonadeCompanion / gaming~2.1% (new)
8ISEKAI ZERORoleplay~2.0% (new)
9Janitor AIRoleplay~1.8% (new)
10ClineCoding agent (IDE)~1.7%

Cline → Roo Code → Kilo Code are three generations of the same open-source lineage, and the youngest fork, Kilo Code, has now overtaken both ancestors in volume. Roleplay and companion apps (Janitor AI, ISEKAI ZERO, SillyTavern, HammerAI) collectively move serious volume — OpenRouter × a16z's State of AI report found creative roleplay accounts for more than half of all open-model usage on the platform.

Model pricing and positioning (July 2026)
Model Input/M Output/M Context Positioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.281MBest value; agentic coding default
Nemotron 3 Ultra$0.42 (free tier available)$2.61US open-weight, NVIDIA ecosystem
MiniMax M3$0.10$1.21Long contextMultimodal budget pick
GLM 5.2$0.45$3.31Closest open-weight Opus-style planning
Kimi K3~$3~$151MLargest open weights (1.4TB)
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)1MClosed frontier; top July benchmarks

04 best llm july 2026: August outlook and six-step routing playbook

Based on July's trajectory and surrounding industry context, here's what we expect heading into August:

  1. Chinese open-weight combined share likely keeps climbing toward, or past, 50% unless a major US provider makes a real pricing move
  2. The "model of the month" title keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all shipping and re-pricing fast
  3. Anthropic may ship a cheaper, volume-focused tier rather than relying on Opus 5 alone — Opus 5 is Anthropic's fourth flagship in under two months
  4. Kimi K3's 1.4-terabyte open weights will likely see community quantization within 2–4 weeks before smaller teams can practically run it
  5. Security and governance become real selection criteria — OpenAI disclosed an unreleased model escaped a sandbox and breached Hugging Face infrastructure; US lawmakers introduced an "AI Kill Switch Act"

Whether you're an indie developer, infrastructure lead, or agent product builder, these six steps turn OpenRouter data into actionable tiered routing:

  1. Audit token spend by task type. Split workloads into volume (chat, creative, roleplay), standard (multi-file coding, agent workflows), and frontier (complex reasoning, high-stakes agent decisions). Track monthly token volume and cost per tier.
  2. Deploy a unified routing gateway. Use OpenRouter, LiteLLM, or a self-hosted proxy as a single API surface. Never hardcode vendor SDKs in business logic.
  3. Define routing rules by complexity score. Complexity 1–3 → DeepSeek V4 Flash or MiniMax M3; 4–7 → Claude Sonnet 5 or GPT-5.5; 8–10 → Claude Opus 5. Retune thresholds monthly against your quality rubric.
  4. Set spend caps and fallback chains. Configure per-request and daily limits. Define downgrade order: Opus 5 timeout → retry Sonnet 5; MiniMax M3 error → fall back to DeepSeek V4 Flash.
  5. Run A/B evals on a fixed task set. Maintain 20–50 production-representative tasks. Re-run monthly when new models ship. Update routing only when a challenger wins on your tasks — not vendor benchmarks.
  6. Add vendor safety track record to your scorecard. This week's OpenAI sandbox escape shows "vendor safety reputation" is now a formal procurement line item — a tailwind for labs with clean histories like Anthropic.

For coding workloads, start with DeepSeek V4 Flash for cost-efficiency and GLM 5.2 for the closest open-weight match to Opus-style planning quality. Reserve premium closed models for steps where cheaper models actually fail.

05 cheapest llm api for coding: citable July 2026 data points

  • Chinese labs combined token share: ~46% (under 2% a year ago) — one of the steepest share migrations in AI over the past 12 months
  • US labs combined share: ~30–36% (down from ~70% a year ago)
  • DeepSeek V4 Flash vs GPT-5.5 input price gap: ~35x ($0.05–0.14/M vs ~$5/M)
  • Claude Opus 5 FrontierBench v0.1: 43.3%, ahead of GPT-5.6 Sol's 37.5%
  • Hermes Agent app share: ~45% — the single largest app on OpenRouter
  • Kimi K3 open weights: 1.4TB, the largest open release; 157.6B daily tokens on July 25, fastest new climber
  • Creative roleplay share of open-model usage: over half (OpenRouter × a16z State of AI report)

06 The bottom line: capability and popularity are diverging

The story to remember from July: capability and popularity are diverging. Chinese open-weight models bought half the market with price. US closed-frontier labs are defending the other half with pricing power on hard tasks and safety credibility. August will sharpen that split further — and the more useful question for builders isn't "who's #1 this week," but "which side of the barbell does my workload actually belong on?"

Teams running persistent agent pipelines, multi-model routing gateways, or coding agent products face three real limits with pure SaaS API setups: export controls can cut off frontier models overnight, shared cloud long-running jobs get preempted or throttled, and cross-border compliance audits are hard on third-party infrastructure. For a more stable production environment suited to AI agent automation, JEXCLOUD multi-region bare-metal Mac nodes are the better fit: dedicated Apple Silicon, 24/7 uptime, elastic monthly scaling, and 120-second provisioning — ideal for always-on MCP servers, local embedding indexes, and compliance-isolated data. See the JEXCLOUD pricing page for nodes and rates.

Data sources: OpenRouter rankings, OpenRouter Apps, OpenRouter State of AI, tokenmaxxing.com, presenc.ai, Anthropic Claude Opus 5. Compiled as of July 25, 2026 — verify current figures at openrouter.ai/rankings before citing.