Claude Opus 5 Cuts the Price in Half — Meanwhile Kimi K3 Gets Caught Calling Itself Claude
Two stories broke this week that look unrelated on the surface. On July 24, Anthropic shipped daily-driver model Claude Opus 5 at half the cost of flagship Fable 5. On July 16, Moonshot AI released open-weight Kimi K3 — then the White House accused the lab of industrial-scale distillation, and independent researcher Ryan Greenblatt found K3 literally identifying itself as Claude, deployment IDs included.
For developers and tech leads choosing models in production, this article covers: ① Opus 5 pricing, benchmarks, and data-retention policy changes; ② the full K3 distillation arc from political accusations to Greenblatt's technical evidence; ③ a six-step framework for deciding between Opus 5 and K3 in the current price war. For earlier K3 architecture context, see our Kimi K3 deep-dive review.
01 This Week's Two AI Stories: Three Model-Selection Pain Points
Both events point at the same industry theme: performance-per-dollar and where model capability actually comes from.
- Pain point one: flagship models are too expensive for daily work. Fable 5 leads on hard tasks, but at roughly $10/$50 per million tokens, scaling coding agents and automation pipelines gets expensive fast.
- Pain point two: open models look suspiciously strong. K3 claims 2.8T parameters and 93.5% on GPQA-Diamond, but full weights were not scheduled to drop until July 27 — external researchers could not independently verify architecture or scores during the controversy.
- Pain point three: compliance and supply-chain risk stack up. White House accusations cover distillation and export-restricted chips; Anthropic flagged over 3.4 million anomalous API interactions from Moonshot back in February — enterprise selection can no longer rely on benchmarks alone.
Opus 5 answers price pressure by closing half the gap to flagship intelligence; K3 answers it with open weights, low price, and fewer refusals. The distillation fight is really the industry asking: is your low cost genuine engineering, or are you quietly riding someone else's model?
02 Claude Opus 5 Release: Not the Flagship, Maybe the Most Useful Claude
On July 24, 2026 (US Pacific time), Anthropic released Claude Opus 5 and immediately made it the default model on Claude Max, plus the strongest model available to Claude Pro subscribers. Model ID: claude-opus-5, available on Claude API, AWS Bedrock, Google Vertex AI, and Microsoft Foundry.
- Pricing unchanged: $5 per million input tokens, $25 per million output tokens — identical to Opus 4.8. Fast mode runs about 2.5x faster at 2x the base price.
- Context: 1M tokens (default and only tier); max output 128K tokens; Thinking enabled by default with Effort parameter controlling reasoning depth.
- Frontier-Bench v0.1: beats every other model on software engineering tasks at more than 2x Opus 4.8's score, with lower cost per task.
- CursorBench 3.2: at max effort, within 0.5% of Fable 5's peak score at half the per-task cost; best performance-per-dollar at high, xhigh, and max effort tiers.
- ARC-AGI 3: scores 3x the next-best model on novel problem-solving.
- OSWorld 2.0: beats Fable 5's best result using barely a third of the cost.
- Zapier AutomationBench: roughly 1.5x the pass rate of the next-best model; even at lowest effort, passes more tasks than any other model. Zapier reports Opus 5 hit 100% pass rate on an end-to-end account-health workflow no prior model could complete.
Beyond coding, Opus 5 shows meaningful gains in structural biology, organic chemistry, and bioinformatics — a 10.2 percentage-point improvement on internal spectroscopy-to-structure benchmarks and 7.7 points higher on protein variant function prediction. Enterprise customer Box reported data-analysis workflows up 11%, due-diligence workflows up 17%, and 8% overall accuracy gain.
Safety and alignment: Anthropic's automated behavioral audit found Opus 5 to be its most aligned model to date — lowest deceptive-behavior rate, hardest to trick into misuse. But Anthropic deliberately did not push Opus 5 to the frontier on dual-use cybersecurity or biology; that spot stays with limited-access Mythos 5. Cyber safety classifiers intervene about 85% less often than Fable 5's, enabling source-code vulnerability discovery while still blocking binary scanning, penetration testing, and exploit generation.
Data retention: like every prior Opus model, no data retention requirement for general access. Fable 5 and Mythos 5 require opting into a 30-day retention policy — for compliance-sensitive workloads, Opus 5 may be the more practical pick regardless of the benchmark gap.
Cursor team's take: "Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it's just under Fable 5 and has many of the same behaviors."
03 Claude Opus 5 vs Fable 5: Which Should You Pick?
| Dimension | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Input / output pricing | $5 / $25 per million tokens | ~$10 / $50 per million tokens |
| Relative cost | About half of Fable 5 | Flagship pricing |
| CursorBench 3.2 (max effort) | 0.5% below Fable 5 peak | Peak benchmark |
| Context window | 1M tokens | 1M tokens |
| Data retention | No 30-day retention required by default | Requires 30-day retention opt-in |
| Claude Max default | Yes (since 2026-07-24) | No |
| Dual-use risk capability | Deliberately below Mythos 5 frontier | Stronger, but Mythos 5 is restricted top tier |
| Best fit | Daily coding agents, automation, compliance-sensitive teams | Frontier SWE, maximum reasoning depth |
If Fable 5 was the model you wanted but not the price you could justify, Opus 5 is Anthropic's direct answer: give up almost nothing on hard agentic and coding tasks for half the cost.
04 Kimi K3 Distillation Controversy: White House Accusations and Timeline Pushback
Moonshot AI released Kimi K3 on July 16, 2026: 2.8 trillion total parameters — the first open-weight model to cross the 3T-class mark. Sparse MoE with 16 of 896 experts activated per token (~50B active-parameter equivalent); 1M-token context plus native vision; Kimi Delta Attention (KDA) architecture with ~2.5x training efficiency gain over K2. Full weights committed for July 27 — during the controversy, external researchers could not independently verify claims.
| Benchmark | Kimi K3 | Notes |
|---|---|---|
| GPQA-Diamond | 93.5% | Best open-weight score at launch |
| Terminal-Bench 2.1 | 88.3% | 0.5 points behind GPT-5.6 Sol |
| BrowseComp | 91.2% | Category best at launch |
| Program Bench | 77.8% | Overall best |
| SWE Marathon | 42.0% | Overall best |
| DeepSearchQA (F1) | 95.0% | — |
On July 22–23, 2026, White House OSTP Director Michael Kratsios posted on X accusing Moonshot of "large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology" from Anthropic's Fable model — and separately alleged export-restricted Nvidia GB300 chips obtained via servers in Thailand. Treasury Secretary Scott Bessent backed the claim, saying officials were "finding watermarks of our U.S. large language models on many of the Chinese models" without specifying what that meant.
This was not the first accusation. In February 2026, Anthropic publicly named Moonshot, DeepSeek, and MiniMax, claiming over 3.4 million anomalous API interactions attributed to "deliberate capability extraction," with some activity traced to senior Moonshot staff via request metadata. Moonshot has never publicly confirmed or denied.
TechCrunch (July 23) interviewed multiple independent researchers skeptical that distillation explains K3 — mainly because Fable 5 was publicly available only since July 1, leaving just two weeks before K3's launch:
"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation... Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks." — Braden Hancock, Laude Institute researcher, co-founder of Snorkel AI
"Distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to reinforcement learning... if it were the case, everyone would be easily able to catch up to a GLM or a K3 by using its data for distillation. But we have not." — Nathan Lambert, Allen Institute for AI
Industry context: Elon Musk testified that xAI distilled OpenAI models while building Grok, calling the practice common. Distillation itself is not necessarily illegal — the dispute is where the line sits between normal technique-borrowing and covert industrial-scale extraction.
05 Why Does Kimi K3 Say It's Claude? Greenblatt's Technical Evidence
The most technically substantive angle in this saga: Redwood Research Chief Scientist Ryan Greenblatt published a statistical analysis around July 24 (GitHub: rgreenblatt/which_claude_is_k3) comparing how models respond when asked to identify themselves.
- Disproportionate Claude self-identification: Kimi K3 frequently identifies as Claude and sometimes emits exact internal Anthropic deployment ID strings like
claude-opus-4-5-20250929andclaude-sonnet-4-5-20250929. - More accurate than real Claude: an actual Claude Sonnet 4.5 just says "I'm Claude Sonnet 4.5"; Opus 4.5 often does not volunteer internal version strings at all — K3 reproduces deployment metadata more accurately than the teacher model states about itself.
- Generation-by-generation pattern: K3's signal points to the "Claude 4.5 era" (late 2025), not the current Fable/Mythos generation; prior Kimi K2 pointed to earlier Claude Sonnet 4 (mid-2025).
- Greenblatt's read: training on Claude data labeled with deployment metadata — API logs or metadata-tagged synthetic data — is a specific, harder-to-dismiss form of distillation than conversational mimicry.
- Important caveat: Greenblatt emphasizes this does not prove distillation occurred — identity confusion could come from data contamination, leaked system prompts, or synthetic datasets from public sources.
Citeable hard data (sources: Anthropic, Moonshot, Greenblatt analysis, 2026-07-25):
- Opus 5 pricing: $5 input / $25 output per million tokens — same as Opus 4.8, roughly half of Fable 5
- CursorBench gap: Opus 5 max effort within 0.5% of Fable 5 peak
- K3 parameter count: 2.8T total, 896 experts / 16 active per token, ~50B active-parameter equivalent
- Anthropic February accusation: 3.4M+ anomalous API interactions from Moonshot
- Weight release: K3 full weights scheduled 2026-07-27
- Timeline: Fable 5 public 7/1 → K3 launch 7/16 → White House accusation 7/22–23 → Greenblatt analysis 7/24 → Opus 5 release 7/24
Reaction on r/LocalLLaMA splits three ways: excitement that open-closed gaps are now measured in days; jokes that almost nobody can run 2.8T parameters locally; and a grounded view that K3's real selling point is price and lack of refusals, not beating Fable 5.
06 Six-Step Model Selection Guide, FAQ, and Production Close
Both stories reflect the same battlefield: performance-per-dollar. Opus 5 closes half the gap to flagship at half the price; K3 pushes open weights and low cost; the distillation fight is a public reckoning over where capability actually comes from.
- Map your workload type: daily coding agents and business automation favor Opus 5; extreme FrontierSWE or deep reasoning still points to Fable 5; ultra-long-context open scenarios warrant watching K3 after the July 27 weight drop.
- Evaluate compliance and data retention: for sensitive enterprise data, Opus 5's no-retention-by-default policy is a material advantage over Fable 5 and Mythos 5's 30-day opt-in requirement.
- Compare hard ROI metrics: 0.5% CursorBench gap plus 50% cost savings — if your agent stack runs Cursor or Claude Code workflows, Opus 5 is likely the default upgrade for H2 2026.
- Wait for independent K3 verification: before the July 27 weight release, K3 architecture and scores are "official claims plus external speculation" — do not bet compliance-sensitive production on unverified models.
- Factor supply-chain and provenance risk: distillation accusations, chip export controls, and policy-level restrictions belong in your selection documentation if you serve government or cross-border clients.
- Choose a stable deployment host: whether you route through Opus 5 API or K3, 24/7 agent gateways need dedicated compute and persistent launchd services — avoid oversubscribed shared VPS memory limits and long-connection jitter.
Q: How much cheaper is Claude Opus 5 than Claude Fable 5?
A: At comparable benchmark tiers, Opus 5 costs about half ($5/$25 vs ~$10/$50 per million input/output tokens), with CursorBench peak scores within 1%.
Q: Is Claude Opus 5 the default model on Claude Max now?
A: Yes. Since the July 24, 2026 release, Opus 5 is the Claude Max default and the strongest model available to Claude Pro users.
Q: Did Moonshot AI actually distill Kimi K3 from Claude?
A: Unconfirmed and disputed. The White House accusation lacked public evidence; researchers question the two-week timeline for deep distillation; Greenblatt's finding that K3 self-identifies as Claude with internal deployment IDs is the strongest technical indirect evidence so far, but does not conclusively prove distillation.
Q: When do Kimi K3's full weights release?
A: Moonshot committed to July 27, 2026. As of publication, weights were not yet available.
Q: Is Kimi K3 open source?
A: K3 is open-weight (downloadable once released), not fully open-source in the strictest sense — license details were pending at launch, expected to follow Moonshot's prior modified-MIT approach.
API-only access lets you swap Opus 5 or K3 quickly, but production agents still face long-connection drops on oversubscribed shared hosts, no stable 24/7 environment for multi-model A/B tests, and compliance audits requiring dedicated infrastructure — costs unrelated to token pricing but decisive for uptime. For teams running Claude Code or Kimi Code agents around the clock, JEXCLOUD multi-region bare-metal Mac hosts are the more stable layer: dedicated Apple Silicon unified memory, no oversubscription jitter, launchd-persistent gateways, 120-second provisioning. See the JEXCLOUD pricing page for specs and nodes.