Claude Opus 5: полцены от Fable 5 и Kimi K3, который называет себя Claude
Последняя неделя июля 2026 — два события, которые на первый взгляд orthogonal, но по сути про одно: price/performance и provenance модельных capability. Anthropic 24 июля выкатила daily driver Claude Opus 5 — near-Fable-5 intelligence при ~50% token cost. Kimi K3 (Moonshot AI, релиз 16 июля) попал в distillation controversy: White House accusation, Anthropic February claim про 3.4M+ anomalous API calls, и independent finding от Ryan Greenblatt — K3 self-identifies as Claude и иногда эмитит internal deployment ID strings.
Для тех, кто выбирает model stack: ① Opus 5 pricing/benchmarks/safety policy; ② полный pipeline K3 drama от political accusation до Greenblatt stats; ③ six-step framework Opus 5 vs K3. Архитектура K3: deep dive Kimi K3.
01 Два headline'а недели и три bottleneck'а при model selection
Opus 5 — «ship faster, pay half». K3 — «open weights, cheap API, fewer refusals». Distillation scandal — industry finally asking: откуда реально взялась capability?
- Flagship pricing не масштабируется. Fable 5 top-tier, но $10/$50 per 1M tokens убивает economics coding agents и automation pipelines at scale.
- Open-source model «слишком strong» без independent verify. K3: 2.8T params, GPQA-Diamond 93.5%, но full weights только 27 июля — на момент скандала external researchers не могли reproduce architecture или scores.
- Compliance + supply chain risk stack. Distillation claims, GB300 export control, Anthropic Feb 2026 fingerpointing Moonshot за 3.4M+ suspicious API interactions — enterprise selection больше не «просто смотрим leaderboard».
Opus 5 = legitimate price war response. K3 controversy = question whether cheap frontier is engineering win или covert teacher-model extraction. Same market, different accusation vector.
02 Claude Opus 5: release notes для тех, кто читает specs, не press release
Ship date: 24 июля 2026 (US Pacific). Model ID: claude-opus-5. Endpoints: Claude API, AWS Bedrock, Google Vertex AI, Microsoft Foundry. Day-one: Claude Max default + strongest Claude Pro tier.
- Pricing unchanged vs Opus 4.8: $5 / 1M input tokens, $25 / 1M output. Fast mode ~2.5× speed, 2× price.
- Context: 1M tokens (single tier only). Max output 128K. Thinking on by default; Effort knob controls reasoning budget.
- Frontier-Bench v0.1: SOTA on SWE tasks, >2× Opus 4.8, lower cost per task.
- CursorBench 3.2 max: within 0.5% of Fable 5 peak at half cost; best perf/$ across high/xhigh/max tiers vs all models on market.
- ARC-AGI 3: 3× next-best on novel problem-solving.
- OSWorld 2.0: beats Fable 5 best run at <⅓ cost.
- Zapier AutomationBench: ~1.5× pass rate vs runner-up; 100% on account-health workflow (prior models: 0%).
Life sciences delta vs Opus 4.8: spectroscopy→structure internal bench +10.2 pp; protein variant function prediction +7.7 pp. Box enterprise telemetry: data analysis +11%, due diligence +17%, overall accuracy +8%.
Alignment: Anthropic calls Opus 5 their most aligned model yet — lowest deception rate, hardest to jailbreak into misuse. Dual-use frontier (offensive cyber, biosecurity) intentionally not refreshed — Mythos 5 keeps that slot under restricted access. Cyber classifier triggers ~85% less vs Fable 5: source-level vuln discovery OK; binary scanning, pentest, exploit gen blocked.
Data retention: Opus line = no mandatory 30-day retention on default access. Fable 5 / Mythos 5 require opt-in 30-day policy — meaningful delta for compliance-sensitive workloads.
Cursor team verbatim: «Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it's just under Fable 5 and has many of the same behaviors.»
03 Opus 5 vs Fable 5: comparison table для decision makers
| Dimension | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Input / output pricing | $5 / $25 per 1M tokens | ~$10 / $50 per 1M tokens |
| Relative cost | ~50% of Fable 5 | flagship pricing |
| CursorBench 3.2 (max) | −0.5% vs Fable 5 peak | peak reference |
| Context window | 1M tokens | 1M tokens |
| Data retention | no mandatory 30-day retention | 30-day retention required |
| Claude Max default | yes (from 2026-07-24) | no |
| Dual-use risk capability | deliberately below Mythos 5 | stronger; Mythos 5 = restricted apex |
| Use case fit | daily coding agents, automation, compliance-sensitive | FrontierSWE, limit reasoning |
If Fable 5 was your benchmark but not your budget — Opus 5 is the «−0.5% score, −50% bill» swap Anthropic needed for the current price war.
04 Kimi K3 distillation scandal: White House enters chat
Moonshot shipped Kimi K3 16 июля 2026: 2.8T total params, first open-weight «3T-class» model; sparse MoE 896 experts / 16 active per token (~50B active-equivalent); 1M context + native vision; KDA + Attention Residuals + Stable LatentMoE (~2.5× training efficiency vs K2). Full weights promised 27 июля 2026 — during controversy, zero independent architecture verification.
| Benchmark | Kimi K3 | Note |
|---|---|---|
| GPQA-Diamond | 93.5% | best open-weight score at launch |
| Terminal-Bench 2.1 | 88.3% | −0.5 vs GPT-5.6 Sol |
| BrowseComp | 91.2% | launch high score |
| Program Bench | 77.8% | overall best |
| SWE Marathon | 42.0% | overall best |
| DeepSearchQA (F1) | 95.0% | — |
22–23 июля 2026: OSTP director Michael Kratsios posted on X accusing Moonshot of «large-scale, covert industrial distillation» to steal Anthropic Fable capabilities, plus alleged export-restricted Nvidia GB300 chips (possibly routed via Thailand). Treasury Secretary Scott Bessent: «watermarks» of US LLMs found on Chinese models — no specifics.
Not new ground: February 2026 Anthropic publicly named Moonshot, DeepSeek, MiniMax for «industrial-scale distillation attacks» — 3.4M+ anomalous API interactions from Moonshot, «clearly deviating from normal usage patterns, reflecting deliberate capability extraction», partial trace to senior Moonshot staff via request metadata. Moonshot never confirmed or denied.
TechCrunch (23 July) interviewed researchers skeptical of the two-week distillation timeline — Fable 5 public only since 1 July, K3 launch 16 July:
«I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation... Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks.» — Braden Hancock, Laude Institute / Snorkel AI co-founder
«Distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to reinforcement learning... if it were the case, everyone would be easily able to catch up to a GLM or a K3 by using its data for distillation. But we have not.» — Nathan Lambert, Allen Institute for AI
Elon Musk testified xAI distilled OpenAI models training Grok — «common industry practice». Distillation happens; dispute is boundary between technique borrowing and covert industrial extraction.
05 K3 calls itself Claude: Greenblatt forensic analysis
Most technically interesting angle of the whole saga. Redwood Research chief scientist Ryan Greenblatt (~24 July) published statistical analysis (GitHub: rgreenblatt/which_claude_is_k3) comparing identity responses to «who are you?» across models.
- Anomalous Claude self-ID rate: Kimi K3 disproportionately identifies as Claude — sometimes emitting exact Anthropic internal deployment ID strings like
claude-opus-4-5-20250929,claude-sonnet-4-5-20250929. - More accurate than actual Claude: real Claude Sonnet 4.5 says «I'm Claude Sonnet 4.5»; Opus 4.5 doesn't volunteer internal version strings or gets them wrong. K3 reproduces deployment metadata more accurately than the teacher model states about itself.
- Generation tracking: K3 signal locks to «Claude 4.5 era» (late 2025), not current Fable/Mythos. Prior K2 pointed at Claude Sonnet 4 (mid-2025) — generation-over-generation chase pattern.
- Greenblatt read: hard to explain as conversational mimicry; more consistent with training on Claude data labeled with deployment metadata — API logs, metadata-tagged synthetic data — specific distillation form harder to hand-wave.
- Critical caveat: Greenblatt explicitly states this does not directly prove distillation occurred. Identity confusion could come from data contamination, system prompt leak, or public dataset synthesis.
Citable hard numbers (Anthropic, Moonshot, Greenblatt — 2026-07-25):
- Opus 5 pricing: $5 / $25 per 1M tokens; ~50% of Fable 5
- CursorBench gap: Opus 5 max vs Fable 5 peak = 0.5%
- K3 params: 2.8T total; 896/16 MoE; ~50B active-equivalent
- Anthropic Feb accusation: Moonshot 3.4M+ anomalous API interactions
- Weight release: K3 planned 2026-07-27
- Timeline: Fable 5 public 07/01 → K3 07/16 → White House 07/22–23 → Greenblatt ~07/24 → Opus 5 07/24
r/LocalLLaMA split three ways: hype («open/closed gap now measured in days»), memes («2.8T — nobody runs this locally»), pragmatists («K3 sell is price + fewer refusals, not beat Fable 5»).
06 Six-step model pick, FAQ, production wrap
Main battlefield H2 2026: perf/$. Opus 5 halves cost near flagship. K3 pushes open-weight pricing. Distillation drama = public fight over capability provenance.
- Map workload: daily coding agents / biz automation → Opus 5; limit FrontierSWE / deep reasoning → Fable 5; long-context open source → K3 post-07/27 weights.
- Audit compliance + retention: Opus 5 no mandatory 30-day retention = real enterprise advantage vs Fable/Mythos.
- Benchmark perf/$: CursorBench −0.5% + 50% cost save — if your stack is Cursor / Claude Code shaped, Opus 5 is default H2 2026 upgrade.
- Wait for K3 independent verify: pre-07/27 architecture/scores = vendor self-report only; don't bet compliance-critical prod on unverified model.
- Document supply chain risk: distillation claims, chip export, policy restrictions — log provenance in architecture docs.
- Deploy on stable host: 7×24 agent gateways need dedicated Apple Silicon unified memory + launchd persistence — not oversubscribed shared VPS with long-connection jitter.
Q: Насколько Claude Opus 5 дешевле Claude Fable 5?
A: Same benchmark tier — Opus 5 ~50% ($5/$25 vs ~$10/$50 per 1M input/output); CursorBench peak gap <1%.
Q: Claude Opus 5 — default на Claude Max?
A: Да, с 24 июля 2026 — также strongest Claude Pro tier.
Q: Moonshot дистиллировала Kimi K3 из Claude?
A: Unconfirmed/disputed. White House без public evidence; timeline skeptics; Greenblatt stats = strongest technical circumstantial evidence, not conclusive proof.
Q: Когда полные веса Kimi K3?
A: Moonshot: 27 июля 2026 — на дату статьи not yet released.
Pure API switching быстро, но prod agents hit hidden costs: shared cloud overselling → long-connection drops, no 7×24 stable host for multi-model A/B, compliance audit needs dedicated environment. For always-on Claude Code / Kimi Code agent stacks, JEXCLOUD multi-region bare-metal Mac — dedicated Apple Silicon unified memory, no oversell jitter, launchd-persistent gateway, 120s provisioning. Nodes/pricing: JEXCLOUD pricing page.