Latest News

Five Frontier Models in One Week: What Actually Changes in Your Agent Stack

Rushil ShahRushil Shah
8 min read
Share

Anthropic, OpenAI, xAI and Xiaomi shipped five frontier or near-frontier models inside a single week. The leaderboard shuffle is the least interesting part. The repricing of cache reads, the 50% cut at the mid and low tiers, and the quiet arithmetic of fallback rates are what actually land in your invoice next month.

TL;DR

Between September 21 and 23, four labs shipped five frontier or near-frontier models. The intelligence-index reshuffle matters less than three pricing facts: Opus 5.5 cut cache reads 60% to $0.20/M, OpenAI halved Sol and Luna list prices, and Luna now sits at $0.10/$0.50 per million. If your routing table and eval suite were tuned before last Monday, they are now mispriced — probably by a lot, and possibly in the direction of overspending on tasks a 20x cheaper model closes.

What shipped, and when

This was not a coincidence week. It was four labs landing inside 72 hours of each other, which tells you something about how release calendars are now set.

Sep 21

xAI ships Grok 4.7, a larger base model with a longer RL run weighted toward multi-hour tasks, at $2/$6 per million tokens.

Sep 22

Anthropic releases Claude Opus 5.5 — Fable 5.1-level performance, 40% cheaper to run than Opus 5 on typical workloads by Anthropic's own measure.

Sep 22

OpenAI expands the GPT-6 line with Sol and Luna at a 50% price cut versus their GPT-5.6 equivalents, plus a materially better caching story.

Sep 23

Xiaomi open-sources the MiMo-V2.6 series, with Pro scoring 46 on the Artificial Analysis Intelligence Index — the strongest open-weights model available, by their claim.

Pat McGuinness's roundup of the week is the best single read on the shape of it, and his framing is right: four of the five sit on the Pareto intelligence-cost frontier, and all five have a use case. But roundups optimise for "which model is best." That is the wrong question if you operate agents. The right question is which line items on your invoice just moved.

The real headline is the cache line

Everyone quotes input and output prices. In a production agent loop, neither is the dominant cost. Cached context reads are — and Anthropic said so explicitly, describing cache reads as making up the majority of agentic and coding work costs before dropping them from $0.50 to $0.20 per million, a 60% cut.

$0.20Opus 5.5 cache reads per 1M tokens, down 60% from Opus 5Source: Anthropic, 2026
$0.10 / $0.50GPT-6 Luna input/output per 1M tokens after a 50% cutSource: OpenAI, 2026
90%discount on cached input-token reads for GPT-6Source: OpenAI, 2026
46MiMo-V2.6-Pro AA Intelligence Index, top open-weights scoreSource: Xiaomi MiMo, 2026

OpenAI went further than a discount. Two changes in the GPT-6 caching update are architectural, not commercial: you can now change reasoning effort mid-conversation and enable or disable tools without invalidating the cached prefix, and explicit breakpoints let you choose where a cached prefix ends.

If you have ever built an escalation ladder — run cheap, detect low confidence, re-run harder — you know why that matters. Until now, bumping effort or swapping the tool list rewrote the prefix and you paid full freight for context you had already sent. Dynamic effort escalation was theoretically obvious and economically stupid. It is now neither. That single change probably alters more production architectures this quarter than any benchmark on any chart.

Cost per token is not cost per task

Here is where the price war gets slippery. Grok 4.7 lists at $2/$6 — cheaper per token than GPT-6 Sol's $2/$10. McGuinness's read is that Sol is more token-efficient and ends up cheaper on balance, and that Grok 4.7 burns more tokens than 4.6 for little gain. Anthropic makes the same argument in its own favour: Opus 5.5 costs less per token and uses fewer tokens per task, which is how a 20% list cut becomes a claimed 40% workload cut.

ModelList price (in/out per 1M)Where it earns its slot
Claude Opus 5.5$4 / $20 (cache reads $0.20)Long-horizon migrations, codebase-wide audits, anything where rework cost dwarfs token cost
GPT-6 Sol$2 / $10Mid-tier agentic workhorse; strong on business-workflow benchmarks per OpenAI's AutomationBench numbers
GPT-6 Luna$0.10 / $0.50High-volume routine steps: extraction, classification, transformation, routing decisions
Grok 4.7$2 / $6Niche knowledge work — it leads the comparison set xAI published on the Harvey legal benchmark at 19.6%
MiMo-V2.6-ProOpen weights; ~$0.435 / $0.87 hosted, per McGuinnessCost-sensitive volume where you want weights you control

What this breaks in a stack you are already running

!

Re-routing everything to the cheapest tier that passed your eval

Luna at a tenth the price of Sol makes the spreadsheet sing. But cheap-tier models fail differently, not just more often: they fail late, mid-trajectory, after eleven tool calls, and the cleanup cost never shows up in a per-token comparison. A 3% increase in failed multi-step runs can erase a 90% token saving once you count retries, human triage and downstream data corrections.

Fix: measure cost per successfully completed task on your own traces, with retries and human-intervention minutes priced in. Migrate one workflow class at a time, not the router default.
!

Comparing benchmark numbers across vendor pages

Every score in this week's releases carries an effort level — medium, high, xhigh, max — and cost swings by an order of magnitude across them. Anthropic reported Terminal-Bench at xhigh for its own model and high for the competitor. OpenAI compared Sol at xhigh against Opus 5 at max. Both disclosed it. Neither disclosure survives a screenshot on social media.

Fix: treat vendor charts as hypothesis generators, never as procurement evidence. The only number that decides your routing table is the one from your own harness on your own task distribution.

At these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences.

— Anthropic, Claude Opus 5.5 announcement

The fallback rate nobody is pricing

This is the most operationally important detail of the week, and it is buried in footnotes on two different vendor pages.

Anthropic notes that Opus 5.5 ships with safeguards similar to Fable 5.1's, and that in its own benchmark runs, when safeguards intervened, cybersecurity tasks were completed by Opus 4.8 and biology tasks by Opus 5. It also notes that Zapier's AutomationBench runs were performed without fallback models, so safeguard interventions were considered failures. OpenAI, comparing on the same benchmark, observes that a competitor datapoint understates real cost because Opus 5 fallbacks occurred on roughly 40% of tasks.

Strip away the vendor sniping and the lesson is the same for both: at frontier tier, a non-trivial fraction of agent steps may not be served by the model you selected. That has consequences a benchmark table cannot express — different latency, different token accounting, different behaviour under the same prompt, and an audit trail that says one model while another did the work. If you run agents in a regulated context, your model-attestation logging needs to record what actually served the request, not what your config requested.

What I'd do this week

A two-day re-benchmark that is actually worth running

1
Pull 200 real traces per workflow class

From production logs, not a curated set. Include the ones that failed.

↓
2
Replay against Luna, Sol and Opus 5.5 at default effort

Default, not max. Default is what you will run in production, and it is what Anthropic benchmarked its cost advantage on.

↓
3
Score cost per completed task, not accuracy

Include retries, fallback events, and the human minutes spent on failures.

↓
4
Instrument cache hit rate before you touch the router

OpenAI now ships a caching dashboard and diagnostics. If your hit rate is under 60%, fixing prefix ordering will beat any model swap.

↓
5
Move one class, watch for two weeks, then move the next

Big-bang routing changes make regressions impossible to attribute.

We run this pattern for clients as a standing quarterly exercise rather than a reaction to release weeks — see how we structure it under automation and model selection.

The honest caveats

Three things this week is carefully not saying. First, the cost claims are workload-dependent: Anthropic's 40% figure is measured at default settings on typical workloads, and McGuinness reports that real user metrics suggest a cost closer to Opus 5. Your mileage depends entirely on your cache hit rate and effort settings.

Second, availability is not uniform. OpenAI shipped Sol and Luna to ChatGPT Work, Codex and the API, but not yet to the standard Chat experience. Opus 5.5 is available across AWS, Google Cloud and Azure on day one, with Sonnet 5.5 and Haiku 5.5 still to come — meaning the cheap end of the Claude family has not yet been repriced.

Third, open weights are not free inference. MiMo-V2.6-Pro is genuinely impressive and genuinely open, with a published technical report and weights on Hugging Face. It is also described in user reports as token-hungry and slow to think. Self-hosting a 1T-parameter MoE to save on a workload that GPT-6 Luna handles for $0.10 per million input tokens is a decision that needs a spreadsheet, not an ideology.

The frontier moved this week. Your unit economics moved further. If you want a second pair of eyes on the routing table before the invoice arrives, start here.

Frequently Asked Questions

Which model should be our default after this week?

There is no single default any more, and that is the point. Route by task class: Opus 5.5 for long-horizon coding and migrations where rework cost dominates, GPT-6 Sol for mid-complexity agentic workflows, GPT-6 Luna for high-volume routine steps like extraction and classification. Decide with a replay of your own production traces measured on cost per completed task, not with vendor benchmark tables.

Is Claude Opus 5.5 really 40% cheaper to run than Opus 5?

Anthropic's figure applies at default settings on typical workloads and combines a 20% list-price cut with lower token consumption per task. Cache reads dropped 60% to $0.20 per million, which is where most agentic spend actually sits. Independent commentary suggests real-world costs land closer to Opus 5 for some users. The gap is mostly explained by cache hit rate and effort level, both of which you control.

Does GPT-6 Luna change the unit economics of high-volume automation?

Materially, yes. At $0.10 input and $0.50 output per million tokens, workflows that were marginal at GPT-5.6 pricing become comfortably profitable. The caveat is failure mode, not capability: cheaper models fail later in multi-step trajectories, and cleanup cost is invisible in per-token math. Pilot on one workflow class, price retries and human triage into the comparison, then expand.

What is the fallback-rate issue and why does it matter?

At frontier tier, safeguard systems can route individual requests to a different underlying model. Anthropic disclosed that safeguard interventions during benchmarking shifted some tasks to Opus 4.8 and Opus 5; OpenAI pointed to roughly 40% fallback rates in a competitor's benchmark run. For production teams this affects latency, cost accounting and auditability. Log which model actually served each request, not which one your config requested.

Should we self-host MiMo-V2.6-Pro?

Only if data residency, weight control or a specific fine-tuning plan drives the decision. MiMo-V2.6-Pro scores 46 on the AA Intelligence Index, the best open-weights result available, and it is available via hosted API and on Hugging Face. But it is reported as token-hungry and slow, and serving a trillion-parameter MoE has real infrastructure cost. For pure cost savings on routine work, hosted cheap tiers are usually the shorter path.

frontier modelsmodel routingagent architectureAI cost optimizationClaude Opus 5.5GPT-6open weightsLLM benchmarking

Published

AI-assisted writing · Reviewed by the Twarx research team

Share:
Share

Research digest

AI Research Briefing

Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.

  • No hype, just measurable outcomes
  • Read by 2,400+ engineers
  • Unsubscribe anytime