Anthropic, Meta, Google and OpenAI each shipped a frontier model between September 1 and September 3, according to the companies’ own announcements.
The four launches landed within 72 hours. Anthropic and OpenAI matched headline token rates to the dollar, while standard cache rates across the four ranged from $0.075 to $1.00 per million tokens.
That spread matters because software agents reread cached instructions on every cycle. CNBC reported that the pace has left executives and IT managers spending disproportionate time comparing costs and capabilities.
Anthropic and OpenAI landed on the same headline rate two days apart
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1 at $10 per million input tokens and $50 per million output tokens.
The company held those rates from Fable 5 and cut cache reads by 75% to $0.25 per million tokens, according to Anthropic.
OpenAI released GPT-6 Astra two days later at the same $10 and $50, with cached input at $1.00 per million and cache writes at $12.50, according to OpenAI.
Those are short-context rates. OpenAI’s long-context tier charges $20 for input, $2.00 for cached input, and $75 for output.
More on how AI is reshaping bank economics:
Agentic AI Is About to Move Deposits for Customers. Banks Have Months, Not Years, to Respond
Meta priced cache at $0.15 and offered a discount for training rights
Meta released Muse Spark 1.3 on September 2 and published no prices in the announcement, citing internal comparisons showing roughly 20% fewer tool calls than version 1.2, according to Meta Superintelligence Labs.
Its API documentation lists $1.25 input, $4.25 output, and $0.15 cached input, and states that no long-context premium applies.
Meta also sells a Contributor tier at $0.10 input and $0.002 cached input, in exchange for permission to train on customer prompts, according to Meta’s pricing documentation. That tier discounts token charges in exchange for training rights.
Google prices cache lowest but bills stored caches by the hour
Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on September 2. Flash carries introductory rates of $0.75 for input, $3.75 for output, and $0.075 for context caching.
Those rates expire on December 31, 2026, and then double, according to Google’s API pricing.
Google runs two caching modes. Implicit caching applies automatically and passes on savings with no storage charge.
Explicit caching, where the developer sets how long tokens persist, adds storage billed at $0.50 per million tokens per hour, rising to $1.00, according to Google’s caching documentation.
No other lab bills cache by duration. Flash Cyber carries no published rate and is restricted to vetted defenders through the Fairwind program, according to Google.
Two IPO candidates are setting the release pace
Model developers are competing for share of wallet and racing to signal that they innovate as fast as their rivals, Notre Dame business professor Ahmed Abbasi told CNBC.
“We’re all moving to faster cadences.” — Sam Altman, chief executive, OpenAI, via CNBC

Anthropic and OpenAI are pushing hardest as both move toward public listings, CNBC reported.
Anthropic was valued at $965 billion post-money in its May 28 Series H, according to Anthropic. OpenAI closed a round at $852 billion post-money on March 31, according to OpenAI.
Key pricing and spending figures from the September release week
- Standard-tier cached input priced at $0.075 by Google, $0.15 by Meta, $0.25 by Anthropic, and $1.00 by OpenAI per million tokens (company pricing pages)
- Google alone charges storage on explicit caches at $0.50 per million tokens per hour, doubling on January 1 (Google)
- Gemini 3.8 Flash introductory rates of $0.75 and $3.75 rise to $1.50 and $7.50 on January 1, 2027 (Google)
- OpenAI’s long-context tier doubles input to $20 and raises output to $75 (OpenAI)
- Hugging Face deal structured as $11.9 billion to stockholders plus up to $1.0 billion retention equity, closing expected H1 2027 (Nvidia Form 8-K)
- Worldwide AI spending forecast at $2.59 trillion in 2026, up 47% (Gartner, May 19, 2026)
Cached tokens carry most of an agent’s bill
Agents reread the same instructions, code, and tool history on every cycle. More than 85% of the tokens they burn come from the cached prompt, according to OpenRouter data published by a16z.
The same data put agent token consumption at nearly five times human usage, growing roughly 14 times since February.
What that means for the invoice is a separate question. A category’s share of spending is its volume multiplied by its rate, and no published figure measures how cached tokens split the bill.
Gartner used the same week to warn corporate legal departments about consumption billing. Senior Director Analyst Shannon Nakamoto said in a September 3 release:
“Pricing transparency should be a critical buying criterion for GC. It’s essential to understand how AI consumption is measured, what drives costs and how spending will change as adoption grows.” — Shannon Nakamoto, Senior Director Analyst, Gartner Legal and Compliance Practice, via Gartner
Gartner predicts consumption-based pricing will account for more than 35% of net new corporate legal technology spend with major vendors by 2028.
Nvidia bought the distribution layer in the same week
Nvidia entered a definitive agreement on September 2 to acquire Hugging Face, the platform where open models are hosted and downloaded, announcing it the following day.
The deal is approximately $11.9 billion payable to stockholders plus up to about $1.0 billion in retention equity, expected to close in the first half of 2027 subject to regulatory approvals, according to Nvidia’s Form 8-K.
Nvidia closed at $230.36 on September 4, up 0.84% in the session and 6.23% over the month, according to TradingView.
Its $5.562 trillion market capitalization is the largest of any listed company, ahead of Apple at $4.669 trillion, according to CompaniesMarketCap.

Alphabet Class C closed at $335.31, down 1.11% on the session and 11.76% over the month, according to TradingView.

What four price sheets mean for companies buying AI
Four frontier models arrived in 72 hours. Anthropic and OpenAI matched headline rates to the dollar, while standard cache rates ranged from $0.075 to $1.00 per million tokens. Google’s introductory pricing doubles on January 1, 2027, and it alone bills stored caches by the hour. Companies running agents should model their own mix of cached input, uncached input, cache writes, output, and storage rather than compare headline rates.
Two things to track: Google’s January price step, and the approvals Nvidia’s Hugging Face purchase must clear before its expected close.





