AI Models in August 2026: Claude Opus 5, Kimi K3, and the Month the US-China AI Gap Closed
The complete August 2026 AI model guide. Claude Opus 5 lands near Fable 5 at half the price, Moonshot's 2.8 trillion parameter Kimi K3 rattles Wall Street, GPT-5.6 gets a price cut, and Grok 4.5 ships under the new SpaceXAI brand. Every release, every benchmark, explained.

TL;DR
Claude Opus 5 launched July 24 at the same price as Opus 4.8, $5/$25 per million tokens, and lands near Claude Fable 5 on intelligence for half the cost. On the Artificial Analysis Intelligence Index it now sits at the top, just ahead of Fable 5 and GPT-5.6 Sol.
Moonshot AI released Kimi K3, a 2.8 trillion parameter open weight model, on July 27. It beats Claude Opus 4.8 and GPT-5.5 on several coding and agent benchmarks at roughly a third of Opus 4.8's cost, though it still trails Fable 5 and GPT-5.6 Sol. The release briefly moved Nasdaq chip stocks.
OpenAI's GPT-5.6 family, Sol, Terra, and Luna, reached general availability July 9 after a government-coordinated preview. On July 30, OpenAI cut Luna's price by 80% and Terra's by 20%, pushing the entry cost of frontier-adjacent intelligence lower than ever.
xAI shipped Grok 4.5 on July 8 under a new SpaceXAI brand following SpaceX's acquisition of Cursor. Built on a Musk-claimed 1.5 trillion parameter V9 foundation, it targets coding and agentic work at $2/$6 per million tokens, roughly 60% cheaper than Opus 4.8.
Sam Altman spent late July on Capitol Hill previewing OpenAI's next model family to federal officials and confirmed OpenAI deactivated a model after a Hugging Face security breach. Five national cybersecurity agencies jointly published guidance on agentic AI risk the same week.
The assumption that Chinese labs trailed the US frontier by six to twelve months no longer holds for open weight coding and agent models. Kimi K3 and GLM-5.2 have closed that gap to single digits on several benchmarks, even as the very top of the leaderboard stays with Anthropic and OpenAI.
AI Models in August 2026: Claude Opus 5, Kimi K3, and the Month the US-China AI Gap Closed
Six weeks ago, a wave of export control drama took two of Anthropic's newest models offline worldwide and forced OpenAI to launch its next flagship behind a government-approved guest list. That should have been the biggest AI story of the summer. It was not even close.
Since then, Anthropic has shipped a fourth model in under two months, xAI has rebranded around a new corporate structure and a bigger foundation model, and a Beijing startup with a research paper and a Friday press release moved the Nasdaq. This is the complete picture of where AI models stand as of early August 2026, what actually changed, and what it means for the ongoing argument about whether the United States or China is winning the AI race.
Looking for the full picture from earlier in the summer? See our guides to AI Models in July 2026 and AI Models in June 2026.
The Leaderboard: Where Things Stand Right Now
Here is the current picture from the Artificial Analysis Intelligence Index, the most widely cited composite benchmark in the industry, as tracked through early August 2026.
| Rank | Model | Creator | AA Index | Blended price (1M tokens) | Context | Released |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 60.7% | $5.00 in / $25.00 out | 1M | Jul 24 |
| 2 | Claude Fable 5 | Anthropic | 59.9% | $10.00 in / $50.00 out | 1M | Jun 30 (restored Jul 1) |
| 3 | GPT-5.6 Sol | OpenAI | 58.9% | $5.00 in / $30.00 out | Unconfirmed | Jul 9 (GA) |
| 4 | Grok 4.5 | SpaceXAI | 54.0% | $2.00 in / $6.00 out | 500K | Jul 8 |
| N/A | Claude Sonnet 5 | Anthropic | Not separately indexed | $2.00 in / $10.00 out (through Aug 31) | 1M | Jun 30 |
| N/A | Kimi K3 | Moonshot AI | Not yet indexed | Open weight, self-hosted or via API | Unconfirmed | Jul 27 |
| N/A | GLM-5.2 | Zhipu / Z.ai | Not yet indexed | $1.40 in / $4.40 out | 1M | Jun 13 |
| N/A | Mythos 5 | Anthropic | Restricted access | Not publicly priced | 1M | Restricted since Jun 12 |
A note on how to read this table. The Artificial Analysis Index reflects models with enough independent evaluation volume to score reliably. Kimi K3 and GLM-5.2 are recent enough, and evaluated across different benchmark suites, that they have not yet settled into a single directly comparable index number the way the closed frontier models have. That does not mean they are weak. It means the industry's standard measuring stick has not fully caught up to them yet, which is itself part of this month's story.
Claude Opus 5: Near Fable 5 Intelligence at Half the Price
Anthropic released Claude Opus 5 on July 24, its fourth model release in under two months following Mythos 5, Fable 5, and Sonnet 5 in June. The pitch is direct: intelligence that approaches the Fable 5 frontier, at exactly half of Fable 5's price.
Standard pricing is $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8 and half of Fable 5's $10/$50. A Fast mode runs at roughly 2.5 times the speed for double the base price. The model ships with a 1 million token context window, a 128,000 token maximum output, and a knowledge cutoff of May 2026, the most current of any Claude model to date. A new low, medium, and high effort toggle lets developers trade reasoning depth against token spend on a per-request basis.
On Anthropic's own Frontier-Bench v0.1 evaluation, Opus 5 more than doubles Opus 4.8's score, and on ARC-AGI-3 it scores roughly three times higher than the next-best model tested. It also posts the lowest rate of deceptive behavior Anthropic has measured in any of its models, which the company is positioning as evidence of improved alignment alongside the capability gains.
The honest caveats matter here. Opus 5 does not win everything. GPT-5.6 Sol still edges it out on one agentic coding benchmark, and Fable 5 keeps its lead on legal and health and biology evaluations. One detail that will matter more than any benchmark for regulated buyers: Opus 5 carries no mandatory data retention requirement for general access, unlike Fable 5 and Mythos 5, which both require 30-day retention with no zero-data-retention option, even on EU endpoints.
Opus 5 is now the default model on Claude Max and the strongest model offered on Claude Pro, so most paying Anthropic subscribers already have access without changing anything.
Claude Fable 5 and Mythos 5: Life After the Export Ban
The strangest three weeks any AI model has had this year belong to Fable 5. Launched June 9 as the first publicly available model in Anthropic's new Mythos tier, a class of model sitting above Opus, it was suspended by US export control order on June 12 after researchers reported a jailbreak that let the model generate functional cyberattack code. Because Anthropic had no way to verify the citizenship of API and web users in real time, both Fable 5 and its more capable, fully unrestricted sibling Mythos 5 went dark globally.
The directive against Fable 5 was fully lifted June 30, and the model returned to worldwide availability on July 1 with an enhanced safety classifier designed to detect and block the specific jailbreak pattern that triggered the ban. Mythos 5 did not get the same treatment. It remains restricted to a list of more than 100 vetted US critical-infrastructure organizations, government agencies, and private companies working on defense and security applications. There is no confirmed timeline for broader release.
Fable 5 today costs $10 per million input tokens and $50 per million output tokens, with a 1 million token context window, and it still holds the top published score on several of the hardest evaluations Anthropic tracks, including legal reasoning and health and biology. For teams that need the single most capable model available and can absorb the cost and the mandatory data retention terms, it remains the choice. For everyone else, Opus 5 at half the price is now the more sensible default.
GPT-5.6 Reaches General Availability, Then Gets Cheaper
OpenAI's GPT-5.6 family, structured into three tiers named Sol, Terra, and Luna, previewed June 26 behind a government-coordinated access list of roughly 20 approved organizations, a direct consequence of a June 2026 executive order requiring federal review of new frontier models before wide release. OpenAI has been candid that it participated voluntarily and does not want this kind of arrangement to become permanent, but it complied for this launch.
General availability followed on July 9. Then, on July 30, OpenAI cut prices again: Luna dropped 80% and Terra dropped 20%, a significant move that pushes the cost of frontier-adjacent intelligence to some of the lowest levels the market has seen from a major US lab.
The headline benchmark result is on Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields. GPT-5.6 Sol sets a new high score there, beating Claude Fable 5 by 13.1 points. Even Sol running at medium reasoning effort beats Fable 5 by 11.4 points, at roughly a quarter of the estimated cost. That efficiency carries down the lineup: OpenAI says Terra and Luna both outperform Fable 5 on the same evaluation at around one-sixteenth the cost.
Sol Ultra, the highest-compute tier, has also started shipping inside OpenAI's Codex client, putting the top reasoning configuration directly into the hands of developers who work inside that environment day to day.
Grok Becomes SpaceXAI: Grok 4.5 and the Cursor Acquisition
xAI shipped Grok 4.5 on July 8, the first model released under the increasingly common SpaceXAI branding that followed SpaceX's acquisition of the AI coding tool Cursor in June 2026. The strategic logic is straightforward: fold a widely used coding product directly into the model training pipeline.
Grok 4.5 is built on a foundation Elon Musk has publicly described as 1.5 trillion parameters and codenamed V9, trained on tens of thousands of Nvidia GB300 GPUs at the Colossus cluster in Memphis. It is worth being precise here: xAI's own official launch materials and model documentation have not confirmed a parameter count. Artificial Analysis's own model page states explicitly that SpaceXAI has not disclosed the model size. Treat the 1.5 trillion figure as Musk-attributed rather than independently verified.
What is confirmed: the model was co-trained with Cursor on trillions of tokens of real developer session data, not just static code samples, which is a genuinely different training approach from anything else on the market. It supports a 500,000 token context window, smaller than Grok 4.3's 1 to 2 million, and is priced at $2 per million input tokens and $6 per million output tokens, roughly 60% cheaper than Claude Opus 4.8.
On the Artificial Analysis Intelligence Index it scores 54, placing it fourth overall at launch, behind Fable 5, GPT-5.5, and Opus 4.8 in the pre-Opus-5 ranking. Its published results are mixed by design: it trails badly on SWE-Bench Pro (64.7% versus Fable 5's 80.4%) but leads outright on SWE Marathon, a benchmark measuring whether a model can hold a coherent plan across dozens of steps without drifting, beating both Opus 4.8 and Fable 5 on that specific test. For agent builders who care more about long-horizon reliability than single-bug fixes, that is the number that matters.
The Kimi K3 Shock: Open Weight, 2.8 Trillion Parameters, and a Market Reaction
If one release defines August 2026 so far, it is this one.
Moonshot AI, the Beijing-based startup behind the Kimi model family, released Kimi K3 on July 27 as a fully open weight model, meaning anyone can download it, modify it, and run it on their own infrastructure for free. Financial Times reporting ahead of the release pegged the parameter count between 2 and 3 trillion, and Moonshot has since confirmed 2.8 trillion, making it the largest open weight model released by a Chinese lab to date.
The benchmark claims are the part that got attention outside the AI industry. Moonshot's own published results show K3 beating Claude Opus 4.8 and GPT-5.5 on several coding and general agent benchmarks, while costing roughly a third of what Opus 4.8 charges. By the company's own account, K3 still trails Claude Fable 5 and GPT-5.6 Sol on overall performance, a distinction Moonshot was careful to state publicly rather than overclaim.
The market reaction was immediate and telling. The Nasdaq dropped roughly 1% on the news, with investors selling shares in chipmakers including Nvidia and Intel. The logic, whether or not it holds up over time, was straightforward: if a Chinese startup can release a model this capable, this cheaply, and give it away for free, the assumed moat around premium US compute and premium-priced closed models looks less secure than markets had priced in.
Industry reaction split along predictable lines. AI analyst Kim Isenberg described the release as changing "the entire game," while others pushed back on the framing. Patrick Moorhead, a technology analyst, told CNBC that some of the reaction was overblown and rooted in Washington politics rather than technical substance, noting the irony of debating whether US companies should be allowed to help Chinese labs improve their models when "the Chinese seem to be doing fine with their models" already. Simon Koser, chief product officer at AI startup Tzafon, called K3 legitimately impressive on its own terms.
Moonshot is reportedly raising a fresh funding round that would value the company at $31.5 billion, up sharply from the $20 billion valuation it carried after a $2 billion raise in May.
For US policymakers, K3 lands at an uncomfortable moment. American AI leaders had taken comfort in estimates that Chinese labs remained six to twelve months behind the frontier, even as Chinese open weight models gained ground steadily through the spring and early summer. K3 suggests that cushion has collapsed faster than expected, at least in the specific categories of coding and general agent tasks where it was benchmarked.
GLM-5.2 and the Rest of the Open Weight Field
Kimi K3 is the headline, but it did not arrive in a vacuum. Zhipu's GLM-5.2, released under a permissive MIT license on June 13, had already pushed open weight coding performance further than most Western observers expected, scoring 62.1% on SWE-bench Pro at a blended price of $1.40 input and $4.40 output per million tokens, with a 1 million token context window. For several weeks in June, while Fable 5 and Mythos 5 sat suspended and GPT-5.6 remained gated behind government approval, GLM-5.2 briefly led the publicly accessible coding leaderboard outright, alongside MiniMax's M3 model, released June 1.
The gap between GLM-5.2's 62.1% and Fable 5's 80.4% on SWE-bench Pro is still real, roughly 18 percentage points on that specific benchmark. But it is also the smallest that gap has ever been between an MIT-licensed open model and the best closed model available, and for teams that need to self-host for data sovereignty reasons, GLM-5.2 remains the strongest option that does not require any API dependency at all.
DeepSeek has kept pace on its own timeline. DeepSeek-V4-Flash-0731, released July 31, is the latest point release in that lineup, continuing the company's pattern of frequent, incremental open weight updates rather than the larger, less frequent jumps other labs favor.
Microsoft's Quiet Cybersecurity Play
Away from the flagship model race, Microsoft made its first major cybersecurity push since a leadership shake-up earlier this year, unveiling MAI-Cyber-1-Flash in late July. Paired with OpenAI's general-purpose GPT-5.4, Microsoft says the combination outperforms Anthropic's Mythos 5, Google's Gemini 3.5 Flash Cyber, and OpenAI's own GPT-5.5 Cyber on the CyberGym benchmark, at half the cost of those alternatives. Mustafa Suleyman, CEO of Microsoft AI, put it bluntly at a San Francisco event: "We have world-leading performance at 50% of the cost."
The model will power Project Perception, Microsoft's vulnerability discovery and patching tool, which begins early testing in August. The timing is notable given Microsoft shares have fallen 19% so far in 2026, with analysts citing investor concern that the rise of capable, cheap open weight models undercuts the value of Microsoft's heavy investment in OpenAI specifically.
The Regulatory Backdrop: Pacing, Hearings, and a Joint Security Warning
The policy story running underneath all of this is worth understanding on its own terms, because it shapes what gets released, when, and to whom.
Sam Altman spent the final week of July on Capitol Hill, previewing OpenAI's next model family to senior Trump administration officials including Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick. In interviews around those meetings, Altman confirmed that OpenAI deactivated a model after last month's Hugging Face security breach, and when asked directly whether other OpenAI systems could have been compromised by the same incident, his answer was blunt: "I mean, there could be, yeah." He also said he broadly agrees with the principles behind the "Pacing the Frontier" letter, a call for government help building an international framework to manage the speed of AI development, and confirmed OpenAI's own researchers were involved in drafting it. The company separately raised its 2026 capital expenditure outlook to a range of roughly $130 to $145 billion.
The same week, cybersecurity and intelligence agencies from the United States, Australia, Canada, New Zealand, and the United Kingdom jointly published guidance titled "Careful Adoption of Agentic AI Services," addressing security risks in agentic AI deployed across critical infrastructure and defense environments. The guidance identifies five categories of risk, covering privilege, design and configuration, behavior, structural issues, and accountability, and recommends incremental deployment with strong governance and continuous human oversight.
Taken together, these are not isolated stories. They are the same trend from different angles: governments are moving from watching AI development to actively shaping its release cadence, and every major lab is now operating with that reality as a planning input rather than a hypothetical.
The US-China Gap, Honestly Assessed
It is worth resisting both extremes of the current commentary. Kimi K3 did not erase the gap between Chinese and US AI labs. Claude Fable 5 and GPT-5.6 Sol remain ahead of every Chinese model tested on the hardest, most comprehensive benchmarks, and that lead is real and measured in double-digit percentage points on evaluations like SWE-bench Pro.
But the assumption that carried US policy and market pricing through most of 2025 and early 2026, that Chinese labs trailed by six to twelve months and that gap was stable or widening, no longer holds for the specific categories of coding and general agent capability. GLM-5.2 in June and Kimi K3 in July both demonstrated that a fully open, freely downloadable model from a Chinese lab can beat the previous generation of US flagship models on real benchmarks, at a fraction of the operating cost, within weeks of a comparable US release rather than months.
What has not changed: the very top of the market, the newest, most expensive, most restricted models like Fable 5, Mythos 5, and GPT-5.6 Sol, remains a US-led category. What has changed: the gap between that top tier and the best available open weight alternative has compressed from a wide margin to something measured in single digits on several important benchmarks. For any organization making a build-versus-buy decision on AI infrastructure this month, that compression is the single most consequential fact in this entire article.
What to Actually Use Right Now
Practical guidance based on what is confirmed and priced today.
For the best all-around model most teams should run daily: Claude Opus 5 at $5/$25 per million tokens. Near-Fable-5 intelligence, half the price, no mandatory data retention, and it is already the default on Claude Max.
For the single highest ceiling on hard reasoning, legal, and biology work: Claude Fable 5 at $10/$50, if the cost and 30-day retention terms work for your use case.
For agentic coding at the lowest cost per correct answer: GPT-5.6 Sol or Terra, especially after the July 30 price cuts. Sol's Agents' Last Exam lead over every Anthropic model is the most concrete evidence available right now.
For long-horizon agent reliability on a budget: Grok 4.5 at $2/$6, particularly for teams already inside the Cursor ecosystem.
For self-hosted, open weight deployment with genuine frontier-adjacent capability: Kimi K3 or GLM-5.2. Test both against your specific codebase before committing, since Moonshot's benchmark claims are self-reported and independent verification is still accumulating.
For everyday text and coding at Sonnet pricing: Claude Sonnet 5 at introductory pricing of $2/$10 through August 31, closing in on Opus 4.8's coding performance at a fraction of the cost.
What Comes Next
The Ai4 2026 conference runs August 4 through 6 in Las Vegas, and it is likely to be the next venue where a major lab makes news, given how consistently 2026 has produced a significant AI announcement every few weeks rather than every few months.
Beyond that, watch for independent third-party benchmark verification of Kimi K3's claims, which should firm up through August as evaluation platforms like Artificial Analysis and LMArena accumulate enough votes and test runs to assign it a stable composite score. Watch also for whether Mythos 5 sees any expansion beyond its current restricted list, which would be one of the clearest signals yet about how the export control situation is evolving.
The pattern this year has been consistent: whatever model leads the leaderboard in any given month rarely holds that position for more than a few weeks. Plan your infrastructure accordingly, and keep two or three models on hand rather than betting everything on a single provider.
Compare Every Model Side by Side
The table above is a starting point. For full benchmark scores, live pricing, context windows, and deployment notes across every model in this article:
Compare AI models on Renovate QR
The /tools directory updates as new benchmark data and independent evaluations arrive, including the ongoing Kimi K3 verification and any Mythos 5 access changes.
Published August 3, 2026. We update this article as new benchmark data, pricing changes, and access decisions are confirmed. For the full June and July context, see our archives on AI Models in July 2026 and AI Models in June 2026.
Frequently Asked Questions
What is the best AI model in August 2026?
Claude Opus 5, released July 24, 2026, currently leads the Artificial Analysis Intelligence Index, just ahead of Claude Fable 5 and GPT-5.6 Sol. Opus 5 is notable because it approaches Fable 5's intelligence at half the price, $5 per million input tokens versus Fable 5's $10. For pure peak reasoning on health, biology, and legal tasks, Fable 5 and the restricted Mythos 5 still lead. For coding specifically, GPT-5.6 Sol beats every Anthropic model on the Agents' Last Exam benchmark. There is no single best model. The right choice depends on whether you need peak intelligence, coding depth, or cost efficiency.
What is Kimi K3 and why did it move the stock market?
Kimi K3 is a 2.8 trillion parameter open weight AI model released by Chinese startup Moonshot AI on July 27, 2026. Moonshot published benchmark results showing K3 beating Claude Opus 4.8 and GPT-5.5 on several coding and general agent benchmarks, while running at roughly a third of Opus 4.8's cost. Because the model is freely downloadable and modifiable, it undercut assumptions that US labs held a comfortable lead over Chinese open weight AI. The Nasdaq dropped around 1% on the news, with investors selling shares of chipmakers including Nvidia and Intel on concerns about demand for premium US compute if cheaper open models close the capability gap.
How does Claude Opus 5 compare to Claude Fable 5?
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, exactly half of Fable 5's $10/$50 pricing. Anthropic says Opus 5 approaches Fable 5's intelligence on coding and knowledge work, and on the company's own Frontier-Bench v0.1 evaluation it more than doubles Claude Opus 4.8's score. It is now the default model on Claude Max and the strongest model available on Claude Pro. Fable 5 still leads on legal and health and biology evaluations, and GPT-5.6 Sol edges out Opus 5 on one agentic coding benchmark, so Opus 5 is not a strict upgrade on every axis, but it is the model most teams should run day to day given the price difference.
Is Grok now called SpaceXAI?
xAI increasingly operates under the SpaceXAI brand following SpaceX's acquisition of the AI coding tool Cursor in June 2026. Grok 4.5, released July 8, 2026, is the first model shipped under this restructured entity. It is built on a foundation Elon Musk has described as 1.5 trillion parameters and codenamed V9, though xAI has not officially confirmed the parameter count in its own documentation. The model was co-trained with Cursor on real developer session data and is priced at $2 per million input tokens and $6 per million output tokens, roughly 60% cheaper than Claude Opus 4.8.
What happened to Claude Fable 5 and Claude Mythos 5?
Both models were suspended by US export control order on June 12, 2026, just days after launch, following reports that a jailbreak technique allowed Fable 5 to generate functional cyberattack code. Because Anthropic could not verify user nationality in real time, both models went offline globally. The directive against Fable 5 was fully lifted June 30, and the model returned to worldwide availability July 1 with an enhanced safety classifier and a mandatory 30-day data retention policy, even on EU endpoints. Mythos 5 remains restricted to a list of more than 100 vetted US critical-infrastructure organizations and has not returned to general availability.
Why did OpenAI's GPT-5.6 launch behind a government-approved list?
Following a June 2026 executive order directing federal agencies to benchmark new frontier AI models before wide release, OpenAI's GPT-5.6 family, Sol, Terra, and Luna, launched June 26 in a preview restricted to roughly 20 government-approved partner organizations. OpenAI has publicly said it participated voluntarily and does not believe this kind of government access process should become the permanent default. GPT-5.6 reached general availability on July 9, 2026, and OpenAI cut prices on the smaller Terra and Luna tiers again on July 30.
Are open weight AI models catching up to closed frontier models?
Yes, meaningfully, though not uniformly. Zhipu's GLM-5.2, released under an MIT license in June 2026, scored 62.1% on SWE-bench Pro, and Moonshot's Kimi K3 pushed further in July, beating Claude Opus 4.8 and GPT-5.5 on several coding and agent benchmarks at a fraction of the cost. Both models still trail the very top of the market, Claude Fable 5 and GPT-5.6 Sol, by a meaningful margin on the hardest evaluations. The honest summary is that the gap between the best open weight model and the best closed model has narrowed from many months to single-digit percentage points on several key benchmarks, even as the absolute frontier keeps moving.

