Chinese AI Models in 2026: DeepSeek V4, Kimi K3, GLM-5.2, Qwen 3.8, and the Full US-China Gap Analysis
The complete, fact-checked guide to Chinese AI models in 2026. DeepSeek V4, Kimi K3, GLM-5.2, and Qwen 3.7/3.8 compared on benchmarks, pricing, and licensing, plus the Stanford data behind the US-China AI performance gap.

TL;DR
Stanford's 2026 AI Index, published April 13, found the top US model led the top Chinese model by just 2.7% on Arena as of March 2026, down from a 17.5 to 31.6 point gap in 2023, despite the US outspending China roughly 23 to 1 on private AI investment.
GLM-5.2 from Z.ai led open weight coding benchmarks after its June 13 launch. Moonshot's Kimi K3, a 2.8 trillion parameter model released July 27, then took the title as the largest open weight model ever shipped, beating Claude Opus 4.8 and GPT-5.5 on several coding and agent benchmarks.
DeepSeek V4 Pro launched in April at $1.74 per million input tokens, then dropped to a permanent $0.435 input / $0.87 output rate on May 22. V4 Flash runs at $0.14/$0.28, making it one of the cheapest capable coding models available anywhere.
On July 31, DeepSeek pushed V4-Flash-0731, a re-post-trained update with the same 284B architecture, that outscores DeepSeek's own larger V4-Pro-Preview on nine published agent and coding benchmarks, including a jump from 7.3 to 54.4 on DeepSWE. Pricing did not change and the migration was silent.
Alibaba previewed a 2.4 trillion parameter Qwen3.8-Max on July 19 at WAIC Shanghai, claiming it is second only to Claude Fable 5. No benchmark table, license, active-parameter count, or pricing has been published. Treat the claim as unverified until independent testing arrives.
DeepSeek V4 uses a Hybrid Attention design combining Compressed Sparse Attention and Heavily Compressed Attention to cut long-context compute by roughly 73%. GLM-5.2 uses a mechanism called IndexShare that cuts per-token compute by 2.9x at full 1M context.
Every major Chinese lab covered here ships permissive licenses, MIT for DeepSeek and GLM-5.2, Apache 2.0 for open Qwen releases, which lets Western enterprises self-host on their own infrastructure and satisfy data sovereignty requirements without routing traffic through China-hosted endpoints.
Chinese AI Models in 2026: The Complete Guide
Eighteen months after DeepSeek's R1 briefly matched the top US model and rattled markets in early 2025, the story has moved well past a single disruptive release. Stanford's own 2026 AI Index put a number on what had been mostly anecdotal until this year: the performance gap between the best American and the best Chinese model measured just 2.7% as of March 2026, down from a 17.5 to 31.6 point spread in 2023. And that was before the summer's releases.
Since that report was published in April, the open weight coding race alone has changed leadership twice. GLM-5.2 took the crown in June. Kimi K3, a 2.8 trillion parameter model that briefly moved the Nasdaq on the day its weights went public, took it in July. Alibaba previewed a 2.4 trillion parameter Qwen 3.8 model in the same window, claiming, without evidence yet, that it trails only Anthropic's Claude Fable 5 among all frontier systems.
This guide covers what is actually confirmed about each of the major Chinese model families as of early August 2026, what pricing and licensing look like today, how the US-China gap is actually measured, and what remains an unverified vendor claim versus independently tested fact.
How we evaluate: every benchmark score, price, and specification below is attributed to its source: the model developer's own documentation, an independent evaluator such as Artificial Analysis or LMArena, or named third-party reporting. Where a claim is vendor-reported and has not been independently verified, that is stated explicitly rather than presented as settled fact.
The US-China Gap, By the Numbers
Before getting into individual models, it is worth grounding the broader narrative in the one number everyone cites and rarely sources correctly.
Stanford HAI's 2026 AI Index, released April 13, 2026, found that Anthropic's Claude Opus 4.6 led the Arena Leaderboard at 1,503 Elo, with ByteDance's Dola-Seed-2.0-Preview close behind at 1,464. Stanford's own framing of that 39-point Elo gap is 2.7%, a dramatic compression from the 17.5 to 31.6 percentage point gaps recorded across MMLU, MATH, and HumanEval back in May 2023. By the end of 2024, those same benchmark gaps had already fallen to between 0.3 and 3.7 points.
The investment side of the story is the part that makes the performance number startling. The US spent $285.9 billion in private AI investment in 2025, versus $12.4 billion in China, a gap of roughly 23 to 1. China, meanwhile, leads on AI publication volume (23.2% of global output), patent filings (69.7% of global grants), and industrial robot installations, which ran at close to nine times the US rate in the same period.
Three things are worth being precise about when you see this "2.7%" statistic repeated elsewhere. First, it is a snapshot as of March 2026, measured through Anthropic's Opus 4.6 and ByteDance's Dola-Seed model specifically, not a permanent or universal figure. Second, it predates GLM-5.2, Kimi K3, Claude Opus 5, and GPT-5.6, all of which shipped after the report's data cutoff. Third, it measures general Arena performance, not the narrower coding and agent benchmarks where Chinese open weight labs have made their most aggressive gains this summer.
What can be said with confidence: on the specific categories of agentic coding and open weight availability, the gap between the best Chinese model and the best US model has almost certainly narrowed further since March, given that Kimi K3 is now beating Claude Opus 4.8, a model more capable than the Opus 4.6 Stanford measured, on several coding benchmarks. On raw peak intelligence, measured by the newest closed models like Claude Fable 5 and GPT-5.6 Sol, the US retains a clear lead that neither GLM-5.2 nor Kimi K3 has closed.
Model Comparison at a Glance
| Model | Developer | Parameters (total / active) | Context window | Max output | License | Blended price (per 1M tokens, approx.) |
|---|---|---|---|---|---|---|
| Kimi K3 | Moonshot AI | 2.8T MoE (active not disclosed) | 1M | Not published | Open weight | Self-hosted or API, roughly a third of Opus 4.8 |
| DeepSeek V4 Pro | DeepSeek | 1.6T / 49B | 1M | 384K | MIT | $0.435 in / $0.87 out |
| DeepSeek V4 Flash | DeepSeek | 284B / 13B | 1M | 384K | MIT | $0.14 in / $0.28 out |
| GLM-5.2 | Z.ai (Zhipu AI) | 753B / 40B | 1M | 128K | MIT | $1.40 in / $4.40 out |
| Qwen 3.7-Max / Plus | Alibaba Cloud | Proprietary MoE | 1M | Not fully published | API only | Not fully published |
| Qwen3.6-35B-A3B | Alibaba Cloud | 35B / 3B | 262K | Standard | Apache 2.0 | Free to self-host |
| Qwen3.8-Max Preview | Alibaba Cloud | 2.4T (active not disclosed) | ~984K | 128K | Preview, license TBD | Promotional credits, no standard price yet |
All pricing figures are current as of early August 2026 and subject to change; always verify against the provider's own documentation before committing production budget.
Kimi K3: The Model That Moved the Nasdaq
Moonshot AI's Kimi K3 is the single most consequential release covered in this guide, not because it is unambiguously the best Chinese model, but because of how directly it challenged assumptions that had held through most of the year.
Financial Times reporting ahead of the launch put the parameter count between 2 and 3 trillion. Moonshot confirmed 2.8 trillion on release, making K3 the largest open weight model any Chinese lab has shipped. The company announced the model July 17 and released downloadable weights July 27, allowing any developer to run it on their own infrastructure at no licensing cost.
Moonshot's own published benchmarks show K3 beating Claude Opus 4.8 and GPT-5.5 on several coding and general agent evaluations, while costing roughly a third of what Opus 4.8 charges per token. Moonshot was notably careful not to overstate the result: the company's own materials state that K3 still trails Claude Fable 5 and GPT-5.6 Sol on overall performance. That kind of self-imposed honesty is unusual in AI marketing and is part of why the release was taken seriously rather than dismissed as hype.
The market reaction was immediate. The Nasdaq dropped roughly 1% on the news, with investors selling shares in chipmakers including Nvidia and Intel. The underlying logic, whether or not it fully holds up, was that a freely downloadable model this capable undercuts the assumption that premium US compute and premium-priced closed models carry a durable moat. Technology analyst Patrick Moorhead pushed back on some of the reaction as overblown and politically charged, noting to CNBC that debate in Washington over whether US firms should even engage with Chinese open models is somewhat beside the point, since "the Chinese seem to be doing fine with their models" regardless. Moonshot is reportedly raising a new funding round that would value the company at $31.5 billion, up from a $20 billion valuation after a $2 billion raise in May.
For enterprises evaluating K3, the practical caveat is compute. Running a 2.8 trillion parameter Mixture-of-Experts model, even with sparse activation, requires meaningfully more infrastructure than a 70 billion parameter dense model. This is not a laptop model. Budget for serious multi-GPU cluster infrastructure or a managed inference provider before committing production workloads.
DeepSeek V4: Cheapest Frontier-Adjacent Coding Model on the Market
DeepSeek's V4 series remains the strongest pure economics story in the entire Chinese AI landscape, and the pricing has moved twice since launch, always downward.
DeepSeek V4 Pro launched April 24, 2026, the same day OpenAI shipped GPT-5.5, at 1.6 trillion total parameters with 49 billion active per token. V4 Flash, the lighter sibling, runs 284 billion total with 13 billion active. Both share a 1 million token context window and support up to 384,000 tokens of output per request, well above GLM-5.2's 128K ceiling.
The architecture is the reason the pricing works. DeepSeek V4 introduces a Hybrid Attention design combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), which lets the model hold a 1 million token context while consuming roughly 27% of the compute, a 73% reduction, that a conventional multi-head attention model would need at the same length. Paired with DeepSeek's DSpark inference framework, which the company says improved per-user generation speed by 60 to 85% on V4-Flash and 57 to 78% on V4-Pro over the prior baseline, the result is a model that is both fast and cheap to serve at scale.
On pricing specifically: V4 Pro launched around $1.74 per million input tokens, then DeepSeek announced a permanent cut to $0.435 input and $0.87 output on May 22, 2026, with cached input priced at just $0.0036 per million tokens. V4 Flash sits even lower at $0.14 input and $0.28 output. When DeepSeek shipped its official, fully stable version in mid-July, it introduced peak-hour pricing that doubles the standard rate during high-demand windows, so budget planning should account for time-of-day variability rather than treating the headline rate as fixed. One week later, on July 24, DeepSeek retired its legacy deepseek-chat and deepseek-reasoner API aliases entirely, so any integration still calling those old model names is already broken and needs to move to deepseek-v4-flash or deepseek-v4-pro.
On raw capability, V4 Pro posts a 93.5% score on LiveCodeBench and a Codeforces rating of 3,206, and it leads the field on HMMT February 2026 competition math and Tool-Decathlon, a benchmark for multi-tool agentic execution. Both models are released under an MIT license, and DeepSeek separately open-sourced DeepSpec, its full speculative decoding training stack, also under MIT, making the underlying technology usable by teams building on Qwen3 or Gemma as well.
What Happened With V4-Flash-0731
On July 31, 2026, DeepSeek pushed an update that is worth explaining carefully, because the headline result sounds implausible until you understand the mechanism behind it.
DeepSeek moved its V4-Flash API out of preview status and into official public beta under the build name DeepSeek-V4-Flash-0731. The architecture did not change. The parameter count did not change. It is still the same 284 billion total, 13 billion active Mixture-of-Experts model that launched in April. What changed is the post-training: DeepSeek re-ran the fine-tuning and reinforcement learning phase on top of the identical base model, without touching pretraining or scale in any way. DeepSeek's own API changelog is explicit about this: "DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained."
The result of that re-post-training alone is a genuinely striking jump. On Terminal-Bench 2.1, DeepSeek reports the 0731 build scoring 82.7, up from the April Flash preview's score in the high 50s to low 60s depending on the source cited, and notably higher than V4-Pro-Preview's own score of 72.1 on the same benchmark. On DeepSWE, the jump is even larger: from 7.3 in the preview to 54.4 in the 0731 build. In other words, DeepSeek retrained its smaller, cheaper Flash model until it started beating its own larger, more expensive Pro flagship on nine separate agent and coding benchmarks the company published, without changing the underlying architecture at all.
The update also adds native support for OpenAI's Responses API format and full compatibility with Codex, making it easier to drop into agentic coding tools built around that interface. Pricing did not change: it remains $0.14 input and $0.28 output per million tokens, the same rate that was already in place before the update, despite landing one day after OpenAI cut GPT-5.6 Luna's price to $0.20/$1.20, a coincidence of timing rather than a reactive price cut. For developers, the migration cost is zero: the API call is unchanged, so any application already calling deepseek-v4-flash received the upgraded model automatically and silently.
Two caveats matter here. First, this update only touches the Flash API. The V4-Pro API and the consumer app and web versions were not updated and remain on their earlier builds. Second, DeepSeek ran its published benchmarks using an internal evaluation harness, called DeepSeek Harness, that has not itself been released publicly, run at maximum reasoning effort with top_p 0.95 and temperature 1.0. As of the July 31 announcement, no independent lab had reproduced these figures. Treat the specific benchmark numbers as a strong vendor claim worth testing on your own workload, not yet as an independently confirmed result.
The broader significance goes beyond one model update. It is evidence that a meaningful share of the capability gains happening across Chinese open weight labs right now are coming from better post-training and reinforcement learning, not from bigger base models or new architectures. That has direct implications for cost: post-training a 284 billion parameter model is dramatically cheaper than pretraining a new one from scratch, which is part of why DeepSeek can ship this kind of improvement without raising the price at all.
GLM-5.2: Z.ai's Long-Horizon Coding Specialist
Zhipu AI, now operating under the Z.ai brand, released GLM-5.2 on June 13, 2026, first through a Coding Plan subscription and then via standalone API on June 16. For several weeks that summer, before Kimi K3 arrived, GLM-5.2 was the strongest publicly accessible open weight coding model on the market.
The architecture is a 753 billion parameter Mixture-of-Experts model with 40 billion active parameters per token, a 1 million token context window, and MIT licensing with weights published on Hugging Face. The standout technical feature is IndexShare, a sparse attention mechanism that reuses the same attention indexer across every four sparse attention layers, which Z.ai's own model card says cuts per-token compute by 2.9 times at full 1 million token context.
On Artificial Analysis's Intelligence Index, GLM-5.2 running at max effort scores 51 against DeepSeek V4 Pro's 44, and it holds a 6.7 point lead over DeepSeek V4 Pro on SWE-bench Pro specifically, along with larger advantages on long-horizon coding evaluations: plus 45.4 on FrontierSWE, plus 38.2 on DeepSWE, and plus 17.0 on Terminal-Bench 2.1. Jefferies technology strategist Christopher Wood described GLM-5.2 as "almost equal to Anthropic as a competitor for the corporate market" while costing roughly a quarter as much per token, a striking claim from an equity research desk rather than an AI marketing team.
Pricing sits at $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's standard API, or a flat $18 per month Coding Plan subscription for teams that prefer predictable billing over metered usage. The model is designed to drop into existing agentic coding tools with minimal friction, including Claude Code, OpenCode, and Z.ai's own ZCode, through an Anthropic-compatible endpoint.
The honest trade-off against DeepSeek V4: GLM-5.2 wins clearly on long-horizon, repository-scale coding tasks, but DeepSeek V4 Pro is roughly five times cheaper per output token and supports three times the maximum output length. Route measurable, high-volume work to DeepSeek, and reserve GLM-5.2 for the harder tasks that need the extra reasoning depth.
Qwen 3.7 and the Unverified Qwen 3.8 Preview
Alibaba's Qwen family has followed the fastest release cadence of any lab covered in this guide, and it is also the family where the gap between announcement and independent verification is currently widest.
Qwen3.7-Max and Qwen3.7-Plus reached stable release on May 18, 2026, following the open weight Qwen3.6-35B-A3B (April 15) and Qwen3.6-27B (April 22) releases earlier in the spring. The 3.6 generation open models remain the practical self-hosting option in the Qwen family today: Qwen3.6-35B-A3B activates just 3 billion of its 35 billion total parameters per token under an Apache 2.0 license, making it genuinely deployable on a single high-end consumer GPU or a modest cloud instance, a meaningfully different proposition from the trillion-parameter flagships covered elsewhere in this guide.
Then, on July 19, 2026, two days after Moonshot's Kimi K3 announcement, Alibaba's Qwen team previewed Qwen3.8-Max at the World AI Conference in Shanghai. The claims are bold: a 2.4 trillion parameter multimodal Mixture-of-Experts model that Alibaba describes as "one of the most powerful models available today, comparable to leading frontier AI models, second only to Fable 5." Qwen developer Shuai Bai noted it is the team's first multimodal model above one trillion parameters, capable of processing text, images, and reportedly video and documents.
Here is what is genuinely confirmed as of early August: the model exists, is accessible as Qwen3.8-Max-Preview through Alibaba's Token Plan, Qoder, and QoderWork platforms, and carries 2.4 trillion total parameters per Alibaba's own announcement. Here is what is not yet confirmed: any benchmark table, the active parameter count, a license, or standard per-token API pricing. Alibaba says open weights are "coming soon" without naming a date. Given that every previous Max-tier Qwen release has stayed API-only while a separate line handled open weights, whether Qwen 3.8 actually follows through on an open release is a real open question, not a formality.
The timing of the preview, landing exactly two days after Kimi K3's open weight launch, reads as a deliberate competitive response, an attempt to reclaim news cycle attention before independent benchmarks could settle the question of which lab actually leads. Treat the "second only to Fable 5" framing as a marketing claim until Artificial Analysis, LMArena, or another independent evaluator publishes results. Unconfirmed rumors, sourced to leaked roadmap chatter rather than any Alibaba statement, point to a full Qwen 3.8 release landing in August and an entirely separate Qwen 4 arriving around September, reportedly pushing into 3D coding and design applications. None of that should be treated as scheduled until Alibaba confirms it directly.
Visual and Multimodal Generation
Chinese labs have also built a genuine specialization in visual generation, particularly around workflows that depend on East Asian typography, cultural context, and tight integration between image and video generation pipelines. ByteDance's Seedream line and Zhipu's CogView models are the most cited examples, both of which are frequently reported to handle complex Hanzi character rendering, traditional clothing and architectural detail, and multi-step image-to-video handoffs more reliably than Western general-purpose image generators, which remain optimized primarily for Latin-script typography and Western visual conventions.
Precise, independently verified accuracy figures comparing Western and Chinese image models on typography and photorealism are not consistently published across evaluators, so specific percentage claims in either direction should be treated with caution until a named benchmark, such as an Arena-style blind comparison, backs them up. The safer summary: for APAC-market creative work, multilingual packaging design, and workflows requiring native video pipeline integration, Chinese visual models are a strong default worth testing directly against your specific use case before committing.
Deploying Chinese Models Safely: The Data Sovereignty Question
The single most important operational decision for any Western enterprise evaluating these models is how you access them, not just which one you pick.
Calling a Chinese lab's hosted API directly means your request data transits to infrastructure subject to Chinese data regulation. For teams handling regulated or proprietary data under SOC 2, HIPAA, or GDPR obligations, that path generally carries more compliance risk and should go through legal review before any production commitment.
The self-hosting path is the more common enterprise pattern precisely because every major model in this guide, DeepSeek V4, GLM-5.2, and the open Qwen releases, ships under a permissive license: MIT for DeepSeek and GLM-5.2, Apache 2.0 for Qwen's open weight line. Downloading the weights and running them on your own AWS, Azure, GCP, or on-premise infrastructure keeps data entirely inside your existing security boundary, with no dependency on any China-hosted endpoint. This is the standard recommended path for regulated industries that still want the cost and capability benefits these models offer.
One practical note specific to GLM-5.2: Z.ai has faced scrutiny from the US Bureau of Industry and Security in past coverage of Chinese AI labs, which is worth factoring into procurement conversations even though the model itself carries an open MIT license that permits self-hosted deployment regardless of the vendor's regulatory status.
Which Model Should You Actually Use
Practical routing guidance based on what is confirmed and priced today.
For the lowest cost per token at genuinely competitive coding quality: DeepSeek V4 Flash or Pro, depending on task complexity. At $0.14 to $0.87 per million tokens, nothing else in this guide comes close on pure economics.
For the strongest long-horizon, repository-scale coding agent: GLM-5.2, especially if your workflow already runs through Claude Code, OpenCode, or another Anthropic-compatible tool.
For the largest available open weight model and cutting-edge agent swarm orchestration: Kimi K3, if you have the infrastructure to run a 2.8 trillion parameter model and your workload benefits from its coordination capabilities.
For lightweight, single-GPU self-hosted deployment: Qwen3.6-35B-A3B under Apache 2.0, the most practical genuinely local option covered here.
For anything involving Qwen 3.8: wait. Test the preview if you are curious, but do not build production infrastructure around unverified performance claims until Alibaba publishes a real benchmark table and license.
Compare These Models Side by Side
The table earlier in this guide is a snapshot. For live pricing, updated benchmark scores, and deployment notes as new independent evaluations of Kimi K3 and Qwen 3.8 arrive:
Compare AI models on Renovate QR
The /tools directory is updated as new benchmark data lands, including ongoing independent verification of the claims covered in this guide.
Last updated August 3, 2026. Pricing, benchmark scores, and licensing terms for fast-moving open weight models change frequently. Always verify current terms directly with the provider before committing production budget or infrastructure.
Frequently Asked Questions
How wide is the performance gap between US and Chinese AI models in 2026?
According to Stanford HAI's 2026 AI Index, published April 13, 2026, the gap between the top US model and the top Chinese model on the Arena Leaderboard had narrowed to 2.7% as of March 2026, specifically Claude Opus 4.6 at 1,503 Elo versus ByteDance's Dola-Seed-2.0-Preview at 1,464. That is down from a 17.5 to 31.6 percentage point spread across major benchmarks in 2023. The US still leads on private AI investment by roughly 23 to 1, while China leads on AI publication volume, patent filings, and industrial robot deployment. Developments since March, including GLM-5.2 and Kimi K3, suggest the gap on open weight coding and agent benchmarks has likely narrowed further, though no updated Stanford figure has been published to confirm that.
What is Kimi K3 and why did it move markets?
Kimi K3 is a 2.8 trillion parameter open weight AI model from Chinese startup Moonshot AI, announced July 17, 2026 with weights released for public download on July 27. It is the largest open weight model shipped by any Chinese lab to date. Moonshot published benchmarks showing K3 beating Claude Opus 4.8 and GPT-5.5 on several coding and general agent tasks, while running at roughly a third of Opus 4.8's cost, though it still trails Claude Fable 5 and GPT-5.6 Sol overall. Because the model is freely downloadable, the release triggered a roughly 1% Nasdaq dip with selling in chipmakers including Nvidia and Intel on concerns that cheap, capable open models could reduce demand for premium US compute.
How much does DeepSeek V4 cost and how has pricing changed?
DeepSeek V4 Pro launched April 24, 2026 at roughly $1.74 per million input tokens. On May 22, 2026, DeepSeek announced a permanent price cut to $0.435 input and $0.87 output per million tokens, with a cached input rate of $0.0036. DeepSeek V4 Flash, the smaller and faster variant, runs at $0.14 input and $0.28 output. When DeepSeek shipped its official, non-preview version in mid-July, it introduced peak-hour pricing that doubles the standard rate during high-demand periods, so effective cost varies by time of day. Both models carry a 1 million token context window and are released under an MIT license.
Is Qwen 3.8 better than Claude Fable 5?
That claim comes from Alibaba itself, not from independent testing. Alibaba's Qwen team previewed Qwen3.8-Max on July 19, 2026 at the World AI Conference in Shanghai, describing the 2.4 trillion parameter model as second only to Claude Fable 5 among the systems it evaluated. No benchmark table, model card, active-parameter count, license, or standard API pricing has been published as of early August 2026. The preview is accessible through Alibaba's Token Plan, Qoder, and QoderWork platforms. Until Artificial Analysis, LMArena, or another independent evaluator publishes results, this ranking should be treated as an unverified vendor claim rather than a confirmed fact.
What is GLM-5.2 and how does it compare to DeepSeek V4?
GLM-5.2 is Z.ai's, formerly Zhipu AI, open weight reasoning model, launched June 13, 2026 through a Coding Plan subscription and June 16 via standalone API. It is a 753 billion parameter Mixture-of-Experts model with 40 billion active parameters, a 1 million token context window, and MIT licensing. On Artificial Analysis's Intelligence Index, GLM-5.2 at max effort scores 51 versus DeepSeek V4 Pro's 44, and it leads DeepSeek V4 Pro on SWE-bench Pro by 6.7 points. DeepSeek V4 Pro wins decisively on price, at roughly a fifth of GLM-5.2's per-token cost, and offers a larger 384K maximum output versus GLM-5.2's 128K. The practical choice depends on workload: GLM-5.2 for long-horizon repository-scale coding agents, DeepSeek V4 for high-throughput, cost-bound production API traffic.
Is it safe and legal for Western enterprises to use Chinese AI models?
Yes, provided deployment is handled correctly. DeepSeek, Z.ai, and Alibaba's open Qwen releases ship under permissive licenses, MIT for DeepSeek and GLM-5.2, Apache 2.0 for open Qwen models, which means Western enterprises can download the weights and self-host them on their own private cloud or on-premise infrastructure. Self-hosting keeps data inside your own security boundary and avoids routing traffic through China-hosted API endpoints, which is generally the safer path for teams operating under SOC 2, HIPAA, or GDPR obligations. Calling a Chinese lab's hosted API directly, without regional isolation, carries more compliance risk and should be evaluated case by case with your legal team.
What changed in DeepSeek-V4-Flash-0731?
On July 31, 2026, DeepSeek moved its V4-Flash API out of preview into official public beta under the build name DeepSeek-V4-Flash-0731. The base architecture and parameter count are unchanged, still 284 billion total parameters with 13 billion active. The update comes entirely from re-post-training, redoing the fine-tuning and reinforcement learning phase without touching pretraining. The result is a smaller, cheaper model that DeepSeek says now outscores its own larger V4-Pro-Preview on nine published agent and coding benchmarks, including a jump from 7.3 to 54.4 on DeepSWE. Pricing stayed at $0.14 input and $0.28 output per million tokens, and existing integrations calling deepseek-v4-flash received the upgrade automatically with no code changes. The benchmarks were run on DeepSeek's own unreleased evaluation harness and had not been independently reproduced as of the announcement.
What makes DeepSeek V4's architecture different from earlier models?
DeepSeek V4 introduces a Hybrid Attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). This lets the model maintain a 1 million token context window while using roughly 27% of the compute, a 73% reduction, that a standard multi-head attention architecture would require at the same context length. Combined with DeepSeek's DSpark inference framework, which improved per-user generation speed by 60 to 85% on V4-Flash and 57 to 78% on V4-Pro over the prior baseline, this is the core reason DeepSeek can offer 1M-token context at a fraction of the price of comparable Western models.

