For the first time since the AI race began in earnest, Chinese firms are delivering consistent competition to the best the US and EU have to offer. National Technology News associate editor Rory Bathgate looks at the technical and financial appeal of models from firms such as DeepSeek, Moonshot AI, and Z.ai, and speaks to experts to discover the fundamental pros and cons of the Chinese AI ecosystem.
Enterprise AI still lives or dies on cost. As businesses move from proof-of-concept systems and chatbots to organisation-wide rollout and AI agents, they are scrutinising how much they pay for each AI token and what the value of each output is in real terms.
To date, Western AI labs have argued that their models justify their higher costs through superior performance on real-world enterprise tasks such as coding.
Chinese AI developers have increasingly undermined this model, by releasing free and cheap-to-run AI models that have edged closer to the best proprietary offerings from US labs. As these Chinese AI models begin to compete with those of Western counterparts on performance as well as price, enterprise leaders are being forced to consider how much they are willing to pay for their AI, and whether US offerings best suit their needs.
Since June, four major Chinese model releases have raised eyebrows in the AI community: Alibaba’s Qwen-3.8-Max, DeepSeek’s V4 Flash, Z.ai’s GLM-5.2, and Moonshot AI’s Kimi K3. All boast benchmark-topping performance, competing with frontier models by Anthropic and OpenAI, with lower price tags.
Kimi K3 is notable for offering performance at almost the same level as Claude Fable 5, just one month after Anthropic launched its flagship model. The rapid succession of these releases suggests the performance gap between Chinese and Western frontier models is narrowing.
In this piece, NTN looks at how we got here, why Chinese AI models may or may not be attractive to Western enterprises, and how much of an upset they could be for the US ecosystem.
The DeepSeek effect
The global launch of DeepSeek R1 in January 2025 was a landmark moment for the AI ecosystem. The model, which lagged only slightly behind the performance of Claude Sonnet 3.5 and GPT-o1, challenged the narrative that only American firms can achieve frontier results.
At the consumer level, DeepSeek’s rise was meteoric. By the end of January 2025, it had eclipsed ChatGPT to take the number one spot on Apple's App Store free downloads chart. But at the enterprise level, where every cent spent on AI outputs is weighed against the potential return, DeepSeek had more long-term effects.
The surprise for those in the enterprise AI space was twofold: first, that China could train such a powerful model using domestic hardware; and second, that it was willing to open-source the leap in performance.
Allan Dabre, technology compliance & AI lead at PwC, tells NTN that DeepSeek proved capable AI models could be produced without the data centre infrastructure available to US firms.
“For years the assumption was that frontier AI capability needed massive compute budgets, and that gave a handful of US labs, and the chipmakers behind them, a durable moat,” he explains. “DeepSeek R1 challenged that directly. It showed competitive reasoning performance at a fraction of the reported training cost, and that's what made it more than just another model release.”
Without access to the massive compute clusters used to train US AI models, the researchers behind DeepSeek were forced to optimise it at a software level. Breakthroughs included architecting the model for low resolution from the ground up by training on lower-precision data than US frontier models, allowing DeepSeek to compress the model to reduce its memory footprint without impacting its performance.
DeepSeek also rewrote the underlying code for its models to optimise them for Huawei chips, a key step to circumventing the bans on importing Nvidia hardware into China, and open sourcing five of its code repositories to allow AI researchers around the world to replicate its GPU-optimisation techniques.
Emilie Colker, chief executive at European data and AI consultancy ADC, says: “I like Azeem Azhar's reporting and when his team visited 14 Chinese labs in May, they concluded that American sanctions created America's toughest competition.”
“That means that chip constraints forced those labs to innovate on efficiency, and the breakthroughs now benefit everyone.”
These improvements are beginning to be felt at the enterprise level, as businesses gain access to free models that can compete with paid offerings. Kimi K3 is a notable example, as the model is on par with Anthropic’s Claude Fable 5 for tasks such as long-horizon software engineering, and DeepSeek V4 Flash offers firms a frontier reasoning model at an exceptionally low price point.
These models are open-weight rather than open-source, meaning that firms share their post-training parameters and allow users to download them freely, but do not provide any insight into the code and training data used to create the models.
The exact usage rights also vary by model. Kimi K3’s licence, for example, specifically notes that companies that run Model as a Service (Maas) offerings and earn more than $20 million in annual revenue have to enter into a separate commercial agreement with Moonshot.
The firm defines MaaS as "giving a third-party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data".
Open questions remain about the extent to which this has been achieved using purely Chinese hardware. The US Department of Commerce maintains that Chinese firms are smuggling Nvidia chips into China to use for AI training and in March, the US Department of Justice charged the co-founder of Super Micro Computer and two other individuals with allegedly conspiring to smuggle $2.5 billion worth of Nvidia chips into China via Taiwan.
Nvidia has responded by halving its list of approved Asian customers and stepping up its manual checks on data centres and customers.
Regardless of the extent to which chips have been smuggled in, the architectures and open-weights of Chinese models make clear that these are ruggedly optimised products designed to reach performance thresholds under marked compute restrictions.
Model distillation
One of the most prominent points of discussion in the “West vs China” dialogue around AI model development is model distillation.
This is a common machine learning technique in which one model is trained on the output of another, exemplar model. In recent years, the US government and certain AI developers have accused Chinese AI labs of illicitly using model distillation to steal technological advances.
In February 2026, Anthropic alleged that DeepSeek, Minimax, and Moonshot AI had carried out millions of interactions with its LLM Claude to extract its reasoning capabilities. It claimed developers at the labs had exchanged a combined 16 million messages with Claude via 24,000 fraudulent accounts. As a result, Dario Amodei, co-founder and CEO of Anthropic, has called for a sector-wide crackdown on distillation.
Such practices are explicitly banned in the terms of service for most AI platforms, particularly those accessed via API.
Some critics have argued that the developer’s protests are hypocritical, given the fact that every LLM currently on the market was trained on trillions of data points, in intense training runs that drew on data from across the public internet, as well as scans of physical media such as books.
James Hall, co-founder and tech director at Parallax, a UK tech consultancy specialising in AI, tells NTN that complaining about distillation is “a bit rich given the source of frontier models was copyrighted information in the first place”. At the same time, he says, it’s a problem for frontier AI developers as they are forced to spend a lot of money on training powerful models that are then copied for much lower sums.
Brock adds that if firms such as OpenAI or Anthropic were to pursue claims that distillation breached their licences, it could backfire.
“If a court was to decide it was a copyright infringement, then I suspect the whole training and fair use debate would be re-opened in the countries where it’s already settled,” she tells NTN.
The main risk here is the potential for IP and licensing battles down the line. If enterprises use AI models that are found to have been built through contractual breaches and illicitly accessed output data, they could find themselves in breach of compliance and supply chain rules.
Regardless of whether this practice is legitimate, or how extensively Chinese AI developers rely on it, the broader result is the same: open-weight AI models are becoming increasingly efficient, allowing enterprises to achieve similar results at a lower cost.The price war
Indeed, the intense focus Chinese AI developers have placed on efficiency has allowed them to offer inference of models at lower prices than Western rivals – often far lower. This is a core appeal of Chinese models for businesses, a calculation of performance over pedigree, as leaders are asked to consider the potential return on investment of slightly lower performance at a much lower price.
Arena AI tracks the price of using DeepSeek’s latest model, V4 Flash, at just three cents per task. The official price of using the model via DeepSeek API is $0.14 per million input tokens and $0.28 per million output tokens.
Compare this with Claude Opus 4.8, an Anthropic model that became a workhorse model for code generation at the enterprise level this Spring, and one starts to understand the appeal of Chinese AI models. The independent AI analysis firm Artificial Analysis has recorded Opus 4.8 scoring 84.6 per cent at the agentic coding benchmark Terminal Bench 2.1, versus 78.7 per cent by DeepSeek V4 Flash. But firms can expect to pay $5 per million input tokens and $25 per million output tokens for Anthropic’s offering.
Even for cached hits, a term used for outputs generated on inputs saved in server memory for repeat use, Opus 4.8 costs $0.50 per million output tokens versus $0.0014 for V4 Flash. Cache hits are most often used for repetitive tasks that draw on a large system prompt, including for generating code based on a repository or producing documents according to a schema, and DeepSeek is selling V4 Flash as the optimal model for meeting these needs.
This is just one example of how AI optimisation is now allowing Chinese labs to meaningfully undercut Western competition.
Jose Lejin PJ, IEEE senior member and principal engineer at Salesforce, tells NTN that the primary risk Chinese AI models pose to the US AI market is one of commoditisation.
“If the open-weights model is good enough for 90 per cent of the applications and runs on your computer for cents, then the entire pricing umbrella under which the US-based labs operate becomes worthless,” he says. “We've seen how the market has reacted to the first moments of DeepSeek.”
Dabre agrees that US firms are facing a real competitive threat. He tells NTN that Chinese models are increasingly competing with US models on quality and that this upsets their entire business model.
“Frontier labs have built enormous valuations on the assumption that capability commands a premium,” he says. “If capability becomes commoditized while the price collapses, that premium erodes, and it erodes for reasons that have nothing to do with whether the domestic models got worse.”
Amanda Brock, chief executive of the open tech industry body OpenUK, tells NTN that open models offer services at a price point that the proprietary developers cannot compete with in the public cloud.
“There is a clear advantage from a cost and access perspective to these open models, so it does feel like the battle on openness is coming to its natural conclusion and that the open models will win,” Brock says.
This has not stopped them from trying. OpenAI is already moving to stay more competitive with the Chinese competition. In July, it lowered the API costs for flagship model GPT-5.6 Luna by 80 per cent, to $0.2 per million input tokens and $1.20 per million output tokens.
Google has similarly pivoted to cost effectiveness, with its model Gemini 3.5 Flash Lite marketed as a cost-effective option for AI agents at $0.3 per million input tokens and $2.5 per million output tokens.
For now, Chinese AI models are still the cheaper option, and evidence suggests Western businesses are increasingly considering them out as an option. In July, Bloomberg reported that Chinese models accounted for 60 per cent of US enterprise AI usage on the model aggregation platform OpenRouter, which allows firms to access a wide range of AI models via an API key.
The San Francisco-based agentic AI startup Lindy AI also made headlines in June, when it announced that it had cut its inference costs by 90 per cent by switching from Claude Sonnet to DeepSeek V4 Flash for its underlying model.
“The model is the engine. It is also one of the biggest costs in the business,” wrote Bruno Škvorc, staff software engineer at Lindy AI, in a blog post explaining the change. Škvorc said that Lindy AI designed its pricing around the premise that models would get cheaper fast enough to lower operational costs without degrading the product.
“We did not want to spend the same money forever and merely give users more intelligence they did not always need,” Škvorc added. “You do not need God to write your emails.”
An inherent risk?
For all the potential benefits of Chinese AI models, enterprises looking to deploy them must take certain risks into account.
Many of the risks associated with using Chinese AI models come from using them via API and in the cloud. Thomas Cloud, founder and principal at fractional CTO and technology advisory firm Vertex CTO Advisory, tells NTN that data residency is the first concern for firms assessing whether to use Chinese AI.
“The 2021 PRC Data Security Law and PIPL create real exposure for any US company whose inference or fine-tuning data could be compelled by Chinese authorities,” he says. “CISOs and boards treat that as a live risk, not a theoretical one.”
One way to mitigate these risks is to run the model on premises. Using enterprise-owned hardware is a reliable method to maintain complete oversight of data it processes and ensure it complies with sovereign requirements, but it does come with additional up-front costs.
DeepSeek V4 Flash, for example, can be run on eight Nvidia H100 GPUs, an expensive but feasible setup to run in a colocation facility. This deployment ensures all data passed to the model remains truly private, essential if it is being used for sensitive tasks such as generating proprietary code.
A cheaper option is to rent servers via a GPU as a service offering from a neocloud, a private cloud infrastructure provider. This is a middle ground suitable for firms that have strict compliance requirements.
The key here is the approach to deployment and data management, rather than the models themselves. Indeed, Dabre says that business leaders simply cannot approach the task of ensuring their AI is secure as simply as trusting or distrusting AI models based on their country of origin.
“The actual issues are training data provenance, which is undisclosed for most open-weight models regardless of who builds them, jurisdictional data handling if a company uses a hosted API rather than self-hosting the open-weights, documented differences in content and refusal tuning shaped by different regulatory environments, and the standard supply chain diligence, checkpoint integrity, fine-tune vetting, that applies to any third-party model,” he tells NTN.
“Self-hosting the open-weights sidesteps most of the jurisdictional concern entirely. None of that adds up to a reason to distrust a model simply because of where it was built.”
At the same time, however, this eats into the cost argument for Chinese AI models, which are at their cheapest when accessed via API. Whichever approach to local infrastructure an enterprise takes, the short-to-medium term costs will be far higher.
James Hall, co-founder and tech director at Parallax, a UK tech consultancy specialising in AI, tells NTN that running a Chinese AI model locally can be a worse financial option for firms looking to use them for “spiky” tasks such as generating code from 9am to 5pm.
“If you bought your own GPUs, what are you doing with them when they’re sat idle?” he says. “The maths doesn’t always make it make sense from a total cost of ownership point of view.”
Cloud adds that the cost of accessing AI models from US labs via a hyperscaler includes legal guarantees, which must be factored into the higher costs of Western models.
“OpenAI, Anthropic and Microsoft all indemnify enterprise customers on training-data IP claims now,” he says. “Chinese labs don't. If a model output produces a copyright complaint, the buyer eats the liability.”
Cloud adds that for firms with access to government or defence data “US supply-chain and export-control rules make Chinese-origin models a non-starter regardless of how they benchmark”.
Sections of the US government and associated agencies including the Department of Commerce, the US Navy, and NASA have all banned employees from operating DeepSeek on their official devices.
In June, Reuters reported that the Trump administration was holding off from adding DeepSeek, as well as more than 100 other Chinese firms, to the Department of Commerce’s Entity List over fears of eroding diplomatic relations with China. If DeepSeek was added to the list in the future, US firms would be prohibited from using its API.
This is an issue of AI sovereignty, forcing enterprises to consider the potential pitfalls – and regulatory restrictions – associated with sending their data to be processed overseas. For example, an enterprise may determine it can only use Chinese AI models for certain tasks, or that even with security guarantees in place it cannot stake its reputation on non-US models.
Ultimately, these are strategic risks to be weighed in line with organisational goals, rather than inherent risks baked into any one model.
Indeed, while early criticism of models such as DeepSeek and Qwen zeroed in on the fact that the models could introduce vulnerabilities into generated code or were more susceptible to hacking such as AI prompt injection, there is little evidence for this.
In Veracode’s 2026 GenAI Code Security Report, researchers found that both US and Chinese AI models produced code containing vulnerabilities 44 per cent of the time on average. DeepSeek V4 Flash was found to pass just 51 per cent of Veracode’s security tests compared to GPT-5.5’s 68 per cent pass rate while Moonshot AI’s Kimi K2.6 scored 57 per cent, ahead of Google’s Gemini 3.1 Pro at 52 per cent.
Researchers at Cisco have similarly reported that vulnerabilities in open-weights models do not depend on national origin. In Death by a Thousand Prompts: Open Model Vulnerability Analysis, they compared the vulnerability of open-weight models to single-turn and multi-turn attacks revealing “pervasive vulnerabilities across all tested models” with multi-turn attacks yielding a two to 10 times increase in attack success.
This included analysis of models from Alibaba, DeepSeek, and Z.ai, as well as Google, Microsoft, OpenAI, and French sovereign AI firm Mistral.
A pillar of the community
In a recent blog post Anthropic’s Amodei wrote that although he is not against open-source AI models in principle, he thinks they “potentially present a higher risk than closed models, because it is very difficult to apply guardrails to them or monitor their usage”.
He is in direct disagreement with some of the world’s largest technology companies.
In July, some of the most prominent technology companies in the world put their names to an open letter penned by Microsoft, Nvidia, and Meta which called on the US government to hold back from applying “premature restrictions” on open-weight models.
Signatories included Amazon, AMD, Cisco, Cloudflare, Google, IBM, Intel, Hugging Face, Lenovo, OpenAI, Rackspace, Red Hat, and Uber.
Mark Zuckerberg, chief executive and co-founder of Meta, has argued that open-weight models are overall positive for cybersecurity. On 10 August, Meta published a blog post in which Zuckerberg argued that any government effort to block open-source AI would negatively impact the whole ecosystem.
“I do not believe restricting access to foreign open-source models is an effective solution,” Zuckerberg wrote. “Our goal should be for American open-source models to be the best globally. This requires removing the hurdles that make it harder for American open-source models to compete.”
Zuckerberg also proposed that widespread access to cybersecurity capabilities would counteract potential harms, stating that “widely deployed open-source systems have proven more secure because more people can identify vulnerabilities, harden the systems, and easily upgrade to the latest most secure versions”.
Indeed, a core argument of proponents for open-weight models is that they make the AI community safer, not less safe, by lowering the bar to entry for automated security systems and advanced security analysis.
This isn’t a purely theoretical problem, as the recent incident in which a rogue AI agent run by OpenAI hacked the infrastructure of US AI platform Hugging Face. When they detected the breach, the firm said its security teams tried to use proprietary AI models to contain and analyse the threat but were unable to as the models declined their requests.
In response Hugging Face carried out its forensic analysis using Z.ai’s GLM-5.2, running on its own infrastructure.
In a post-mortem blog post, Hugging Face said the attacker “was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried”.
None of this is to suggest that Chinese AI models are inherently more trustworthy or capable at security than Western models. The very measures that Hugging Face circumvented using Z.ai’s model are those that prevent LLMs from producing harmful or dangerous content such as code vulnerabilities or malware.
The open-source community’s argument that these tools can improve the overall safety of the software ecosystem is predicated on enterprises implementing their own guardrails and harnesses that turn open-weight models into useful tools with secure restrictions. Chinese AI models can play a role here but are not off-the-shelf solutions.
Putting the toothpaste back in the tube
The inescapable fact for AI developers, and for enterprises is this: Chinese AI models are effective and freely available.
Enterprises in the US are facing either a radical revision of the AI pricing and options to which they have become accustomed, or a potential ban on using Chinese AI models that will isolate them from the international AI ecosystem.
One route through is to reduce the focus on the models underpinning software and instead time perfecting the additional proprietary layers on top of models that make them truly useful for both enterprises and consumers.
In this way, Chinese AI competition could have a positive effect on the entire AI economy. Where US AI companies are already focusing their attention on the layers that go on top of the model, such as with Anthropic’s Claude Cowork productivity harness or OpenAI’s agentic harness Codex, they may still claim a unique advantage over Chinese labs.
Steve Povolny, VP for AI strategy and security research at Exabeam, tells NTN that Anthropic is still a frontrunner in this regard:
“Specifically, Anthropic’s annual revenue run rate has rocketed from about $9 billion at the end of 2025 to nearly $50 billion in mid-2026,” he says.
“They’ve continued to innovate and provide foundational experiences that differentiate the products, such as Plugins, Skills, MCP libraries, prebuilt agent frameworks and enticing purpose-driven models such as Fable. This innovation is what is driving growth and adoption, even while similar foundational features and capabilities exist in foreign counterparts.
Microsoft also describes its flagship AI offering Copilot as “model agnostic” and has repeatedly emphasised that it is the interaction between the foundation model, data, and enterprise layer that make Copilot useful to businesses, not the model alone.
“On one side of the spectrum, the US government will want to protect their IP and close off access to other countries thereby further accelerating the breaking up of Pax Americana,” says Edward Jansen, director of technology and innovation at ADC. “On the other side, AI might become a global public good, bringing it back to its academic roots.”
Colker calls on European firms to spend their time “building the capability to govern and deploy whichever models we choose, deliberately, than on picking a side”.
“Competition that halves the cost of intelligence for a hospital or a small manufacturer is hard to argue against. That would mean that doctors are able to spend more time with patients, and small companies can compete at a scale they couldn't reach before.”
Other model developers are seeking to ride the wave of open models themselves. Meta, once the frontrunner in open AI, has returned to the field with its new model Muse Glimmer. The 30-billion-parameter model is optimised to run on-device agentic workflows, with benchmarks competitive with the likes of Qwen 3.6 and Google’s Gemma 4.
The firm has also committed to releasing an open-weight version of Muse Spark 1.2, its frontier model that competes with the OpenAI and Anthropic mid-range offerings. This is a full circle moment for Meta which, under the leadership of Zuckerberg and chief AI officer Alexandr Wang, appointed in 2025, is seeking to share advanced models and lower the bar for AI-assisted cybersecurity work.
This is one of the first measures taken by US AI labs in direct response to the increasing market share for Chinese LLMs. While others lean on their USPs, Meta is trying to beat China at its own game and ensure the cheap models enterprises use are American.
Beyond the AI labs, businesses have an opportunity to make the most of a price war. If enterprises around the world can use the disruption Chinese AI models are causing for their own benefit, these releases could end up spurring a new wave of AI investment.
Indeed, the potential losses incurred by Western AI labs could translate directly into the gains of the wider economy, as enterprises achieve their AI productivity goals at a far lower cost. The security argument is also hard to overlook: firms named in the industry open letter back open-weights models as the future of AI-driven security. Notable companies such as Hugging Face have also demonstrated the value of Chinese AI models for analysing cyber incidents.
As US and Chinese labs compete, enterprises have more options for how they deploy AI. Enterprise teams can choose models depending on the workload, with the freedom to switch between trusted proprietary and open-weight models according to their performance, data governance, and security requirements.
With the right architecture built on top, none of this need be set in stone. Model agnostic systems allow enterprises to balance their budgets with their risk appetite in choosing whichever model suits them best, and the freedom to run systems without ever having to nail their colours to the mast of any one AI developer.
If US labs have revealed the limits of AI performance, Chinese labs have shown just how much AI performance can be squeezed from each penny. For enterprises willing to navigate the security, sovereignty and licencing risks, that creates a new opportunity to make AI meet their specific needs.







Recent Stories