Alibaba’s Contest: Fresh Contender is Bringing Significant Firepower
The artificial intelligence competitive race escalated sharply on Monday when Chinese e-commerce and cloud giant Alibaba challenged the United States’ tech superiority with the introduction of Qwen 3.8-Max, a 2.4 trillion-parameter model that stands shoulder to shoulder with the leading offerings from Anthropic and OpenAI. This latest model debuted mere days after the unveiling of DeepSeek V4 Flash 0731, which, as per independent evaluations by Artificial Analysis, operates within just one point of OpenAI’s economically accessible GPT-5.6 Luna while being 40 percent cheaper per task. Furthermore, at 284 billion parameters, it is compact enough to function on fairly standard enterprise servers and workstations.
Chinese model developers like Moonshot, Alibaba, and DeepSeek are actively competing against their American rivals on both cost and efficacy. This dual-pronged approach emerges as US developers such as Anthropic and OpenAI amplify concerns regarding the provenance and security of Chinese-manufactured AI systems.
US Apprehensions and the Strength of Chinese Contestants
In a recent blog entry, Anthropic’s CEO Dario Amodei asserted his opposition isn’t directed at open models generally, but specifically at those produced in China, those derived from proprietary designs, and those that fall short of stringent safety benchmarks. To put it another way, any model that presents a real challenge to Anthropic’s own offerings. The safety argument is particularly misleading, as the company’s chief fearmonger has gone out of his way to incite worry among US government officials. While proprietary models are manageable, open weights, once unleashed, cannot be retrieved.
Nonetheless, Amodei’s remarks and promises from prominent American and European tech companies don’t alter the reality that China is the sole substantial competitor in the open weights space. China is in a leading position here. Speaking on CNBC on Monday, Clément Delangue, CEO of Hugging Face, the world’s largest model repository, stated this. “They are clearly leading the charge on open models at the moment, and I wouldn’t be surprised if they begin to lead at the frontier by the end of this year or the next, given their pace of development,” he told the financial news network.
Alibaba Showcases Qwen 3.8-Max
The most capable open weights model from the US, Inkling, is just below a billion parameters, and still falls short compared to DeepSeek’s elite offerings. Alternatives like Moonshot’s Kimi K3 and Z.ai’s GLM 5.2 operate in an entirely separate league. In reality, the only viable options for enterprises seeking alternatives to proprietary models and their questionable security practices are Chinese models. DeepSeek, Alibaba, Moonshot, MiniMax, and Z.ai are keen not to miss out on this chance.
Unleashing the Kraken: Model Weights Go Live
Among the Chinese developers, Alibaba is arguably the closest match to its American adversaries. Its open-weight models are highly regarded, and their variety of sizes and flexible licensing have made them popular for fine-tuning application-specific systems. However, similar to Google and OpenAI, its premium models have been gated behind an API until this moment. Seizing the opportunity, Alibaba is now offering its most powerful model weights for download for the first time with the release of Qwen 3.8-Max.
The blog contains the familiar assortment of somewhat obscure bar charts illustrating how the model stacks up against the offerings of OpenAI and Anthropic, along with a series of demos that would have made Billy Mays proud. For those seeking details, we suggest checking out the launch blog here. One doesn’t need to examine too closely to grasp Alibaba’s message: whatever OpenAI and Anthropic can achieve, we can do equally well, if not superiorly, at a lower cost, and on your own hardware.
Hardware Challenges for Large Enterprises
That said, managing 2.4 trillion parameters is a significant undertaking for most enterprises, likely necessitating 48-64 Nvidia B200-class GPUs for customer-facing roles. For in-house tasks, 8-16 B300 or AMD MI355X GPUs would suffice. If this seems too extravagant, Alibaba has not overlooked its origins and will be introducing a 27-billion parameter iteration of the model alongside the Max version.
With benchmarks, model developers can generally find a set of tests that favorably depict their model. Unsurprisingly, Artificial Analysis’ own Intelligence leaderboard presents a slightly altered narrative compared to Alibaba’s, with Qwen 3.8-Max equaling Anthropic’s less robust, yet still impressive, Claude Sonnet 5. Beyond the marketing and demonstration pitches, Qwen’s blog is surprisingly lacking in specifics. What is known is that it is a multi-modal mixture of experts (MoE) model. This indicates that of the total 2.4 trillion parameters, only 95 billion are actually utilized to generate tokens for a specific request.
Pricing and the Marketplace Medley
We can also assume the model employs the same hybrid Transformer+Mamba architecture as earlier Qwen designs to maintain efficiency over large contexts. As for context capabilities, the model will accommodate context windows up to 1 million tokens, although it remains unclear if this is achieved through methods like rope scaling or not.
Currently, the model is accessible via Alibaba’s API service QwenCloud at a price of $2 per million input tokens and $6 per million output tokens. Cached tokens incur a sliding scale fee with $0.25 assessed per million implicit cached tokens, $0.17 per million explicit cache token reads, and $2.5 per million explicit cache tokens created. For comparison, Anthropic’s Claude Sonnet 5 costs $2/M for input tokens, $0.20/M for cache hits, and $10/M for output tokens. Notably, Sonnet 5 pricing is scheduled to rise by 50 percent beginning September 1. Meanwhile, OpenAI’s GPT 5.6 Luna, which ranks just below Qwen 3.8-Max and Sonnet 5 on the Artificial Analysis leaderboard, will charge $0.20/M for input tokens, $0.02/M for cached input, $0.25/M for cached writes, and $1.20/M for output tokens concerning short context lengths under 272,000 tokens, doubling for tasks exceeding that length.
The model weights are expected to be available on major model repositories, including Hugging Face, starting next week.
DeepSeek’s New Offering: An Alternate Strategy
DeepSeek undercuts OpenAI with a vibrant new V4 update. While Alibaba accompanies Moonshot.AI’s offensive on frontier models, DeepSeek has opted for a decidedly different approach with its latest open weights model: extracting maximum performance from a minimal number of weights. The outcome is DeepSeek V4-Flash-0731, a 284-billion parameter model capable of fitting into approximately 142 GB of GPU memory (at FP4). This enables enterprises to seamlessly run this model at scale on a single system. Surprisingly, despite its smaller size, the Flash model outperforms the 1.6 trillion-parameter DeepSeek V4 Pro by nearly 14 percent on the Intelligence leaderboard by Artificial Analysis.
Efficiency and Pricing: Engaging with Smaller Giants
That being said, the enhancements incorporated into DeepSeek V4 Flash will undoubtedly make their way into the Pro model in due time. Like Qwen 3.8-Max, DeepSeek V4 Flash offers competitive pricing against OpenAI and Anthropic – this time significantly so. Currently, for API access, DeepSeek is pricing at $0.14/M for input tokens, $0.0028/M for cached tokens, and $0.28/M for output tokens. However, in a market filled with reasoning models, API pricing doesn’t offer a complete perspective. A model might seem less expensive, but if it consumes double the tokens of another higher-priced model, it may not genuinely be more affordable.
Thus, it is crucial to assess how effectively the model addresses real-world challenges. Worryingly for US LLM developers, according to Artificial Analysis, DeepSeek V4-Flash not only holds a cheaper per-token price; it is also remarkably effective at its tasks. In contrast to OpenAI’s GPT 5.6 Luna, which ranks among the most efficient and cost-effective models in Sam Altman’s current portfolio, DeepSeek’s latest offering is a full 40 percent cheaper, costing only three cents to solve versus five cents. This discrepancy may seem minor, but for developers utilizing tens or hundreds of millions of tokens daily, this difference can quickly accumulate.
Key to Success: Speculative Decoding
A secret behind DeepSeek’s efficiency is the incorporation of DSpark speculative decoding directly within the model weights. We’ve examined speculative decoding previously, but essentially it entails employing a smaller draft model to forecast the outputs of a larger,