Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch






On July 19, Alibaba’s Qwen team previewed Qwen3.8-Max-Preview, the next flagship in the Qwen family. The research team describes it as a 2.4 trillion-parameter model, ‘second only to Fable 5’ among the systems it benchmarked. The preview is live now. The benchmark table, model card, and license are not.

The July 19th 2026 announcement landed during the World AI Conference (WAIC) in Shanghai. It also arrived two days after Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-weight model. The timing is the story as much as the model.

This article separates what Alibaba confirmed from what it only claimed. Every performance figure below carries that caveat.

What Qwen announced

The Qwen account posted that Qwen3.8 is launching and going open-weight soon. It called the model ‘one of the most powerful available today, comparable to leading frontier systems.

The preview build is real and purchasable. Access runs through Alibaba’s Token Plan subscription. The preview is offered at 10% of standard pricing.

Qwen developer Shuai Bai added technical detail. He described Qwen3.8 as the team’s first multimodal model above 1 trillion parameters. It processes text, images, video, and documents. Alibaba team states the model should beat Qwen3.7-Max on coding, full-stack development, data analysis, and office workflows.

Interactive Explainer



Qwen3.8 Explainer

Marktechpost Interactive Explainer

Qwen3.8-Max-Preview: what Alibaba confirmed, what it only claimed

A 2.4T-parameter multimodal preview shipped before any benchmark, model card, or license. Explore the facts below.

Confirmed vs Claimed

Parameter scale

Serving-cost calculator

Verified 3.7-Max baseline

Preview is live and purchasable. Qwen3.8-Max-Preview is sold through Alibaba Token Plan, Qoder, and QoderWork at 10% of standard pricing.

Sparse Mixture-of-Experts, multimodal. Developer Shuai Bai says it is the team’s first multimodal model above 1T parameters, handling images, video, and documents.

OpenAI and Anthropic protocol compatibility. Existing coding agents can point at Qwen3.8 without rebuilding their harnesses.

2.4 trillion parameters. This is Alibaba’s own figure. No model card or specification confirms it.

“Second only to Fable 5.” No benchmark table has been published. The ranking rests on internal evaluation.

Open weights “soon.” No date, no license, no Hugging Face repository. Alibaba’s last two Max flagships shipped closed.

Active parameters per token. Undisclosed — the single number that decides real serving cost for a sparse MoE model.

Status as of July 19, 2026. Toggle a fact type using the tabs above.

Total parameters, publicly disclosed frontier models, July 2026. Qwen3.8’s 2.4T is Alibaba’s claim, not a verified figure.

Total count is not usable compute. A sparse MoE model activates only a fraction of these parameters per token.

If the full 2.4T model shipped open-weight, what would it take to load?



Nvidia H200 (141GB) cards

Estimate for weights alone; add roughly 20–30% for the KV cache and runtime overhead. Active-parameter serving would be far cheaper, but Alibaba has not published that number.

The verified predecessor. Every published Qwen3.8 “capability” number is really a Qwen3.7-Max number until Alibaba releases a benchmark table.

1M

Context window (tokens)

$3.75

Per 1M output tokens

Qwen3.7-Max, May 2026, closed weights. Alibaba’s historic edge has been price-to-performance, not topping a leaderboard.

Marktechpost

Data verified July 19, 2026 · Figures marked claimed are Alibaba’s own

The 2.4 trillion-parameter question

Total parameter count is not the same as usable compute. This distinction matters more than the main number. Qwen’s own history proves the point.

Qwen3-235B-A22B carries 235 billion total parameters but activates 22 billion per token. Qwen3-30B-A3B activates roughly 3 billion. Both are sparse MoE designs, and Qwen’s Max tier is too.

For Qwen3.8, the active-parameter count is the number nobody has. Without it, the 2.4T highlighted parameters says little about serving cost. As Startup Fortune calculated, a 2.4T model at 4-bit precision needs roughly 1.2 terabytes for weights alone. A single Nvidia H200 carries 141GB. Even eight cards leave awkward math.

That is why the practical question is not leaderboard position. It is whether Alibaba ships a smaller activated-parameter variant, a good quantized checkpoint, or a distilled sibling.

How developers reacted

On the July 19th 2026 Qwen3.8 preview’s community reaction split along predictable lines. Enthusiasm for another open-weight frontier model met fatigue over unverified benchmarks.

On Hacker News, the dominant view was that an open-weight race between Chinese labs benefits everyone. Commenters debated motive and read the timing as a direct response to Kimi K3. A small group questioned the ‘second only to Fable 5’ framing and called Qwen a benchmark specialist next to rivals.

On Reddit’s r/LocalLLaMA, the conversation was practical. The 2.4T serving math dominated, alongside hope for a smaller or distilled variant that a workstation could load. On X, the announcement trended and large accounts, including kimmonismus, amplified the open-weight line.

The dashboard below breaks that reaction down by platform.



Qwen3.8 Sentiment

Marktechpost Social Signal

How X, Reddit and Hacker News reacted to Qwen3.8-Max-Preview

A qualitative read of public discussion in the hours after the July 19 announcement. Filter by platform below.

Overall mood: cautiously positive

Enthusiasm for another open-weight frontier model, checked by fatigue over unverified benchmarks and doubts about who can actually run 2.4T parameters.

All platforms

X / Twitter

Reddit r/LocalLLaMA

Hacker News

Sentiment mix — all platforms

Positive  
Skeptical  
Neutral  

Method: Illustrative, not a statistical sample. Sentiment shares are an editorial read of the Qwen X thread and amplifiers, the r/LocalLLaMA discussion, and a 29-comment Hacker News thread (75 points), captured July 19, 2026. Quotes are paraphrased.

Marktechpost

Snapshot · July 19, 2026 · Sentiment will shift once benchmarks and weights land



Source link

  • Related Posts

    Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model

    A community developer, GnLOLot, has published a 1B model that runs fully on local hardware. The model is MiniCPM5-1B-Claude-Opus-Fable5-Thinking, with GGUF builds for llama.cpp-compatible runtimes. It needs no API key…

    Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

    A single 24GB card is the practical floor for serious local inference. It is enough for genuinely capable models, and small enough to sit on one GPU. An RTX 3090…

    Leave a Reply

    Your email address will not be published. Required fields are marked *