Tencent has quietly put one of the biggest open-weight language models of 2026 into the hands of any developer who wants it, and the numbers behind the release are only now getting widespread attention outside China. The model, Hy4 preview, carries 770 billion total parameters but activates just 49 billion of them for any given request. Built by the Tencent Hunyuan team and shipped under an Apache 2.0 license, it landed on Hugging Face, ModelScope, GitCode, and CNB on August 28, 2026, then spent the following weeks working its way through Western coverage as analysts dug into its architecture and benchmark claims.

What makes Hy4 preview worth a second look this week is not just its size. It is Tencent’s clearest statement yet that China’s open-weight labs intend to compete directly with GLM-5.3, Kimi K3, DeepSeek, and Mistral on engineering and agentic work, not just chat quality. Tencent ran its own blind evaluation against two of those rivals and published the results alongside the weights, an unusually transparent move for a company that has historically kept Hunyuan’s internal benchmarks close to the chest.

What Is Hy4 Preview? The Specs at a Glance

Hy4 preview is a Mixture-of-Experts (MoE) model, meaning it bundles many specialized sub-networks and only switches on a handful of them per token instead of running the whole model every time. According to the official Tencent-Hunyuan GitHub repository, Hy4 preview carries 770 billion total parameters spread across 78 transformer layers, with 49 billion parameters active for each token processed. The first layer runs a standard dense feed-forward network, while the remaining 77 layers each contain 256 routed experts plus one shared expert that is always active. Every token is routed to the top eight experts plus that shared expert.

Tencent’s own launch page frames the release directly. “Today we’re releasing Hy4 preview, our most capable model to date,” the Tencent Hunyuan team wrote, adding that “it’s a 770B-parameter model with 49B active parameters and a 1M-token context window, and it makes the biggest gains where the work is hardest: long-horizon software engineering, document-heavy office work, and scientific research.” (hy.tencent.ai)

Beyond the headline parameter count, the model card details a hidden size of 6,144, a vocabulary of 120,832 tokens, and 64 attention heads, according to specs compiled by MindStudio’s analysis of the release (mindstudio.ai). A separate Multi-Token Prediction module, roughly 10 billion parameters with 0.7 billion active, rides alongside the main backbone to speed up generation through speculative decoding. Tencent shipped two checkpoints: the full-precision Hy4-preview and an FP8-quantized Hy4-preview-FP8 built for cheaper, faster inference. Both checkpoints are listed on the model’s Hugging Face model card.

SpecificationHy4 preview
Total parameters770 billion
Active parameters per token49 billion
Transformer layers78 (1 dense + 77 MoE)
Routed experts per MoE layer256, top-8 selected
Shared experts per layer1 (always active)
Context window1,000,000 tokens
Hidden size6,144
Vocabulary size120,832 tokens
Attention heads64
LicenseApache 2.0
Release dateAugust 28, 2026
DistributionHugging Face, ModelScope, GitCode, CNB

Inside the Architecture: Gated DSA Attention and the iHC Residual Trick

Two architectural choices separate Hy4 preview from a typical dense transformer, and both target the same problem: making a 1-million-token context window actually affordable to run. The first is what Tencent calls Gated DSA, short for Gated DeepSeek Sparse Attention. It extends DeepSeek’s sparse-attention approach, which trims the usual quadratic cost of attention by having the model pick a sparse subset of relevant tokens instead of scanning everything, and adds a gating mechanism on top. Hy4 preview pairs that with IndexCache, a technique that reuses sparse attention indices across layers instead of recalculating them every time, according to the model card details compiled by MindStudio.

The specific numbers behind Gated DSA are unusually granular for a model card: a query compression dimension of 2,048, a key-value compression dimension of 512, 32 indexer heads at a head dimension of 128, and an indexer top-k of 2,048, meaning the sparse attention mechanism selects up to 2,048 tokens to focus on per step rather than the full context. Tencent credits the design lineage explicitly to both DeepSeek’s sparse-attention research and GLM’s related attention work, a rare acknowledgment of cross-pollination between rival Chinese labs.

Identity Hyper-Connections

The second change sits in the residual pathway, the mechanism that carries information between layers. Standard transformers pass that information through a single stream that every layer reads from and writes to. Hy4 preview instead uses what Tencent calls Identity Hyper-Connections (iHC), expanding the residual stream into four parallel paths. The practical effect, per MindStudio’s breakdown, is more distinct channels for information to persist or recombine as it moves through a 78-layer stack, rather than everything getting compressed through one corridor. Tencent is tackling two bottlenecks at once here: attention cost over long sequences, and information flow across a model deep enough to need it.

The Benchmark Numbers: Hy4 vs. GLM-5.3 and Kimi K3

Tencent did not lean on a standard public leaderboard to make its case. Instead, according to MindStudio’s reporting on the release, the company recruited 163 internal experts, including software engineers, game developers, finance analysts, and security specialists, and had them rate model outputs across 203 real engineering tasks without knowing which model produced which answer. Ground News’s aggregation of the coverage confirms the same headline figures circulating across outlets this week. The result: Hy4 preview scored an average of 2.99, edging out GLM-5.3 at 2.92 and Kimi K3 at 2.94.

Win, Tie, and Loss Rates

Broken into head-to-head win rates, the margins are tight. Hy4 preview beat GLM-5.3 in 46.8% of comparisons, tied in 12.8%, and lost 40.4% of the time. Against Kimi K3, it won 51.2%, tied 7.9%, and lost 40.9%. Those are not blowout numbers. A model that loses roughly two out of every five head-to-head comparisons against its closest rivals is ahead on average, not dominant, and Tencent has not published an equivalent blind comparison against DeepSeek V4.1-Flash, Mistral Large 4, or Qwen3.8, so any claim of an outright win across the open-weight field stays unverified.

Open-weight flagshipTotal / active paramsContext windowLicense
Tencent Hy4 preview770B / 49B1,000,000 tokensApache 2.0
Mistral Large 4 (Le Chonk)~675B official / 41BNot independently confirmedOpen weight
GLM-5.3 (Z.ai)Not disclosed in this comparisonNot disclosed in this comparisonOpen weight
Kimi K3 (Moonshot)Not disclosed in this comparisonNot disclosed in this comparisonOpen weight

The parameter figures for Mistral Large 4 in that table reflect an ongoing dispute of their own. Mistral’s officially stated 675 billion parameters clashed with a third-party claim of 1.05 trillion, and the gap was never fully resolved by the company. That messiness is a useful reminder that parameter counts across the current wave of Chinese and European open-weight releases are not always apples to apples, and readers comparing Hy4 preview to its rivals should treat every number here as self-reported unless a named independent lab has verified it.

Where Hy4 Already Ships: CodeBuddy, WorkBuddy, Yuanbao, and ima

Tencent did not treat Hy4 preview as a pure research drop. The model went live simultaneously inside four of the company’s own products: CodeBuddy (its coding-agent tool), WorkBuddy (a workplace productivity assistant), Yuanbao (Tencent’s general consumer AI assistant), and ima (its workspace application). Tencent’s own statement put it plainly: “Hy4 preview is open-sourced and available today on Tencent Cloud and OpenRouter, and you can use it now in Tencent products such as CodeBuddy / WorkBuddy, Yuanbao, and ima.” (hy.tencent.ai)

That four-product rollout tells you more about Tencent’s strategy than the benchmark scores do. Rather than treating Hy4 as a model to license out piecemeal, Tencent is using it to upgrade products it already runs at scale, while simultaneously giving away the weights to build developer goodwill and external testing. A Tencent Cloud developer post also advertised two weeks of free WorkBuddy access as a launch promotion, though that offer does not establish the product’s ongoing price.

Why Open the Weights at All?

Tencent’s official announcement framed the release in broad terms: “Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens.” (tencent.com) That language, paired with the simultaneous product launch, points to a deliberate strategy with several layers. Tencent gets to improve its own software and productivity tools, sell access through its cloud infrastructure and OpenRouter, build developer adoption through downloadable weights, and compete for influence in the coding-agent space all at once.

It is a strategy that echoes what DeepSeek did when it cut output pricing on V4.1-Flash by 70% earlier this year, using aggressive pricing and open weights to pull developer mindshare away from closed, API-only competitors. Where DeepSeek competed mostly on cost, Tencent is competing on integration, betting that owning CodeBuddy and WorkBuddy gives Hy4 a built-in user base that a standalone model release would lack.

The Compute Math: Why 49 Billion Active Parameters Matters More Than 770 Billion

A 770-billion-parameter headline number invites comparisons to supercomputer-scale infrastructure, but that framing misses how MoE architecture actually works. Because Hy4 preview only activates 49 billion parameters per token, its inference cost sits far closer to a mid-sized dense model than a trillion-parameter system that runs every weight on every request. That is the same tradeoff behind recent releases from Mistral, DeepSeek, and GLM: enormous parameter counts for knowledge capacity, with a sparse routing layer keeping the compute bill manageable.

That said, the storage and memory story is less forgiving. Serving a 770-billion-parameter checkpoint, even one that only activates 49 billion parameters at a time, still means loading the full model into memory across a multi-GPU cluster, since any token could route to any of the 256 experts in a given layer. Tencent’s documented deployment path targets an eight-way tensor-parallel setup for the FP8-quantized checkpoint, with the MTP layer enabled for speculative decoding, which gives a rough sense of the hardware footprint required to run it well.

Deploying Hy4 Preview: What Developers Get Out of the Box

Tencent shipped Hy4 preview with day-one deployment recipes rather than leaving integration to the community. Two inference engines are officially supported, each with a prebuilt Docker image tied to the release.

# vLLM deployment image
docker pull vllm/vllm-openai:hy4-preview

# SGLang deployment image
docker pull lmsysorg/sglang:hy4-preview

A full fine-tuning pipeline ships alongside the inference images, supporting both the LLaMA-Factory and ms-swift workflows, plus Tencent’s own AngelSlim toolkit for further quantization and compression. For teams that want to run Hy4 preview without managing their own GPU cluster, the model is also available through Tencent Cloud and through OpenRouter, giving developers outside China a path to test it without standing up local infrastructure.

Limitations Tencent Itself Flagged

Tencent has been unusually candid about what does not work yet. The model card lists known issues directly, including a tendency to spend longer than necessary reasoning through complex tasks and a habit of over-verifying its own work, both of which translate into slower, more expensive inference for a given task, per MindStudio’s review of the documentation. Tencent describes this release as an early checkpoint with headroom left in both pre-training and post-training, not a finished flagship.

That admission matters for anyone deciding whether to build on Hy4 preview today. The company’s stated plan mirrors what it did with the prior Hy3 preview release: ship early, gather feedback on what breaks in production, and iterate quickly toward a final version. Tencent says that approach made Hy3 substantially better between its preview and final release, and the company plans to repeat the pattern here.

Historical Context: From Hunyuan’s Early Models to a Flagship Contender

Tencent Hunyuan has spent the past several years building out a family of models spanning text, image, and video generation, but Hy4 preview marks the clearest attempt yet to put a Hunyuan text model directly in the same conversation as GLM, Kimi, DeepSeek, and Mistral on engineering benchmarks specifically. The jump from Hy3 to Hy4 preview is not just a version bump. It comes with a new attention mechanism, a new residual design, and a public blind evaluation that Tencent was willing to publish even though the margins over GLM-5.3 and Kimi K3 are narrow rather than decisive.

That willingness to publish close results, rather than only touting wins, sets Hy4 preview apart from some of the more aggressive benchmark marketing seen elsewhere in the open-weight race this year. It also lands at a moment when the field is crowded. Mistral’s Large 4 debuted at 675 billion parameters with 41 billion active, Z.ai’s GLM-5.3 made headlines for different reasons after researchers reported a 100% safety-bypass rate on certain exploit-building tasks, and a wave of smaller, efficiency-focused labs like Reflection AI pushed their own open-weight bets with Beam.

Market Impact: What This Means for Enterprise AI Buyers

For enterprise teams evaluating open-weight models, Hy4 preview adds another serious option to a list that already includes GLM-5.3, Kimi K3, DeepSeek V4.1-Flash, and Mistral’s Le Chonk family. The practical calculus for most buyers comes down to three factors: inference cost per token, context length for document-heavy workflows, and licensing terms that allow commercial redistribution. Hy4 preview scores well on the first two, with its 49-billion active-parameter footprint and 1-million-token context window, and the Apache 2.0 license clears the third bar without the ambiguity that has dogged some other open releases.

Where Hy4 preview is weaker, by Tencent’s own admission, is production readiness. A model that over-verifies its own output and runs longer reasoning chains than necessary is not yet the cheapest option to run at scale, even with a favorable active-parameter count. Enterprise teams with latency-sensitive applications will likely wait for a non-preview release before committing, while teams building coding agents or research tools, where Tencent says the model shows its biggest gains, have more reason to test it now.

The Broader China Open-Weight Race

Hy4 preview’s release pattern says something about where China’s open-weight strategy is heading in late 2026. Tencent explicitly credited its attention design to both DeepSeek’s and GLM’s prior research, an acknowledgment that the leading Chinese labs are converging on similar technical approaches to the same problem: making very long context windows affordable to run. At the same time, each lab is differentiating on deployment, with Tencent leaning on its existing consumer and enterprise products, DeepSeek leaning on aggressive pricing, and GLM leaning on raw benchmark positioning.

That convergence-plus-differentiation pattern is likely to keep defining the space through the rest of 2026. No single Chinese lab has established a decisive, independently verified lead over the others, and Hy4 preview’s own blind-eval numbers, with win rates in the high 40s and low 50s against its closest rivals, underline just how close the field actually is right now.

What Comes Next: 5 Predictions

  • A non-preview Hy4 release before year-end. Tencent’s own comparison to the Hy3 cycle suggests a faster-than-usual iteration, with the company gathering feedback from CodeBuddy and WorkBuddy usage to cut down the reasoning-length and over-verification issues flagged at launch.
  • Independent benchmarks will narrow the gap further. Once third-party evaluators run Hy4 preview against GLM-5.3, Kimi K3, DeepSeek V4.1-Flash, and Mistral Large 4 on a shared test set, expect the margins to tighten even more than Tencent’s own 47-51% win rates suggest.
  • More Chinese labs will publish blind, human-judged evaluations. Tencent’s willingness to show narrow wins rather than only touting victories could push rivals toward similarly transparent evaluation formats instead of cherry-picked leaderboard screenshots.
  • OpenRouter and Tencent Cloud usage will be the real adoption signal. Download counts on Hugging Face measure curiosity. Sustained token volume through OpenRouter and Tencent Cloud will show whether developers outside China actually build on Hy4 preview rather than just testing it once.
  • Pricing pressure spreads to adjacent model classes. If Hy4 preview’s active-parameter efficiency holds up under real workloads, expect downward pricing pressure on mid-tier coding-agent models from both Chinese and Western labs trying to match its cost-per-token profile.

How Hy4 Fits Alongside DeepSeek, Mistral, and Qwen

None of the current open-weight flagships look identical on paper, which makes head-to-head reading tricky. Mistral’s Le Chonk family drew its own scrutiny this year after a parameter-count dispute between the company’s official 675 billion figure and a third-party claim of 1.05 trillion went unresolved. DeepSeek, by contrast, has leaned on pricing over parameter bragging rights, and its V4.1-Flash line continues to undercut rivals on cost per token. Reflection AI’s Beam took a different angle entirely, betting on a smaller, more efficient active-parameter footprint to compete with far larger models on raw compute savings rather than headline scale.

Hy4 preview sits closer to GLM-5.3 and Kimi K3 in positioning: large MoE backbones aimed squarely at coding and long-document work, evaluated head-to-head by Tencent’s own blind panel rather than a public leaderboard. That makes it hard to declare an outright winner across the field. What is clear is that the gap between the top Chinese open-weight models has narrowed to single-digit percentage points on Tencent’s own numbers, a far cry from the double-digit leads some labs claimed earlier in the year.

Frequently Asked Questions

How many parameters does Tencent Hy4 preview have?
Hy4 preview has 770 billion total parameters, of which 49 billion are active per token because of its Mixture-of-Experts architecture. A separate Multi-Token Prediction module adds roughly 10 billion parameters, with 0.7 billion active, for speculative decoding.

When was Hy4 preview released?
Tencent open-sourced Hy4 preview and the FP8-quantized Hy4-preview-FP8 on August 28, 2026, publishing the weights on Hugging Face, ModelScope, GitCode, and CNB under an Apache 2.0 license.

How does Hy4 preview compare to GLM-5.3 and Kimi K3?
In Tencent’s blind evaluation of 203 engineering tasks judged by 163 internal experts, Hy4 preview scored an average of 2.99, versus 2.92 for GLM-5.3 and 2.94 for Kimi K3. Win rates were 46.8% against GLM-5.3 and 51.2% against Kimi K3, both narrow margins rather than clear victories.

What context window does Hy4 preview support?
Hy4 preview supports up to 1 million tokens of context, putting it in the same long-context tier as other 2026-era open-weight flagships.

Can developers outside China run Hy4 preview?
Yes. The weights are open on Hugging Face, ModelScope, GitCode, and CNB, and Tencent also makes the model available through Tencent Cloud and OpenRouter, with official deployment images for vLLM and SGLang.

Is Hy4 preview ready for production use?
Tencent describes it as a preview, not a finished flagship. The company flagged known issues including overly long reasoning chains and excessive self-verification, both of which add inference cost, and says it plans to iterate toward a final release the way it did with the earlier Hy3 model.

Which Tencent products already use Hy4 preview?
Tencent rolled the model into CodeBuddy, WorkBuddy, Yuanbao, and ima at launch, alongside availability through Tencent Cloud and OpenRouter.

Is Hy4 preview’s benchmark lead over rivals independently verified?
No. The 2.99-versus-2.92-versus-2.94 scores come from Tencent’s own internal blind panel. No independent lab has published a matching head-to-head comparison against GLM-5.3, Kimi K3, DeepSeek V4.1-Flash, or Mistral Large 4 using a shared test set.