Tencent Hy4: 770B Open Weights, 82% Cheaper Output Than Kimi K3

Tencent Hy4 preview is a 770-billion-parameter open-weight model released under Apache 2.0 on August 28, 2026, with 49B active parameters and a 1,048,576-token context window. Tencent Cloud prices it at $0.834 per million input tokens and $2.501 per million output — roughly one-eighth the output cost of GPT-5.6 Sol and 82% below Kimi K3. Weights are on Hugging Face.

Tencent open-sourced its largest model to date on Friday, and the interesting number is not the parameter count. It is the price tag attached to weights anyone can download, modify and resell.

That combination — frontier-adjacent scores, permissive licensing, and output tokens at $2.501 per million — is the part that moves money.

What is Tencent Hy4 preview?

Tencent Hy4 preview is a mixture-of-experts language model with 770B total parameters, of which 49B activate per token. Tencent released it on August 28, 2026, published the weights on Hugging Face under the Apache 2.0 license, and shipped it simultaneously into its own consumer and developer products.

According to Tencent’s announcement, the model is live in WorkBuddy, CodeBuddy, Yuanbao and ima, with API access through Tencent Cloud TokenHub and OpenRouter.

WorkBuddy and CodeBuddy are free for two weeks from launch. That is a customer-acquisition subsidy, not a pricing model.

The architecture, briefly

The Hugging Face model card lists 78 layers — the first a dense FFN, the remaining 77 MoE — with 256 routed experts plus one shared expert per MoE layer, and top-8 routing per token.

Vocabulary size is 120,832. Attention is what Tencent calls “Gated Sparse Attention with IndexCache,” reusing sparse indices across layers to keep the million-token window affordable to serve.

There is also a built-in multi-token-prediction layer of 10B parameters (0.7B activated) for speculative decoding. Tencent says system-level optimization lifted end-to-end throughput 31.8% against its own baseline.

How much does Tencent Hy4 cost?

Tencent Cloud lists $0.834 per million input tokens, $2.501 per million output tokens, and $0.042 per million on cache hits. In renminbi terms Tencent quotes ¥6 input and ¥18 output per million tokens.

Those are the headline numbers, and they are aggressive against every comparable model.

Per KuCoin’s summary of the launch materials, that input price is 25% below GLM-5.3 and 70% below Kimi K3; on output it is 36% below GLM-5.3 and 82% below Kimi K3.

Model Input / 1M Output / 1M Context Weights
Tencent Hy4 preview $0.834 $2.501 1,048,576 Apache 2.0
GPT-5.6 Sol (base tier) $4.00 $20.00 n/d Closed
GPT-5.6 Sol (long context) $8.00 $30.00 n/d Closed

On output — the expensive half of any agentic workload, where the model writes code, calls tools and re-reads its own work — Hy4 runs at roughly one-eighth of GPT-5.6 Sol’s base rate.

For a coding agent burning 50 million output tokens a month, that is the difference between about $1,000 and about $125. The comparison only holds if the cheaper model finishes the job in a similar number of tokens, which is exactly the assumption worth testing.

How does Hy4 score on benchmarks?

Hy4 posts strong software-engineering numbers and slightly trails the closed frontier on reasoning. Tencent reports 92.3 on GPQA Diamond, 65.7 on SWE-Bench Pro, 82.9 on SWE-Bench Multilingual, 62.9 on SkillsBench v1.1, and 64.3 on Deep-SWE.

The Deep-SWE result is the standout: 64.3 against 28.0 for the previous generation. That is not incremental.

Against closed models the gap is real but narrow. Hy4 scores 92.3 on GPQA Diamond versus GPT-5.6 Sol’s 94.6, and 85.4 on Terminal-Bench versus Sol’s 88.8.

  • GPQA Diamond: Hy4 92.3 — GPT-5.6 Sol 94.6
  • Terminal-Bench: Hy4 85.4 — GPT-5.6 Sol 88.8
  • SWE-Bench Multilingual: Hy4 82.9
  • SWE-Bench Pro: Hy4 65.7
  • Deep-SWE: Hy4 64.3, up from 28.0

Every one of those figures is vendor-reported. None has been independently reproduced at the time of writing, and the model has been public for about a day.

Is Hy4 better than GLM-5.3 and Kimi K3?

Marginally, on Tencent’s own evidence. Tencent ran a blind evaluation with 163 internal experts across 203 real engineering tasks. Hy4 preview averaged 2.99 out of 4.00, against 2.94 for Kimi K3 and 2.92 for GLM-5.3.

A 0.05-point spread on a four-point scale is not a capability gap. It is a tie with a favorable rounding.

The win-rate breakdown is more honest about how close this is. Against GLM-5.3, Hy4 won 46.8% of comparisons, drew 12.8% and lost 40.4%. Against Kimi K3: 51.2% wins, 7.9% draws, 40.9% losses.

So Hy4 loses roughly two of every five head-to-head comparisons against models that were already open. And per KuCoin’s read of the same materials, Hy4 “did not lead comprehensively in public benchmarks” and lags GLM-5.3 in code and cybersecurity tests.

The differentiator here is price and license, not raw capability. That is worth saying plainly, because the launch framing does not. If you are choosing between cheap open coders, our GLM-5.3 vs DeepSeek V4 Pro comparison and the GLM-5.3-Flash vs Qwen3.8-Flash-Next breakdown cover the alternatives.

Should you self-host Hy4 or use the API?

Self-hosting is viable at a scale that would have required a cluster a year ago. Tencent documents deployment on a single eight-GPU node using the FP8 quantized variant, with prebuilt containers for vLLM and SGLang at tensor-parallel size 8.

The full BF16 checkpoint and an FP8 build are both published, along with the AngelSlim toolkit for further compression.

The API case is stronger for anyone under roughly 100 million tokens a month. One eight-GPU node of current-generation accelerators plus the engineer who babysits it will not come in under $2.501 per million output tokens at low volume.

The self-host case is stronger for three groups: teams with data-residency constraints, teams already running GPU capacity at low utilization, and teams that want to distill Hy4 into something smaller. Apache 2.0 permits all three without a negotiation. Our earlier analysis of Qwen3.8-Max open weights versus API walks through that math in detail.

Who wins and loses financially?

The clearest loser is anyone selling mid-tier closed inference. If a 770B open model at $2.501 output holds up in production, the price umbrella over $20-per-million output tiers gets thinner.

Kimi K3 is the most directly exposed. An 82% output-price gap against a model that wins only 51.2% of blind comparisons is a hard position to hold.

Tencent wins on distribution, not on API margin. At ¥6 per million input, the API is a loss leader that routes developers toward Tencent Cloud, WorkBuddy and CodeBuddy — the same playbook that made cheap Chinese inference a strategic instrument rather than a business line. We covered the last turn of that cycle when DeepSeek raised prices by up to 1,100%.

The GPU vendors win either way. Open weights that need eight accelerators per node create hardware demand that a closed API never surfaces on anyone else’s balance sheet.

What are the catches?

Three, and Tencent names one of them itself.

The company acknowledges in its release notes that the model spends “longer than necessary reasoning” and over-verifies its own work. On a consumption-priced endpoint, verbosity is a bill. A model that is 82% cheaper per token but writes twice as many tokens is 64% cheaper, not 82%.

Second, serving is thin. OpenRouter shows a single provider — Tencent Cloud — with 3.16-second P50 latency, 38 tokens per second, and 98.65% availability over three days. There is no failover route.

Third, the context window is a spec, not a guarantee. The model accepts 1,048,576 input tokens but caps completions at 64,000, and nothing in the release claims uniform recall across the full window. For a like-for-like look at long-context pricing, see our cheapest 1M-context model comparison.

And it is called “preview” for a reason.

Frequently asked questions

Is Tencent Hy4 preview free to use?

The weights are free under Apache 2.0. The hosted API is not — it costs $0.834 per million input tokens and $2.501 per million output. WorkBuddy and CodeBuddy are free for two weeks from the August 28 launch.

Can I use Hy4 commercially?

Yes. Apache 2.0 permits commercial deployment, modification, distillation and redistribution without a separate license negotiation or revenue threshold.

What hardware do I need to run Hy4?

Tencent documents a single eight-GPU node using the FP8 quantized build, served through vLLM or SGLang at tensor-parallel size 8. Minimum memory figures are not published.

How big is the context window?

1,048,576 tokens of input, with completions capped at 64,000 tokens.

Is Hy4 better than GPT-5.6 Sol?

Not on published benchmarks. Hy4 scores 92.3 on GPQA Diamond against Sol’s 94.6, and 85.4 on Terminal-Bench against 88.8. It is cheaper by roughly 8x on output.

Where can I download the weights?

Hugging Face at tencent/Hy4-preview, with code and deployment instructions on GitHub.

Have the benchmarks been independently verified?

No. All published scores are vendor-reported as of August 29, 2026.

The bottom line

Hy4 preview is not a capability breakthrough. It wins its own blind evaluation by 0.05 points and loses 40% of head-to-head comparisons against models that were already open-weight.

It is a pricing event. Tencent shipped near-parity performance under Apache 2.0 at 82% below Kimi K3’s output rate and roughly one-eighth of GPT-5.6 Sol’s, and put the weights on Hugging Face the same day.

Tencent’s own README calls it “another step change in capability — the largest generation-over-generation gain we’ve measured.” The blind-evaluation table does not support that framing. The invoice does.

For anyone running high-volume agentic workloads, Hy4 is worth a benchmark run this week — with token-consumption logging turned on, because the verbosity Tencent admits to is where the savings go to die. For anyone selling inference above $20 per million output tokens, the floor moved again.

Sources

Comments

Leave a Reply

Discover more from Wealth Engine

Subscribe now to keep reading and get access to the full archive.

Continue reading