In 2026, after an April preview, Tencent released Hy3 with 295 billion parameters under Apache 2.0. Commercial users can modify and self-host it without returning to Tencent’s API, leaving the company no guaranteed toll on downstream use.

Key takeaways

  • Tencent released the 295-billion-parameter Hy3 under Apache 2.0 on July 7, 2026, after an April preview.
  • xAI published Grok-1’s 314-billion-parameter weights and architecture under Apache 2.0 in March 2024.
  • Z.ai released GLM-5.2 with MIT-licensed weights and a one-million-token context window.
  • DeepSeek priced V4 Flash at $0.14 per million input tokens and $0.28 per million output tokens.
  • Tencent and Meituan led Even Realities’ $150 million pre-Series B round.

Apache lets Hy3 travel beyond Tencent’s API

Apache gives developers and enterprises broad rights to customize and redistribute Hy3 commercially. A builder can place it behind a private application, adapt it to proprietary data or package it inside another service, all outside Tencent’s direct control.

Quarterly coverage volume: TencentCoverage of Tencent by quarter, 2024 Q4 to 2026 Q3: from 10 to 32 articles per quarter, peaking at 32.322024 Q42026 Q3
Quarterly coverage · Tencent · 2024 Q4–2026 Q3 · current quarter projected

Tencent also said Hy3 competes with GLM-5.1 and GLM-5.2. The company supplied that comparison; the available evidence does not establish independent benchmark leadership or meaningful downstream adoption. Tencent’s strategic bet rests on a release mechanism that can spread Hy3 without benchmark leadership.

Other model providers had already chosen the same route. xAI published Grok-1’s 314-billion-parameter weights and architecture under Apache 2.0 in March 2024. In March 2025, Alibaba released Qwen2.5-VL-32B under the same license. By the time Tencent joined the open-weight model ecosystem, model makers were repeatedly using permissive licenses to recruit downstream development.

From 2024 through 2026, developer framing in coverage of Tencent rose from 6.0% to 16.1%, while consumer framing fell from 42.0% to 26.8%. Attention was already moving toward builders who could carry models into their own systems.

Providers now compete through license terms

China’s model providers now compete across a spectrum of licenses. Z.ai released GLM-5.2 with MIT-licensed weights, a one-million-token context window and an emphasis on agentic coding and long-horizon tasks. Moonshot AI released Kimi K3 under its own Kimi K3 License instead of Apache or MIT. Alibaba launched the 2.4-trillion-parameter Qwen3.8-Max and said it planned to release weights for that model and Qwen3.8-27B.

Each provider chose a different contract with its downstream ecosystem. Apache and MIT minimize legal negotiation and give commercial builders room to modify and redistribute. A custom license can preserve restrictions or bargaining power after the weights leave the provider’s servers. Alibaba reportedly considered requiring heavy commercial users of its next Qwen model to share revenue, an arrangement that would make adoption broad without making it unconditional.

A provider’s license sets the boundaries of diffusion. An enterprise evaluating two sufficiently capable models must compare benchmarks, token prices, redistribution rights, customization rights and exposure to changing commercial terms. Model providers have placed the contract beside the architecture as part of the technical buying decision.

Open weights move competition into serving costs

A team that downloads Hy3 inherits a serving bill for memory, accelerators, networking, scheduling and idle capacity each time users invoke the model; the expense of training it offers no relief at inference time.

Architecture and operations determine that bill alongside model size. Mixture-of-experts systems can separate total capacity from the compute activated for each token. Operators also change delivered cost through quantization, batching, caching, routing and utilization. At fleet scale, they must coordinate hardware and model architecture; a cheap accelerator that sits idle is merely an expensive sculpture.

Investor projections from OpenAI and Anthropic put inference at the center of their cost bases. OpenAI engineers later reportedly found a method that could more than halve inference cost. Both companies must treat serving optimization as a core business function.

Inference costs as a share of revenue in some OpenAI and Anthropic projections

DeepSeek priced V4 Flash at $0.14 per million input tokens and $0.28 per million output tokens even as the company announced substantial price increases across its AI services. Those prices make its margin sensitive to small changes in utilization, memory traffic and generated-token volume. Operators ultimately compete on cost per useful task.

The inference market matters more as weights spread. A developer can replace a model, change the host or route easier tasks to a cheaper system. A serving provider must keep earning each token instead of collecting a scarcity rent from access alone.

Deployment still carries hard constraints. Chinese AI companies reportedly sought overseas compute access in Southeast Asia and the Middle East for Nvidia Rubin chips while Zhipu and Alibaba warned of a widening gap with the United States. Without enough accelerators, an enterprise may still be unable to self-host a model economically.

Governments can narrow the route as well. OpenAI and Google reportedly sold AI services to Singapore-based subsidiaries of Tencent, Alibaba and Baidu, prompting US policymakers to renew calls for tighter controls. Apache can authorize a deployment, but unavailable hardware and changing access rules can still make it impractical.

Long-running agents move the bottleneck into operations

An enterprise gets value only when an agent completes authorized work, survives failures, leaves an audit trail and hands ambiguous cases to the right person. Long-horizon tasks multiply the opportunities for a model error, tool failure or unsafe action to derail the job.

Google introduced the Gemini Enterprise Agent Platform to manage the full lifecycle of agent fleets on Vertex AI. OpenAI added native sandboxing and an in-distribution testing harness to its Agents SDK. Google and OpenAI built those controls because production agents require evaluation, isolation, permissions, monitoring and recovery alongside model intelligence.

Companies have so far used agents primarily to improve efficiency and reduce costs. Those buyers judge a deployment by completed work and avoided labor, giving operational failures an immediate economic price.

The company that owns deployment-layer control can decide which model handles a task, which tools it may call, how failures are retried and where the resulting action appears. Open weights increase the number of eligible engines inside that system. The surrounding operator still has to supply the system itself.

Tencent needs applications to make Hy3 adoption pay

Alibaba has already shown how a model can meet an existing distribution surface. The company connected Qwen with Taobao, Alipay, Fliggy and Amap while pursuing a one-stop AI application for 100 million users. Alibaba can measure Qwen by purchases, payments, trips and completed actions rather than by chatbot visits alone.

Recent investments point to possible interfaces. Tencent and Meituan led a $150 million pre-Series B round for camera-free smart-glasses maker Even Realities. Tencent was also reportedly discussing becoming Manus’s largest investor while Manus’s owners considered unwinding Meta’s $2 billion acquisition. Neither report established a Hy3 deployment.

Tencent’s own game portfolio supplies the caution. Its Lightspeed studio reportedly spent hundreds of millions of dollars over six years on the troubled development of Last Sentinel. Tencent can fund a category and still fail to produce the end-user experience that makes its components valuable together.

Frequently asked questions

What does it actually cost to self-host Hy3?

The piece provides no Hy3-specific hardware, throughput, or per-token serving-cost estimate. It identifies the relevant cost drivers—memory, accelerators, networking, scheduling and idle capacity—but a deployment budget would depend on the operator’s configuration and utilization.

Are there independent benchmarks showing that Hy3 leads rival models?

No. Tencent said Hy3 competes with GLM-5.1 and GLM-5.2, but the available evidence does not establish independent benchmark leadership.

Has Tencent confirmed that Hy3 will be deployed through Manus or Even Realities products?

No. Tencent’s possible Manus investment and its investment alongside Meituan in Even Realities point to potential interfaces, but neither report established a Hy3 deployment.

When will Alibaba release weights for Qwen3.8-Max and Qwen3.8-27B, and what revenue-sharing terms would apply?

The piece says Alibaba planned to release the weights and reportedly considered revenue sharing for heavy commercial users of a next Qwen model. It gives no release date, threshold for heavy use, or finalized commercial terms.

Permissive open-weight releases cited in the piece

  • March 2024 — xAI published Grok-1’s 314-billion-parameter weights and architecture under Apache 2.0.
  • March 2025 — Alibaba released Qwen2.5-VL-32B under Apache 2.0.
  • April 2026 — Tencent previewed Hy3.
  • July 7, 2026 — Tencent released the 295-billion-parameter Hy3 under Apache 2.0.

Apache 2.0 guarantees Tencent no toll on Hy3. It gives the company a chance to make independent deployments pull demand toward its serving stack, agent controls and application surfaces. Because developers can move the weights, Tencent earns a strategic return only when leaving that surrounding system costs more than replacing Hy3.