Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Hy3 GPTQ Int4

Hy3 GPTQ Int4 is AngelSlim's 4-bit GPTQ quantization of Tencent's Hy3: a 299-billion-parameter Mixture-of-Experts model with 21B active parameters per token, a 256K-token context window, selectable reasoning effort and tool calling, released under Apache 2.0.

Model

Hy3 is a 299-billion-parameter Mixture-of-Experts model from the Tencent Hy Team that activates 21B parameters per token: 80 transformer layers, 192 routed experts with 8 activated per token, grouped-query attention with 64 heads and 8 KV heads, a 120,832-token vocabulary, and a 256K-token context window. A 3.8B-parameter multi-token-prediction (MTP) layer enables speculative decoding. Reasoning effort is selectable per request, from direct answers (no_think) through low to high chain-of-thought, and the model supports tool calling via dedicated parsers. Building on the Hy3 preview, Tencent scaled up RL post-training and reports gains on reasoning, agentic and long-context tasks alongside a lower hallucination rate and more stable tool calls.

This variant is the GPTQ Int4 quantization published by AngelSlim, Tencent's open-source model-compression toolkit: 4-bit weights with a group size of 128 and symmetric quantization, with minimal accuracy loss according to AngelSlim's benchmarks. Both the base model and this quantization are released under the Apache 2.0 license. Full details are in Tencent's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Hy3 GPTQ Int4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model AngelSlim/Hy3-GPTQ-Int4

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute