Tencent Shrank a 1.5TB AI Model to 214GB — Here's the Trick
Explains how Tencent used Sherry quantization to shrink a 770B model from 1.5TB to 214GB.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Hy4Model source: provider1,048,576 tokens max native context | 770Bactive 49B | 462 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 62 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 88 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 88 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 97 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 97 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 97 tok/s4-bit quantization |
Explains how Tencent used Sherry quantization to shrink a 770B model from 1.5TB to 214GB.
Tests Hy4 Preview against GLM 5.3 using Inferencer App Q4-INF quantization on Mac Studio M3 Ultra 512 GiB via thinking mode performance, UI testing, censorship checks, complex code modification, game
Tests Tencent's 700B+ MoE Hy4 using Workbuddy app to evaluate its maximum performance capabilities.
Tests Tencent Hy4 preview MIX-STQ1_0 GGUF locally against API versions via Slappis demo, Splunk CTF, and Estonia tasks, evaluating coding and agentic performance preservation.
Tests Tencent HY4 Preview on coding, 3D game dev, and agentic tasks, comparing it against Qwen 3.8, GLM 5.3, DeepSeek V4 Flash Vision, Kimi K3, Claude Opus 5, and GPT-5.6 using WOAI Bench.
Tests Tencent HY4 preview on browser workflows, C++ game creation, FPS development, frontend design, Blender scene generation, and a wrestling game.
Tests Hy4 preview, a new-generation Mixture-of-Experts (MoE) flagship model developed by the Tencent Hy Team.