가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Tencent AI
@TencentAI_News
The official @Tencentglobal newsroom for AI updates and developer resources.
가입 December 2025
95 팔로잉 중    24K 팬
1.5TB → 214GB. Seven times smaller, barely a dent. That's Hy4 preview. The trick is Sherry — our quantization method that packs weights down to 1.25 bits each. (The image shows how.) It also unlocks a new way to run it: stitch the GPUs you already have across machines, and they work as one. GGUF (quantized): Original model:
더 보기
We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error. Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5 See the details on HF : AngelSlim/Hy4-preview-GGUF Weights & low-bit GGUFs 👇 #LLM# #Quantization# #llamacpp# #Hy#
더 보기
0
86
2.6K
303
커뮤니티로 전달