가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Alexey Fateev
@superalesha
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ - 2xDGX Spark 🚀96GB VRAM | Local AI
가입 January 2026
325 팔로잉 중    3.7K
2.38 bpw and MRCR moves 0.2 points. That sounds like a big win. I want to check that number myself but 214GiB does not fit in 96GB, so the model is out of reach here. Their encoder change is not. It works on any model. @TencentHunyuan ship imatrix.gguf and the layer selection rule and I will run the same recipe on a MoE that fits on 4 3090s. I even have a couple of good candidates for this surgery.
더 보기
We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error. Accuracy barely moves vs BF16 📊 MCP Atlas 83.7→83.2 📊 SWE-Bench multi 82.9→81.3 📊 MRCR 81.3→81.1 📊 IFBench 73.5→72.5 See the details on HF : AngelSlim/Hy4-preview-GGUF Weights & low-bit GGUFs 👇 #LLM# #Quantization# #llamacpp# #Hy#
더 보기