Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
280 Following    415 Followers
TL;DR XPeng released X-AuT, a compression method that shrinks a speech LLM's audio encoder from 18 to 14 layers with barely any accuracy loss โ€” and the 16-layer version actually improves accuracy. Title: X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation URL: Points โœ‚๏ธ It uses "progressive pruning," cutting the audio encoder from 18 to 16 to 14 layers in stages ๐Ÿ” What's cool: single-layer removal scores alone can't predict the best pair, so it explicitly measures layer-pair interactions ๐ŸŽ“ Cross-scale distillation from a 1.7B teacher into a 0.6B student clearly beats same-scale self-distillation ๐Ÿ“ˆ The 16-layer model cuts parameters by 10.35% while improving macro-average error from 5.61% to 5.27% ๐Ÿ“‰ The 14-layer model cuts parameters by 20.70% with only a +0.14pt accuracy drop ๐Ÿš— On an in-vehicle chip, it cuts encoder time by 21.4% and total latency by 4.7% ๐Ÿ”“ Code and models are already public on GitHub and Hugging Face (CC BY-NC 4.0) What stands out is the careful design of "how to prune smart and recover well," not just cutting layers. #SpeechLLM# #ModelCompression#
Show more