TL;DR XPeng released X-AuT, a compression method that shrinks a speech LLM's audio encoder from 18 to 14 layers with barely any accuracy loss — and the 16-layer version actually improves accuracy.
Title: X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
URL:
Points
✂️ It uses "progressive pruning," cutting the audio encoder from 18 to 16 to 14 layers in stages
🔍 What's cool: single-layer removal scores alone can't predict the best pair, so it explicitly measures layer-pair interactions
🎓 Cross-scale distillation from a 1.7B teacher into a 0.6B student clearly beats same-scale self-distillation
📈 The 16-layer model cuts parameters by 10.35% while improving macro-average error from 5.61% to 5.27%
📉 The 14-layer model cuts parameters by 20.70% with only a +0.14pt accuracy drop
🚗 On an in-vehicle chip, it cuts encoder time by 21.4% and total latency by 4.7%
🔓 Code and models are already public on GitHub and Hugging Face (CC BY-NC 4.0)
What stands out is the careful design of "how to prune smart and recover well," not just cutting layers.
#
SpeechLLM# #
ModelCompression#