TL;DR XPeng released X-AuT, a compression method that shrinks a speech LLM's audio encoder from 18 to 14 layers with barely any accuracy loss โ and the 16-layer version actually improves accuracy.
Title: X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
URL:
Points
โ๏ธ It uses "progressive pruning," cutting the audio encoder from 18 to 16 to 14 layers in stages
๐ What's cool: single-layer removal scores alone can't predict the best pair, so it explicitly measures layer-pair interactions
๐ Cross-scale distillation from a 1.7B teacher into a 0.6B student clearly beats same-scale self-distillation
๐ The 16-layer model cuts parameters by 10.35% while improving macro-average error from 5.61% to 5.27%
๐ The 14-layer model cuts parameters by 20.70% with only a +0.14pt accuracy drop
๐ On an in-vehicle chip, it cuts encoder time by 21.4% and total latency by 4.7%
๐ Code and models are already public on GitHub and Hugging Face (CC BY-NC 4.0)
What stands out is the careful design of "how to prune smart and recover well," not just cutting layers.
#
SpeechLLM# #
ModelCompression#