๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
270 ํŒ”๋กœ์ž‰ ์ค‘    313 ํŒฌ
TL;DR Advanced encoders were losing to older models on sparse retrieval for one surprising reason: a vocabulary mismatch. Fixing it with just 20k adaptation steps sets a new BEIR state of the art. Title: Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps URL: Key points ๐Ÿ”ค Raw case-sensitive vocabularies were splitting single meanings across redundant surface forms ๐Ÿงฉ The proposed Vocabulary Transfer method migrates vocabularies in 3 lightweight steps ๐Ÿ“Š ModernBERT-VT hits nDCG@10 52.4 on BEIR, a new SOTA ๐Ÿ’ฅ A collapsed RoBERTa-large jumps from BEIR score 1.4 to 51.3 โšก Uses under 0.2% of the original pretraining token budget ๐Ÿงช Near-optimal performance reached in just 500 MLM steps ๐Ÿงฌ Also works for domain-specific vocabularies like chemistry It's a nice reminder that what looked like an architectural ceiling turned out to be a fixable vocabulary design problem. #InformationRetrieval# #LLM#
๋” ๋ณด๊ธฐ