๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
Interesting HuggingFace post. What if the easiest way to make long-context Local AI faster is to simply STOP feeding it everything? A new HF forum experiment tested 600 distractor-heavy long-context prompts with Qwen2.5-7B 4-bit. Instead of sending the entire ~14K-token context, a MiniLM semantic selector kept only the most relevant ~60%. The result ... ๐Ÿ“š Input tokens 13,124 โ†’ 7,702 โžก๏ธ -41% โšก Prefill 871 ms โ†’ 463 ms โžก๏ธ -47% ๐Ÿš€ End-to-end latency 944 ms โ†’ 519 ms โžก๏ธ -45% ๐Ÿ’พ Peak VRAM 11.71 GiB โ†’ 9.16 GiB And here's the weird part... ๐ŸŽฏ Token F1 actually IMPROVED: 0.534 โ†’ 0.581 So the model got: โœ… less context โœ… lower VRAM use โœ… almost 2ร— faster prefill โœ… lower total latency โœ… BETTER answers Why? Because more context isn't always better context. Coding agents and RAG systems often drag around enormous histories filled with irrelevant junk. โš ๏ธ This was a controlled distractor-heavy HotpotQA stress test, not proof that deleting 40% of every real-world context will improve accuracy. ๐Ÿ”— Link in ALT.
๋” ๋ณด๊ธฐ