๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
๊ฐ€์ž… April 2024
183 ํŒ”๋กœ์ž‰ ์ค‘    16K ํŒฌ
Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.๐Ÿ“œ Apache 2.0. ๐Ÿค– โšก Reaches 23.9โ€“25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. ๐Ÿ† Retains most of the FP16 base modelโ€™s quality, with an average gap of just 2.8 points across five benchmarks. MMLU-Pro rises from 65.8 to 70.1. ๐Ÿง  SSD expert offload keeps most weights outside active memory, while a prerouter predicts which experts each layer will need next. ๐Ÿ›  Recover-LoRA offsets quality loss from 4-bit quantization and expert offloading. The current Preview runs locally through MLX on Apple Silicon.
๋” ๋ณด๊ธฐ