注册并分享邀请链接,可获得视频播放与邀请奖励。

stevibe
@stevibe
LLM. Local AI addict. Building @BenchLocalAI Builds things nobody asked for. Benchmarks things for fun.
加入 July 2009
1.3K 正在关注    27.7K 粉丝
Google dropped MTP versions of Gemma4. Ran them on my DGX Spark. The 31B dense model went from 3.94 → 8.91 tok/s. That's +126%. Full results: [26B A4B] > 25.24 → 31.69 tok/s (+25.6%) > TTFT 755 → 332ms (-56%) [31B] > 3.94 → 8.91 tok/s (+126%) > TTFT 599 → 378ms (-37%) If you're not running MTP, you're leaving free perf on the table.
显示更多
0
19
118
10
转发到社区