注册并分享邀请链接,可获得视频播放与邀请奖励。

DailyPapers
@HuggingPapers
Tweeting interesting papers submitted at Submit your own at and link models/datasets/demos to it!
加入 March 2025
4 正在关注    20.4K 粉丝
V-Zero: answer-label-free visual reasoning It uses on-policy distillation with contrastive evidence gating. It trains 5x faster than SFT and 10x faster than RL. 4B model is on Hugging Face.
显示更多
0
2
85
17
转发到社区