註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

DailyPapers
@HuggingPapers
Tweeting interesting papers submitted at Submit your own at and link models/datasets/demos to it!
加入 March 2025
4 正在關注    20.1K 粉絲
V-Zero: answer-label-free visual reasoning It uses on-policy distillation with contrastive evidence gating. It trains 5x faster than SFT and 10x faster than RL. 4B model is on Hugging Face.
顯示更多
0
2
85
17
轉發到社區