註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
加入 February 2026
1.7K 正在關注    12.4K 粉絲
Qwen-Image-2.0-RL Technical Report Paper: They fine-tune VLMs to score pointwise with CoT, so the reward judges alignment against an explanation and is harder to game. OPD then closes the train-inference gap on the diffusion backbone.
顯示更多
0
4
99
13
轉發到社區