가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
가입 February 2026
1.7K 팔로잉 중    12.4K
Qwen-Image-2.0-RL Technical Report Paper: They fine-tune VLMs to score pointwise with CoT, so the reward judges alignment against an explanation and is harder to game. OPD then closes the train-inference gap on the diffusion backbone.
더 보기