注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在关注    56.9K 粉丝
Using Tinker with an autoresearch loop is a really effective way to reproduce post-training papers at predictable costs Today there are dozens of self-distillation methods all claiming improvements over each other, and it’s hard to establish which claims hold up We gave agents a Tinker budget to reproduce self-distillation results across models and training setups. With just a few user prompts, they reproduced SDFT’s continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds, and investigated SFT’s failure modes. Read more below:
显示更多
0
7
169
24
转发到社区