登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research
参加 November 2023
101 フォロー中    56.5K ファン
Using Tinker with an autoresearch loop is a really effective way to reproduce post-training papers at predictable costs Today there are dozens of self-distillation methods all claiming improvements over each other, and it’s hard to establish which claims hold up We gave agents a Tinker budget to reproduce self-distillation results across models and training setups. With just a few user prompts, they reproduced SDFT’s continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds, and investigated SFT’s failure modes. Read more below:
もっと見る