登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Tinker
@tinkerapi
I tink, therefore I am. Post-training API by @thinkymachines
参加 January 2026
1 フォロー中    13.4K ファン
Parameter-efficient fine-tuning isn't just cheap, it's what makes formal guarantees of model learning possible. Compress an RLVR update into a small LoRA and you can set a floor on how it will generalize to unseen data. Sharp paper from @maxYuxuanZhu , @rohanalur, and @ddkang.
もっと見る
New research from Bridgewater AIA Labs, UIUC, and MIT: we prove what we believe to be the first non-vacuous generalization bounds for reasoning LLMs on real-world problems. RLVR powers frontier reasoning capabilities yet its generalization to unseen data has remained an open theoretical question and deployment blocker for practitioners. Our generalization bounds for RLVR deliver provable high-probability lower bounds of the accuracy for billion-parameter RLVR models on unseen data, which can provide guidance on safely deploying RLVR. 1/9
もっと見る