Register and share your invite link to earn from video plays and referrals.

Search results for PostTraining
PostTraining community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including PostTraining
$GOOGL's next flagship model, Gemini 4, is reportedly performing well in internal pretraining evaluations, but still needs to complete posttraining, per WSJ.
GOOGLE SET TO RELEASE GEMINI 3.8 FLASH THIS WEEK: WSJ Google $GOOGL is expected to release Gemini 3.8 Flash, internally called “Skimaki,” as soon as Wednesday, with a major focus on coding. In internal head-to-head tests inside Google’s coding tool Jetski, engineers reportedly preferred 3.8 Flash over Anthropic’s Opus model. Flash is designed to be smaller, cheaper and faster than Google’s flagship Pro models. Google has also scrapped internal Gemini 3.5 Pro candidates that weren’t sufficiently better than Flash, while Gemini 4 has performed well in pretraining evaluations but still needs to complete posttraining.
Show more
TL;DR: "The model you use is not the model that was pretrained." An 87-slide lecture that maps post-training — turning a base model into a useful policy — across six stages: imitate, compare, explore, verify, transfer, anticipate. Title: Post-Training LLMs (Kawin Ethayarajh, AI and Economics Summer Institute 2026) URL: Key points 📝 SFT: teach by demonstration. Can imitate whole trajectories (plan, tools), but it's fundamentally imitative ⚖️ Offline preference optimization: learn from win/loss. DPO drops the reward model; KTO learns per-outcome, no SFT needed 🎲 Online RL: the current policy explores. REINFORCE/PPO/GRPO; rewards are mostly sequence-level ✅ RLVR: replace the reward model with a checker. Binary pass/fail is enough; scales more predictably than RLHF 🧪 Environments: an RL env is task + data + interface + tests. The bottleneck shifts from examples to environments 🔁 On-policy distillation: the teacher scores the student's own rollout. Copies an RL policy in 7–10x fewer steps 🌐 World adaptation: environments adapt back to agents. "Mecha-nudges" raise machine-readability A systematic map, threaded with an economist's lens. #LLM# #PostTraining#
Show more
No amount of post-training cleanly fixes weak long-horizon foundations: noisy trajectories compound errors, sparse rewards misassign credit, and conflicting teachers trigger forgetting. This paper finds, ff you want agents to stay reliable over long tasks, give them clean world-model and long-trajectory training first, use OPD (on-policy distillation) when reward signals get too sparse, and avoid merging teachers with incompatible planning strategies. Suboptimal trajectories were especially damaging because small mistakes accumulated until middle and long tasks nearly collapsed. For post-training, OPD handled longer, noisier settings better than outcome-reward GRPO because teacher feedback arrived throughout the trajectory instead of only at the end. – arxiv. org/abs/2607.24720v1 Title: "The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation"
Show more
Cosmos 3 Post-Training in Action With Aigen and Linker Vision | Cosmos Labs
Built as a post-training variation on open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price.
Show more
"What is Missing from AI Post-Training AI" The bottleneck for autonomous AI R&D may be knowing when to abandon the current strategy, and not executing it better. So AI agents can post-train models end-to-end, but they rarely rethink the strategy they started with. Across 1,338 trajectories, agents were good at debugging, tuning, and iterating, yet almost all improvement stayed inside the initial training approach. Even 2-8x more inference compute mostly produces more optimization inside the same strategy, while human guidance can change the strategy initially but the agent eventually falls back into local refinement. So whether agents can spontaneously realize when their current strategy is wrong and decide to abandon it should be the next step forward.
Show more
Who’s building @tacobell’s post training team
SLAI T-Rex Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD paper:
Showed these guys my secret post-training cake recipe. Secret ingredient: kale. #BringYourGame# #StriveForGreatness#
0
157
3.4K
684
Forward to community