Register and share your invite link to earn from video plays and referrals.

Chenghao Yang
@chrome1996
Senior AS @Microsoft. Ph.D. @UChicago Ex-SR @google Ex-Scientist @AWS. Ex-RA @jhuCLSP @columbianlp @TsinghuaNLP. Ex-Intern @IBM @AWS. Opinions are my own.
770 Following    2.3K Followers
🎉 Our paper “LLM Probability Concentration: How Alignment Shrinks the Generative Horizon” is accepted by TMLR! The short version: 📉 LLM generation usually self-narrows ✂️ Alignment compresses the horizon further ⚡ Unexpected context can reopen it locally We study these dynamics with Branching Factor (BF)—exponentiated length-averaged entropy, interpreted as the effective number of plausible next steps. Direct evidence first: across tasks, BF typically declines as more tokens are generated, in both base and aligned models. Then we intervene. Replacing the model’s own prefix with equally long random tokens makes BF jump; continued autoregressive generation narrows it again. Self-narrowing is therefore a robust aggregate tendency, not a monotonic token-wise law—and not an artifact caused only by alignment. Alignment is a separate force: it lowers BF by 2–5× overall and up to ~10× at early positions. Reasoning models also remain lower-BF than direct-answer aligned models at matched output lengths, so their concentration is not merely because they generate longer. Most importantly, BF can rise in a structured setting. In synthetic agentic tasks, we hold the task, environment state, and current plan fixed, then change only the next environment feedback. A plan-invalidating surprise raises BF relative to normal progress; continued generation subsequently smooths the local spike. So the dynamics are not simply “entropy collapse.” The model narrows, surprise can reopen its consideration set, and self-conditioning narrows it again. Practical implication: when exploration matters, branch early—or introduce genuinely NEW information before the model has fully committed. 🧵 Learn more from our Branching Factor v1.1 tweet:
Show more