Register and share your invite link to earn from video plays and referrals.

Lorenzo Xiao
@lrzneedresearch
Agentic AI for enterpise/Human-centered NLP Previously @LTIatCMU
529 Following    2.8K Followers
COCO Lab's first meeting with Yejin! Felt like I was introducing my kids to my mom 😂 Very special moment for me. A fun heartwarming meeting with @YejinChoinka @sameerfettes @shubin_kim @bigohofone @Benjamin_eecs and Sunwoo! Thanks for your precious time and looking forward to our exciting projects! 🚀
Show more
Building AI tutors requires feedback from students, but real learner studies are slow and hard to scale. 📈 Just in: StudentSim (Microsoft×UIUC) trains personalized AI student simulators 🤖 from real learner 🧑‍🎓 records and evaluates whether they both match a specific student’s behavior and can respond to tutor guidance. ✏️ Across three learning domains, StudentSim outperforms strong student simulator baselines, including GPT-5.4. In a chess tutor RL study, using StudentSim as guidance helpfulness feedback led to tutor guidance that human experts rated higher in accuracy, guidance quality, and personalization. 📄 Arxiv: 2609.01591 💻 Code: microsoft/StudentSim
Show more
Hello Monday! Today I want to shout out to @NSF and @ACM_SIGWEB for providing $10,000 each ($20,000 total!) to support students and researchers in attending #HCOMP2026# + #CI2026#. All awardees were notified! Regular registration ends on Sep 4 (FRIDAY):
Show more
We’re growing our evaluation team at @thinkymachines. We care a lot about whether models are actually useful and whether our measurements are good enough to serve as the y-axis for scaling research. We’ll work across a wide range of problems and support several workstreams, including building signal-bearing internal evals for predictive scaling, turning real user feedbacks and product workflows into evals that close usability gaps, auditing graders and harnesses, and developing new benchmarks for customizability. If you want to work on this, apply here:
Show more
🎉 Our paper “LLM Probability Concentration: How Alignment Shrinks the Generative Horizon” is accepted by TMLR! The short version: 📉 LLM generation usually self-narrows ✂️ Alignment compresses the horizon further ⚡ Unexpected context can reopen it locally We study these dynamics with Branching Factor (BF)—exponentiated length-averaged entropy, interpreted as the effective number of plausible next steps. Direct evidence first: across tasks, BF typically declines as more tokens are generated, in both base and aligned models. Then we intervene. Replacing the model’s own prefix with equally long random tokens makes BF jump; continued autoregressive generation narrows it again. Self-narrowing is therefore a robust aggregate tendency, not a monotonic token-wise law—and not an artifact caused only by alignment. Alignment is a separate force: it lowers BF by 2–5× overall and up to ~10× at early positions. Reasoning models also remain lower-BF than direct-answer aligned models at matched output lengths, so their concentration is not merely because they generate longer. Most importantly, BF can rise in a structured setting. In synthetic agentic tasks, we hold the task, environment state, and current plan fixed, then change only the next environment feedback. A plan-invalidating surprise raises BF relative to normal progress; continued generation subsequently smooths the local spike. So the dynamics are not simply “entropy collapse.” The model narrows, surprise can reopen its consideration set, and self-conditioning narrows it again. Practical implication: when exploration matters, branch early—or introduce genuinely NEW information before the model has fully committed. 🧵 Learn more from our Branching Factor v1.1 tweet:
Show more
Even within individuals, certain traits are remarkably persistent. Think risk tolerance, values, conscientiousness, among others. Understand these fundamental traits deeply, and you can simulate people faithfully. Everything else is an expression of those traits.
Show more
This Fall at CMU we're teaching a new course on AI Agents! The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL. We'll try to balance theory and practice, and introduce modern frameworks and best practices.
Show more
0
53
2.3K
281
Forward to community
How I feel working on LLM compression / efficiency stuff while looking at GSM results
RL helps models learn how to reason with different strategies, but some strategies are more effective than others. But are the strategies learned by RL the ones that are most effective in improving accuracy? Our new work finds that the answer is "not always"!
Show more
My lab's research is supported by both @thinkymachines and @river_ai_inc, and I asked one of my students to compare the baseline and RL (GRPO) performance of GLM 5.2 through both APIs, to see how much they disagree. The task is memory based agentic long-horizon, and tldr the two APIs have very similar performance! I'm quite shocked how close the final model outcomes are.
Show more
Fun fact, 20% of @simile_ai is made up of our research siblings (i.e., labmates) and relatives from @percyliang’s and @msbernst’s labs! Welcome to the team, @tifding!
Hello Monday! Today is the First Day of School at Penn State and we have great news: #HCOMP2026# + #CI2026# *early registration deadline* was extended! Come join us in Sep!
RIP UIUC If this was a year ago I'll think about how to argue deep learning alpha research is mandatory for iSchool graduation lmaoo
What if LLMs and agents could learn continuously from every interaction without retraining? Our new review on Test-Time Learning explores how deployed models can turn experience into adaptive state through a persistent write–read loop, improving future behavior over time.
Show more
PACE shows that a small, carefully selected set of cheap non agent benchmark questions can accurately predict expensive benchmark performance. Thanks so much Yueqi for leading this project. I’m really glad to contribute to it!
Show more
What if LLMs and agents could learn continuously from every interaction without retraining? Our new review on Test-Time Learning explores how deployed models can turn experience into adaptive state through a persistent write–read loop, improving future behavior over time.
Show more