가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Corey J. Gallon
@CoreyGallon
Sharing insights from the frontiers of AI Engineering. 🇦🇺 Technologist. Investor. Coffee Nerd.
가입 December 2008
319 팔로잉 중    544 팬
Every leaderboard you've seen evaluates language models as if they wake up with no memory of anything they've ever done. @pgasawa, a CS PhD student at UC Berkeley, argues in "Beyond Static Intelligence: Evaluating Continual Learning" that this is why we have no idea whether models actually learn. It's on @aiDotEngineer's YouTube. If you're building systems with memory, notepads, or growing context windows, the talk gives you a way to tell whether any of that is producing real improvement or just a stronger base model. - Independent benchmark instances can't measure learning. Chaining AIME problems together doesn't work, because the instances share no structure, so there's nothing for earlier experience to transfer into. - Three design criteria for a continual learning benchmark. Headroom (tasks that actually require online adaptation), shared latent structure across tasks, and a learning mechanism in the environment: scalar reward, error messages, textual feedback. - Cumulative reward confounds learning ability with base model strength. A system can top the leaderboard while improving zero over its own stateless self. - Gain isolates it. Run the system twice through the benchmark, once holding state and once reset between every instance, and take the difference. Reward, gain, and cost all get measured on Pareto frontiers. - Continual Learning Bench 1.0 spans six domains. Blind spectrum monitoring, codebase adaptation, cohort studies in epidemiology, exploitable poker, database exploration, and sales prediction, with instances validated by domain experts. - Concept drift is built into the tasks. A database migration mid-sequence drops columns and changes formats, testing whether a system can discard stale experience and still update from new. - Vanilla in-context learning topped the first leaderboard. Just putting experience in the context beat the more expensive context management systems on reward, and held up on reward-versus-cost and gain-versus-cost too. - Failure modes land on one side of stability-plasticity. A sales forecaster that over-predicted, then under-predicted, jumped straight back to the overprediction instead of splitting the difference. A notepad system in the epidemiology task wrote off cohort definitions as "from a different study schema that doesn't apply here." The schema applied. - The frozen checkpoint may be the sunk cost. Parth's view is that today's continual learning work bolts mechanisms onto models never designed to learn after training, and that a first-principles design might collapse to a single learning phase, with everything after it being deployment. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
더 보기