The fundamental problem with ML conference is that there is upside to doing bad work & *zero* downside. I've unearthed paper-invalidating problems with nearly every conference's Best Paper and Orals that I've touched
(two papers forthcoming on this)
1/2
I would support even stronger quotas: No more than 5 submissions per author sounds reasonable. An author is someone who has supposedly authored, at least part of a paper. Also, authors should be evaluated by their best papers, not how many they got through the process. When I see faculty candidates or researchers with 10-15 papers in one conference, it now counts as a negative signal to me.
#AAAI2027# reportedly received at least 44,000 abstract submissions.
Opinion: The publish or perish era may be over. We have reached a point where publishing alone carries almost no signal
I will give an invited talk on this result at the NExT-Game workshop at ICML tomorrow
Join me if you are interested in knowing the details of the proof, the mistakes in previous papers, and the ChatGPT proof.
I solved an open problem in optimization with ChatGPT Pro.
Consider min-max problems, where we have to minimize with respect to some variable and maximize with respect to other variables to find a saddle-point.
new post on harness engineering for AI self-improvement:
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
On Monday @natalie_collina, Ira, and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: as well as an interactive demo of online calibration algs. See you there!
Yesterday at the COLT business meeting, one senior core member of this community said that the COLT papers are "disconnected from reality".
The world is really changing.
Our COLT 2026 paper!
A single stepsize with high probability gives both fast and robust rates in TD learning, without using any projection.
IMO, Wei-Cheng put together a very interesting proof, with some new tricks as well.
1/11
New paper: A Single Stepsize Suffices for Unprojected Linear TD(0)
Can TD(0) without projection and knowledge of curvature achieve high-probability rates that are both robust and fast under Markovian sampling?
We show: one curvature-free stepsize + PR average suffice. 🧵
Final version of my book (with a new title)
Online Learning: A Modern Introduction Using
Convex Optimization
Especially proud of the Foreword by @NicoloCB!
It'll be printed by Cambridge University Press.
The end of 7 years of updates :)
I solved an open problem in optimization with ChatGPT Pro.
Consider min-max problems, where we have to minimize with respect to some variable and maximize with respect to other variables to find a saddle-point.