๐ A way to train "think longer, get smarter" models without the gradients exploding.
Title: Thinking with Looped Flows
URL:
Looped models that recurrently update hidden states to "think" have long struggled with unstable BPTT (backprop through time). Looped Flows fixes this by borrowing training principles from diffusion models. Here are 3 highlights.
๐ง Training recurrence without BPTT
By training each step with a local loss at gradually decreasing noise levels, the model learns to keep "thinking" stably, without vanishing or exploding gradients.
โฑ More compute at inference, for free
Just using a finer time grid at test time boosts accuracy, from 74.5% at 8 steps to 97.9% at 128 steps on Sudoku, with no retraining needed.
๐ Beats prior looped models on ARC-AGI
It reaches 58.8% on ARC-AGI-1 and 12.2% on ARC-AGI-2, outperforming previous looped-model approaches on both.
A neat new take on test-time compute scaling: thinking longer at inference genuinely pays off.
#
LLMInference# #
MachineLearning#