RSI WATCH
Self-modification of agent policy via "dreaming" = exploration of previous histories using a replay simulator. Everything we know about our own brains via introspection can be tried for AI improvement...
"the agent can dream over many alternative exploration strategies before redeploying the improved policy online"
Meta-Layer RSI Loop (Dream-RSI): ... a recursive self-improvement loop that continuously collects discovery histories through online exploration, constructs replay simulators from history to refine meta-exploration strategies via dreaming, and redeploys the upgraded policy online
Empirical Validation: We conduct experiments to demonstrate that Dream-RSI improves both discovery effectiveness and efficiency in several settings.
RSI is open source, open source is RSI.
A shape, a seed of an idea of something that could exist, and then many minds in agreement that this hyperobject should exist.
From here, everyone pours in capital: their time, tokens, agents, etc.
Coalescing,
“Dream-RSI: Recursive Self-Improvement through Evolving Worlds”
AI agents can search for better solutions, but they’re usually stuck using the same search strategy over and over.
Dream-RSI lets the agent learn how to search better by turning its past exploration into a simulator, where it can cheaply replay different strategies before spending compute in the real world.
So the agent improves not just its solutions, but also the process it uses to discover them.
fascinating RSI loop for AI discovery in this Google paper
but again, google is great at writing academic papers and somehow much worse at turning that research into a great AI product. i think they need a bit more of the openai/anthropic culture
or a return to the innovative google of the 2000s
i guess RSI is basically a checkbox for AGI
but we'll probably keep debating it because RSI itself is loosely defined
in a broad sense, we already have "weak" or partial RSI loops, depending on how you define them -- openai for example has made it clear that it has systems with some of these weaker capabilities
Bitcoin "Death Cross" + Oversold RSI + RSI rise back to 56 + rising back 50-week moving average =
Bottom always in.
Reaction at 50-week always different, but bottom always in.