Register and share your invite link to earn from video plays and referrals.

Search results for SelfEvolvingAI
SelfEvolvingAI community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SelfEvolvingAI
Most “Self-Evolving AI” Is Not Recursive Self-Improvement Yet AI may generate its own data, rewards, skills, and code. But if humans still define what counts as better and approve deployment, the loop is not closed. Zhihu contributor 一口鸟 (@alsaceym) argues that RSI’s hardest bottleneck is reliable verification. 1️⃣ Moving humans out of the loop RSI progresses through three stages: 🔹 Human in the loop: AI proposes changes; people approve them. 🔹 Human on the loop: data, rewards, and verifiers are automated; people supervise deployment. 🔹 Closed loop: the system generates, verifies, and applies improvements itself. Most “self-evolving” systems remain in the second stage. 2️⃣ What actually improves? Self-refine changes the current answer. Test-time training writes experience into weights. Agent evolution modifies prompts, tools, memory, skills, workflows, or Agent code across tasks. output → weights → the Agent itself Training-time RSI follows another ladder: 🔹 Zero-label: AI generates supervision. 🔹 Zero-data: AI also generates problems and curricula. 🔹 Auto research: AI chooses hypotheses, training recipes, and experiments. The system gradually takes over how to learn, what to learn, and finally how to improve learning itself. 3️⃣ Self-improvement can amplify mistakes A generator and verifier may share the same biases. Wrong outputs can produce biased evaluations, biased learning signals, and stronger errors. Even correct rewards do not guarantee stability: training can improve and later collapse. Self-play may also lose diversity or favor problems that are easy to reward rather than genuinely useful. Automation makes grounding more important, not less. 4️⃣ Verification is the real bottleneck Math and code are RSI-friendly because proofs, unit tests, and execution feedback provide clear signals. Open-ended Agent work is harder. A verifier must judge not only correctness, but novelty, usefulness, importance, and research taste. The next step may be evolving the verifier itself. But if the policy and evaluator change together, what keeps both aligned with reality? ✅ The boundary of true RSI The loop is expanding: answer → experience → learning signal → problem and curriculum → verifier True RSI requires improvement across every layer without losing contact with real objectives. Until evolving verifiers remain reliable without constant human grounding, recursive self-improvement is still an aspiration rather than an achieved capability. 🔗 Full analysis: #RecursiveSelfImprovement# #RSI# #SelfEvolvingAI# #AIAgents# #ReinforcementLearning# #AISafety#
Show more