Register and share your invite link to earn from video plays and referrals.

Search results for SelfImprovement
SelfImprovement community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SelfImprovement
What if letting an agent evolve its own workflow just meant it got better at gaming the test? This paper tackles exactly that overfitting problem. Title: RRSI: Regularized Recursive Self-Improvement of Agent Harnesses URL: ❓ What's a harness? It's everything wrapped around the frozen LLM: prompts, control flow, tooling, memory, context management. The same model can perform very differently depending on this design. ❓ Why does self-improving it overfit? Three failure modes show up: fitting too tightly to the evolve-set benchmark, chasing noise, and letting complexity pile up unchecked. 💡 How does RRSI fix it? It applies classic ML regularization (L0, L1, L2) to harness evolution: an annealed edit budget limits changes per round, benchmark-specific proposals get screened out at selection, and unproductive components get pruned. 💡 What's the payoff? Across 8 benchmarks, RRSI beats baselines by up to 22.9% on unseen tasks, while using 30% fewer tokens. It feels like a genuinely grounded step toward agents that can safely improve their own workflow in production. #AIAgents# #SelfImprovement#
Show more
Every interaction your deployed AI agent handles could double as training data to make it smarter. That is the pitch of this paper. NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness 🧩 Overview The authors turn logs from a "routing harness" (the layer that decides which capability tier handles each request) directly into training data, with no extra synthetic data pipeline, to recursively improve the model. ❓ The problem it solves Recursive self-improvement needs a way for a system to observe its own capabilities and feed that back into training. Most approaches build a separate data pipeline for this, but the material is already sitting inside everyday routing logs. ⚙️ Method Execution trajectories are structured by user turn, quality-checked with six semantic dimensions plus structural validation, and routing scores are used to order a three-stage curriculum for both SFT and on-policy distillation. Evaluation results then steer the next batch of training data toward weak spots. 📊 Results Across 10 agentic, coding, and instruction-following benchmarks, the 4B model's macro-average rose from 58.94 to 64.87, and the 9B model from 65.60 to 69.04, beating public synthetic data by +6.26 macro-average points. #AIAgents# #SelfImprovement#
Show more
Self-improving agents are a top research topic right now. This new survey is a good map of the area. (bookmark it) It splits recursive self-improvement into stages of autonomy. An agent first executes improvements someone else designed. Then it chooses its own improvement strategy, collects its own experience, adapts to new environments, and finally improves the process of improvement itself. That staging makes claims easier to check. When a paper says its agent is self-improving, you can ask which of these stages it actually automates. The survey also uses a Headroom-Closed Index to show where current LLMs fall short, and compares requirements across scientific discovery, embodied agents and software engineering. Paper:
Show more
0
33
799
134
Forward to community
RECURSIVE SELF IMPROVEMENT
0
86
1.6K
172
Forward to community
Most “Self-Evolving AI” Is Not Recursive Self-Improvement Yet AI may generate its own data, rewards, skills, and code. But if humans still define what counts as better and approve deployment, the loop is not closed. Zhihu contributor 一口鸟 (@alsaceym) argues that RSI’s hardest bottleneck is reliable verification. 1️⃣ Moving humans out of the loop RSI progresses through three stages: 🔹 Human in the loop: AI proposes changes; people approve them. 🔹 Human on the loop: data, rewards, and verifiers are automated; people supervise deployment. 🔹 Closed loop: the system generates, verifies, and applies improvements itself. Most “self-evolving” systems remain in the second stage. 2️⃣ What actually improves? Self-refine changes the current answer. Test-time training writes experience into weights. Agent evolution modifies prompts, tools, memory, skills, workflows, or Agent code across tasks. output → weights → the Agent itself Training-time RSI follows another ladder: 🔹 Zero-label: AI generates supervision. 🔹 Zero-data: AI also generates problems and curricula. 🔹 Auto research: AI chooses hypotheses, training recipes, and experiments. The system gradually takes over how to learn, what to learn, and finally how to improve learning itself. 3️⃣ Self-improvement can amplify mistakes A generator and verifier may share the same biases. Wrong outputs can produce biased evaluations, biased learning signals, and stronger errors. Even correct rewards do not guarantee stability: training can improve and later collapse. Self-play may also lose diversity or favor problems that are easy to reward rather than genuinely useful. Automation makes grounding more important, not less. 4️⃣ Verification is the real bottleneck Math and code are RSI-friendly because proofs, unit tests, and execution feedback provide clear signals. Open-ended Agent work is harder. A verifier must judge not only correctness, but novelty, usefulness, importance, and research taste. The next step may be evolving the verifier itself. But if the policy and evaluator change together, what keeps both aligned with reality? ✅ The boundary of true RSI The loop is expanding: answer → experience → learning signal → problem and curriculum → verifier True RSI requires improvement across every layer without losing contact with real objectives. Until evolving verifiers remain reliable without constant human grounding, recursive self-improvement is still an aspiration rather than an achieved capability. 🔗 Full analysis: #RecursiveSelfImprovement# #RSI# #SelfEvolvingAI# #AIAgents# #ReinforcementLearning# #AISafety#
Show more
Recursive self-improvement may not begin with models training models. It can also begin in the infrastructure those models depend on. This AI on Air clip is excerpted from a recent episode on the @MatthewBerman channel. The main speaker is Thibault Sottiaux @thsottiaux, Head of Core Product and Platform at @OpenAI. ▷ Thibault says models can already help improve the inference stack, hardware, CUDA kernels, and more efficient product interactions. Each sits on the critical path for getting value from a model. ▷ Capabilities such as cloud agents can first increase the utility people receive from models, then direct that added capability back toward improving the system. That is also a form of recursive improvement.
Show more
Recursive Self Improvement is a possibility worth taking seriously. Our new paper adds empirical evidence to the understanding of current capabilities, avoiding both extremes — hype based on naive trend extrapolation and knee-jerk dismissal based on hand-wavy arguments.
Show more
Embrace self-improvement and breakthroughs through every morning run. Instead of chasing speed, focus on consistent progress, gradually surpassing your past self and growing steadily. 🔌Milk with honey, vanilla, and nuts
Show more