Register and share your invite link to earn from video plays and referrals.

Search results for SelfEvolution
SelfEvolution community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SelfEvolution
🧭 Can an AI model really teach itself to improve when it's only handed a vague goal, with no task spec and no reward? A team from ByteDance Seed and collaborators built a benchmark to find out. Title: Aspire: Can Models Self-Evolve from Vague Goals? URL: The benchmark hands the model nothing but a vague goal, lets it handle interpretation, training, and verification end-to-end, and then measures real capability gains against a hidden evaluation set. 🎯 Highlight 1: The cost of ambiguity Simply rephrasing a task as vague dropped Claude Opus 4.8's score from 32.90% to 27.07% and GPT-5.6's from 36.23% to 29.58%. Agents burn time figuring out what to optimize for, leaving less time for actual training. 🔁 Highlight 2: Execution isn't the same as improvement In self-evolution runs with Qwen3.5-4B/9B, most final checkpoints underperformed the base model, and 21 of 24 runs were rolled back to their starting state by the safety mechanism. Training against narrow self-evaluations produces local gains that fail to transfer to hidden evaluation. 🛠️ Highlight 3: Harness evolution can't beat the baseline either Even when agents were allowed to rebuild their own execution harness, the best successor harness scored 27.22 versus 28.64 for the reference implementation — falling short due to narrow verification and gaps in output completeness checks. 💡 This research makes clear that deciding what to improve is the next big wall standing between agents and sustained self-improvement. #LLMAgents# #SelfEvolution#
Show more
🧬 For agents to evolve themselves safely, the protocol itself needs a redesign — this work fills the pieces missing from MCP and A2A with a self-evolution protocol. Title: Autogenesis: A Self-Evolving Agent Protocol URL: 🧬 Overview The Autogenesis Protocol (AGP) lets LLM agents dynamically improve their own prompts, tools, and policies during execution. Its central idea is to decouple "what evolves" from "how evolution occurs." ❓ Challenges Solved Existing protocols (Anthropic's MCP, Google's A2A) standardize connectivity and invocation but lack the primitives evolution needs. ・Resources stay tightly coupled to agent code ・No version control or rollback for evolutionary steps ・Heuristic edits, with no standardized operators forming a rigorous control loop 💡 Methodology & Approach AGP has two layers. ・RSPL: models prompts, agents, tools, environments, and memory as "registered resources" with explicit state, lifecycle, and versioned interfaces ・SEPL: formalizes evolution as typed, composable operators, routing every change through RSPL so it's versioned and reversible On top, the AGS system runs sub-agents concurrently on an Agent Bus and, when traces signal failures, self-improves via a Reflect → Select → Improve → Evaluate → Commit loop. 🎯 Use Cases It fits building agents that need long-horizon planning and diverse tool use. Because improvements keep an auditable lineage and can be rolled back, you can run self-evolution safely while avoiding brittle glue code. 📊 Results ・Science & math: weaker models gain most. gpt-4o improves +100% on AIME25, gpt-4.1 +71.38% on AIME24, while saturated strong models see small gains (ceiling effects) ・GAIA: Test 79.07% → 89.04% (+12.61%), with +33.34% on the hardest Level 3 ・Code generation: C++ pass rate 79 → 99, time-limit errors 9 → 0, runtime efficiency +46.4%. Jointly evolving prompts and outputs worked best #AIAgents# #SelfEvolution#
Show more
Another interesting approach to self-evolve agent skills. But it's important to know that skill self-evolution loops fail in two specific ways: 1. Direction instability. Effective corrections get overwritten by iteration-local feedback instead of accumulating, so the loop keeps undoing its own fixes. 2. Fixed update scope. Every revision changes about the same amount regardless of whether recent case-level improvements were consistent or noisy. SkillAdam addresses both by porting Adam's two moment estimates to discrete, non-differentiable skill documents. As a functional analogue of the first moment, an optimization memory records identified problems and the outcomes of prior solution attempts, which stabilizes the update direction. As an analogue of the second moment, a volatility-driven edit budget tracks the history-weighted variation of recent case-level improvements and controls how large each revision is allowed to be. Across seven benchmarks spanning short and long-horizon tasks it reaches state of the art with more stable optimization dynamics, and it gets there in substantially fewer iterations and at lower cost than prior methods. Paper:
Show more
Nice paper showing a better way to evolve agent skills. And they achieve 40–70% less token cost compared to frontier evolving methods. The idea is to let agents improve their skill prompts by ranking candidates with a learned rubric instead of running a full rollout to score every revision. Rollout cost is the reason skill self-evolution usually only patches observed failures. Every candidate edit needs a real agent run to evaluate. SkillLift trains a rubric to agree with real outcomes on which of two skills is better, since ranking needs fewer oracle runs than predicting each score. An inner loop revises skills against the frozen rubric at no rollout cost. An outer loop spends a few real rollouts to re-align the rubric by rank correlation. On SkillsBench and WildClawBench (147 tasks) with three models, it beats SkillOpt and CoEvoSkills in all six combinations, even when those baselines get twice the token budget. It reaches target performance with 40 to 70% fewer tokens. Paper:
Show more
Yes, our latest special guest is Fuli Luo @_LuoFuli . The second battle in the global large model arms race has begun: shifting from the Chat era dominated by pre-training to the Agent era driven by post-training. This marks Fuli Luo’s first-ever interview, as well as her first in-depth technical conversation. We talked systematically about the massive AI upheaval triggered by technological breakthroughs including Claude Opus 4.6 and OpenClaw in 2026, along with its subsequent structural impacts across the industry. Amid the fierce large-model arms race, the world around us is undergoing brutally rapid changes—even for researchers who train models firsthand. “I used to believe our work was highly creative, and could never be simplified into fixed skills or standardized workflows. But now I realize it can be automated after all. If that’s possible, can models train stronger models on their own? Can they achieve iterative improvement through self-evolution? This is exactly what will unfold in the next couple of years,” Fuli Luo says. As human knowledge and wisdom are internalized into model capabilities, what will humanity pursue in the future? Is our society truly ready for this tsunami-scale technological revolution? All in all, this is an information-dense dialogue. It reveals how an AI lab makes strategic technical bets, allocates resources, and adjusts organizational structure and team planning amid a major paradigm shift. At the core of its response to drastic change lies its established culture and core values. Though lengthy and technically intensive, we hope this conversation brings great insights to every viewer. Our podcast, video episode and article are released simultaneously across platforms, with English subtitles provided to assist non-Chinese-speaking audiences. Luo Fuli: OpenClaw, Agent Frameworks — The AI Paradigm Has Already Chang... 来自 @YouTube
Show more
🌐 The key to building strong AI agents may actually be designing the environments they operate in. This 63-page survey systematizes the view of "environment engineering." Title: Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application URL: 📝 Overview LLM agents don't act alone; they operate inside interactive environments. This survey organizes the research landscape through the lens of "environment engineering," the engineering design and construction of those environments themselves. ❓ Challenges Solved Until now, how to build environments was discussed only in fragments. Even though agent capability depends heavily on good environment design, there was no unified framework to organize it. 💡 Methodology & Proposed Approach It classifies environments along the development lifecycle in four pillars. ・Environment modeling: characterizing representative environments and assessing core capabilities ・Environment synthesis: two paradigms, symbolic and neural ・Environment evaluation: domain-specific assessment aligned with the synthesis paradigms ・Environment application: agent-environment co-evolution across four pathways, memory-centric, orchestration-centric, trajectory-centric, and exploration-centric 🎯 Use Cases It helps agent researchers locate their own work on a map and spot missing perspectives, and serves as a starting point when designing environment synthesis, evaluation, and self-evolution. 📊 Trends and Outlook ・It organizes evolution approaches into three families: neural-driven, difficulty-driven, and scaling-driven ・It analyzes across eight attributes and eight application domains ・It points to Environment-as-a-Service, multi-agent systems, and neural-symbolic integration as future directions #AIAgents# #LLM#
Show more
Recent thoughts: The Shift to Long-Horizon Tasks The most likely breakthrough this year will be in long-horizon tasks. We are moving toward a stage where Large Language Models (LLMs) learn to complete extended, complex missions by interacting with Agent environments. This is perhaps where the true value of LLMs lies. Take cybersecurity as an example: imagine a model that continuously hunts for software bugs and vulnerabilities. While it sounds like a search process, it’s actually the model learning the high-level intuition and methodology of a professional hacker. Unlike humans, AI can run 24/7 without fatigue. It could potentially find exploits at a much higher frequwill ency and claim bounties on platforms like HackerOne or BugCrowd. It sounds fun, but fundamentally, it's a revolution that displaces the hacker. If even hackers are being "disrupted," one can only imagine the impact on general programmers. From One-Person to None-Person Companies Building on long-horizon capabilities, Autonomous Agent Systems (AAS) will inevitably become the next frontier. Last year, we were discussing the rise of the "One Person Company" (OPC). I didn't expect us to move so quickly toward the "None Person Company" (NPC). It’s an ironic twist—we might all end up as NPCs in this new ecosystem. Engineering the Impossible: Memory and Learning To realize the vision above, we must solve three technical pillars: Memory, Continual Learning, and Self-Judging. I used to think these would require massive paradigm shifts and years of research. However, the pressure from both the technical and application sides is so intense that we are seeing these capabilities emerge through ingenious engineering "tricks": Memory: Long context windows (1M+) and RAG have significantly bridged the gap. Continual Learning: While true continual learning remains difficult, the release cycles are shrinking. Global models are updated monthly; domestic models are catching up. If we reach weekly updates by next year, it will effectively function as continual learning. Self-Judging: This remains the most elusive, yet models like Opus 4.7 are already demonstrating early self-correction and judgment capabilities. The Self-Evolving Endgame The most difficult—and most promising—path is Self-Evolution. The current wave is incredibly fierce. I suspect that models like Claude may have already achieved a baseline for self-training: writing their own code, cleaning their own data, generating synthetic data, and then training on it. It might "waste" some compute, but it saves the most precious resources: human labor and time. In the LLM era, speed is everything. Rapid iteration is what creates the cognitive gap between leaders and followers. Claude’s rumored 2-million-chip cluster for next year is likely dedicated to exactly this: autonomous model self-training. Technical Summary: 1M Context: Necessary baseline. Memory & Continual Learning: Prerequisites, likely solved first via "tricky" engineering. Harnessing Environments: The breakthrough point. Self-Judging: The tipping point. Full Self-Training: The endgame. Redefining AGI and the Industry If this is the road to AGI, then AGI’s definition should be the sum of all human collective intelligence, not just an individual’s intelligence. It must possess the creative capacity to produce something as profound as the "Theory of Relativity"—meeting the bar set by Hassabis. During this transition, every APP will need to be reconstructed as AI-native. In fact, we might move past the concept of APPs entirely. The most significant challenge will be the reconstruction of the operating system itself. In the future, you won’t see a traditional desktop; you will see an LLM OS, where applications are "generated on demand." This challenges the 80-year-old Von Neumann architecture and represents a total upheaval of the computer science industry. The Irreversible Wave From completing long-horizon tasks to fully autonomous operations, every sector—Security, Finance, Law, E-commerce—will be reshaped. Many friends have reached out lately, asking how to transform their enterprises to keep pace with AI. But few truly realize that this irreversible process has already begun. As this massive technical wave hits, we must be prepared to act, but we must also start thinking seriously about how to regulate it.
Show more
0
39
768
149
Forward to community