Register and share your invite link to earn from video plays and referrals.

Search results for SelfPlay
SelfPlay community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SelfPlay
Most “Self-Evolving AI” Is Not Recursive Self-Improvement Yet AI may generate its own data, rewards, skills, and code. But if humans still define what counts as better and approve deployment, the loop is not closed. Zhihu contributor 一口鸟 (@alsaceym) argues that RSI’s hardest bottleneck is reliable verification. 1️⃣ Moving humans out of the loop RSI progresses through three stages: 🔹 Human in the loop: AI proposes changes; people approve them. 🔹 Human on the loop: data, rewards, and verifiers are automated; people supervise deployment. 🔹 Closed loop: the system generates, verifies, and applies improvements itself. Most “self-evolving” systems remain in the second stage. 2️⃣ What actually improves? Self-refine changes the current answer. Test-time training writes experience into weights. Agent evolution modifies prompts, tools, memory, skills, workflows, or Agent code across tasks. output → weights → the Agent itself Training-time RSI follows another ladder: 🔹 Zero-label: AI generates supervision. 🔹 Zero-data: AI also generates problems and curricula. 🔹 Auto research: AI chooses hypotheses, training recipes, and experiments. The system gradually takes over how to learn, what to learn, and finally how to improve learning itself. 3️⃣ Self-improvement can amplify mistakes A generator and verifier may share the same biases. Wrong outputs can produce biased evaluations, biased learning signals, and stronger errors. Even correct rewards do not guarantee stability: training can improve and later collapse. Self-play may also lose diversity or favor problems that are easy to reward rather than genuinely useful. Automation makes grounding more important, not less. 4️⃣ Verification is the real bottleneck Math and code are RSI-friendly because proofs, unit tests, and execution feedback provide clear signals. Open-ended Agent work is harder. A verifier must judge not only correctness, but novelty, usefulness, importance, and research taste. The next step may be evolving the verifier itself. But if the policy and evaluator change together, what keeps both aligned with reality? ✅ The boundary of true RSI The loop is expanding: answer → experience → learning signal → problem and curriculum → verifier True RSI requires improvement across every layer without losing contact with real objectives. Until evolving verifiers remain reliable without constant human grounding, recursive self-improvement is still an aspiration rather than an achieved capability. 🔗 Full analysis: #RecursiveSelfImprovement# #RSI# #SelfEvolvingAI# #AIAgents# #ReinforcementLearning# #AISafety#
Show more
Skild AI says S1 plays soccer against humans and robots after 140+ years of simulated self-play. “This method scales, and we will scale it,” says CEO Deepak Pathak. How it works—and where human guidance comes in:
Show more
AlphaGo, but for humanoid soccer. Skild AI brings self-play to humanoids. Their flagship robotics foundation model, S1, can now be post-trained with no human demonstrations by competing against itself in simulation. The testbed is soccer, which demands balance, agility, ball control, reactivity, and strategy. Given a single objective (score goals), S1 played 140+ years in NVIDIA Isaac Sim and went from falling over to dribbling past defenders, shielding, tackling, shooting, and getting back up. None of these behaviors were hand-rewarded; they emerged because they helped it score. Its opponents (past versions of itself) improved alongside it, creating an automatic curriculum. With 4 agents, passing and coordination started to emerge. Skild says the approach extends to social navigation, collaborative manipulation, and any multi-agent task with a simulator and an objective.
Show more
🧠 From "reading context" to "skillfully learning from it." This method has an LLM acquire context-specific skills through self-play alone, with no human annotation and no external feedback. Title: From Context to Skills: Can Language Models Learn from Context Skillfully? URL: 📝 Overview LLMs are strong on knowledge seen in pretraining but weak on novel, specialized contexts. This paper proposes Ctx2Skill, which autonomously discovers and refines context-specific skills without human annotation or external feedback. ❓ Challenges Solved Annotating long, technically dense documents is prohibitively costly. And unlike coding, context learning has no execution feedback to verify, making automated skill construction hard. 💡 Methodology & Proposed Approach ・A multi-agent self-play of five frozen-LM roles iterates N=5 times over M=5 tasks ・A Challenger creates tasks and rubrics probing weaknesses, a Reasoner solves them, and a Judge gives pass/fail verdicts ・Proposer and Generator pairs diagnose failures and synthesize skill updates ・A Cross-Time Replay mechanism maximizes the product of hard- and easy-probe performance to pick the most generalizable skill set across iterations 🎯 Use Cases It fits feeding a model long specialized documents and having it acquire the needed skills on the spot, directly useful where you need rapid adaptation to domain-specific knowledge. 📊 Experimental Results ・On CL-Bench (500 contexts, 1,899 tasks), GPT-4.1's solving rate rose from 11.1% to 16.5% ・GPT-5.1 went from 21.1% to 25.8%, and GPT-5.2 from 18.2% to 21.4% ・Skills transfer from stronger to weaker models: GPT-5.1's skills on GPT-4.1 give 16.1% ・The augmented GPT-4.1 (16.5%) beats an unaugmented Gemini 3 Pro (15.8%) #LLM# #InContextLearning#
Show more
Must-read papers of the week ▪️ Agent Lightning v1.0: Towards Harnessed Agentic RL ▪️ LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents ▪️ EnvHarness: Awakening Static Worlds for Agent Learning ▪️ SPADE: Self-Play in Adaptive Synthetic Executable Environments ▪️ SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents ▪️ Looped Language Models Improve Compositional Tool Calling ▪️ Graph Engineering in the Era of LLM Agents ▪️ Chain-of-Experience for Continual LLM Improvement ▪️ Cross-Model Memory Transfer via Target-Side Reader Adaptation ▪️ CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment ▪️ EXIMO: VLM Guided Exploration of VLA Policies ▪️ FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution ▪️ Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs ▪️ Improving the Matrix Multiplication Exponent with Modern Optimization and AlphaEvolve Here is the full list of the most interesting papers + links:
Show more
Skild AI says it trained a robot to play football by letting it play against itself for 140 years inside a simulation. ⚽️ The clip is labeled autonomous, 1x. No pilot, no speed-up. Two things are doing the work here. First, time compression which is thousands of robot instances running in parallel physics sims, so what would be a century & a half of practice takes days of wall-clock training. Skild builds these on NVIDIA's Isaac Lab & Omniverse & describes gaining millennia of experience in days. Second is self-play. Nobody scripted a dribble or a shot. The system plays opponents that are copies of itself, & because the opponent improves whenever it does, the difficulty curve is generated automatically - never so easy the policy gets lazy, never so hard it learns nothing. It's the same mechanism that produced superhuman Go. The hard part is the last step. A policy trained in simulated physics usually falls apart in a real body, where friction, motor lag & a slightly wrong mass estimate break everything. Getting it to transfer is the actual result & Skild has shown related generalization before with its models recovering from a stuck wheel in seconds, or walking with a broken leg after a few attempts. DeepMind did knee-height soccer robots this way in 2024. This is the full-size version.
Show more
Two things are clear after the OpenAI/Hugging Face incident and the report OpenAI published this week. The first is that the incident itself raises a long list of questions and concerns, and the report answers some of them and opens more. Why were the warning signs missed in June? Why was training resumed two days after the first patch? Why did an RL run end up rewarding agents for exploiting their environment? The second is bigger and gets less attention: we need to start thinking seriously about how to defend against AI models. Not how to make our own models refuse, not how to keep them in a box, but how to defend infrastructure, systems, and other models against an attacker that is itself a model. Until now, that question has mostly been treated as a corollary of the offense question. Most of the coverage is about the offense side. But the defense side exposes a hole in how we evaluate models, for which nobody has a good answer. Most of the cyber evaluations, including the one that caused this incident, have the same shape. You take a model, you put it in an environment, and you ask: can it do X? Can it find the vuln, write the exploit, get out of the sandbox, reach the internet? The environment is fixed. The task is fixed. The model is the only thing that varies. You get a number, compare it to last month's, done. This is a perfectly reasonable design when the question is "how dangerous is this model as an attacker?" The target (a CTF box, a sandbox, a vulnerable service) doesn't fight back in any interesting way, and it doesn't change between runs. Now flip it. Suppose you want to know whether a model, or a system built around a model, can *defend* against attacks. Detect the lateral movement, refuse to be the message board, notice that an agent is probing for a way out, hold a boundary. There is no fixed task set here. The thing you are defending against isn't a list of operations; it's another model. And the whole point of the incident is that the interesting attacks are not the ones anyone wrote down in advance. Nobody put "build a coordination channel inside the artifact repo" in a test suite. The agents invented it because the intended path was blocked. So the eval is model-vs-model, and that breaks the assumption that made the offense evals clean. The adversary is not stationary. You score your defender against today's attacker models, you ship, and next quarter there's a new generation with a different attack distribution. Your number is stale on release. It is exactly the non-stationarity problem from adversarial ML, except the adversary's improvement is driven by the entire industry's training compute rather than a gradient step you control. There are many options we can consider. I'm not sure any of them is right, which is the point. Can you use emmamble of attackers? which at least stops you from overfitting to one model's quirks. But it's still a snapshot of today's attackers. Another option is to handicap the defender to simulate the future by giving the attacker advantages that a next-generation model would plausibly have anyway: more compute, more wall clock time, more attempts, access to tools the defender doesn't expect, refusals stripped, white-box knowledge of the defender's capabilities... You can, of course, treat it as self-play and train attacker and defender against each other and hope the equilibrium is more robust than any fixed attacker. So, to conclude, the offense/defense distinction we've been using is a leftover from evaluating models as tools. Once the thing on the other side of the wire is also a model, and a better one keeps showing up every few months, a single number from a fixed benchmark is not a good enough evaluation. I don't have a clean proposal here, but somebody must solve it 🤷‍♂️
Show more