Register and share your invite link to earn from video plays and referrals.

Search results for LLMAgents
LLMAgents community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including LLMAgents
🔄 TL;DR: A search index that diagnoses its own weaknesses, rewrites its keys, and validates the changes — all without any human in the loop. It beats existing methods by a wide margin on BRIGHT. Title: Self-Evolving Search Index URL: Points 🩺 Builds co-retrieval profiles from search results to self-diagnose whether documents are well distinguished, with no human annotation needed ✍️ Selectively revises key sets only for flagged documents, autonomously deciding up to 10 keys per document ✅ Self-validates proposals on faithfulness, specificity, and separation, keeping only keys that pass 🔍 A Query Simulator proactively probes uncovered demand with synthetic queries 📈 Hits 22.8 average nDCG@10 on BRIGHT, +9.2% over RL-Index and +40-57% over the base index 🤖 For search agents, answer accuracy jumps +77.87% while search calls drop -16.80% The idea of making index optimization itself self-evolving, not just the model or the agent, is what makes this interesting. #InformationRetrieval# #LLMAgents#
Show more
A/B tests take weeks to run. What if real-data-driven AI personas could predict the outcome in advance? Title: Data-Driven Persona-Conditioned Agents for A/B Test Simulation URL: ❓ What question format works best for AI personas? 💡 Pairwise evaluation, where the agent directly compares variants, wins clearly: 0.75 accuracy on CTR and 0.80 on subscription tests. Independent scoring drops to around 0.40. ❓ Do you need proprietary data to build good personas? 💡 Open e-commerce data matched proprietary data almost as well. But out-of-domain data (movie reviews) hurt accuracy, showing domain match really matters. ❓ Is it better to go deep on behavior data or wide on demographic diversity? 💡 A "deep pool" of highly engaged users won for CTR prediction, but for subscription tests, a demographically diverse pool performed just as well. ❓ Can you cut costs by using fewer personas? 💡 Yes — subsampling from 935 down to 500 personas kept accuracy nearly unchanged, roughly halving simulation cost. #ABTesting# #LLMAgents#
Show more
🧭 Can an AI model really teach itself to improve when it's only handed a vague goal, with no task spec and no reward? A team from ByteDance Seed and collaborators built a benchmark to find out. Title: Aspire: Can Models Self-Evolve from Vague Goals? URL: The benchmark hands the model nothing but a vague goal, lets it handle interpretation, training, and verification end-to-end, and then measures real capability gains against a hidden evaluation set. 🎯 Highlight 1: The cost of ambiguity Simply rephrasing a task as vague dropped Claude Opus 4.8's score from 32.90% to 27.07% and GPT-5.6's from 36.23% to 29.58%. Agents burn time figuring out what to optimize for, leaving less time for actual training. 🔁 Highlight 2: Execution isn't the same as improvement In self-evolution runs with Qwen3.5-4B/9B, most final checkpoints underperformed the base model, and 21 of 24 runs were rolled back to their starting state by the safety mechanism. Training against narrow self-evaluations produces local gains that fail to transfer to hidden evaluation. 🛠️ Highlight 3: Harness evolution can't beat the baseline either Even when agents were allowed to rebuild their own execution harness, the best successor harness scored 27.22 versus 28.64 for the reference implementation — falling short due to narrow verification and gaps in output completeness checks. 💡 This research makes clear that deciding what to improve is the next big wall standing between agents and sustained self-improvement. #LLMAgents# #SelfEvolution#
Show more
Training data for terminal agents often ships tasks where instruction, environment, solution, and verifier disagree — producing unsolvable tasks. A new synthesis framework cuts that at the root. Title: FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis URL: 📌 Overview Reconstructs related skills into rich scenarios, builds the environment first, and grounds instruction, solution, and verifier in that same executable state. 🧩 Problems it solves ・Multi-stage generation loses source dependencies and intermediate states ・Misaligned instruction/environment/solution/verifier yield unsolvable, unverifiable tasks ⚙️ Method ・Collects 71K+ skills, reconstructs into 5-dimensional scenarios ・Builds the environment in Docker, exposing its real state as a shared channel ・Generates instruction→solution→verifier sequentially; a router pinpoints and repairs only the failing part 📊 Results ・Synthesizes 6,078 validated tasks at 22.77 tests/task on average ・Fine-tuning Qwen3.5: 4B +40.5%, 9B +30.1%, 27B +16.5% ・The 27B (47.57) nears the ~15x larger 397B (49.06) ・70% end-to-end yield vs 15–28% for competitors A clear case that high-quality executable tasks come from careful state grounding, not brute-force generation. #LLMAgents# #TerminalBench#
Show more
The hardest memory failure may be an agent that remembers enough to act but not enough to explain or undo the action. This survey reframes always-on agents as persistent-state systems. Their durable state includes facts and preferences, but also permissions, credentials, task ledgers, triggers, provenance, and external commitments. The key question is whether that state remains authorized, scoped, traceable, mutable, recoverable, and safe to act on. Across a scoped corpus of 435 works, retrieval appears in 269 and writing in 200, while only 27 expose any rollback mechanism; authority appears in 72. The field has built the forward arc of writing, organizing, retrieving, and acting far more thoroughly than the return arc of forgetting, auditing, and recovery. – arxiv. org/abs/2606.30306 Title: "Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents"
Show more
Decouple "search" from "reasoning" in LLM agents and you can cut search costs by up to 98% while keeping accuracy nearly intact 🔌 Title: Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents URL: 🔌 Overview DSG separates search-based grounding from the language model's reasoning. It runs as an independent gateway compatible with the Model Context Protocol (MCP), acting as a vendor-agnostic intermediary layer. ❓ Challenges Solved In production LLM agents, real-time search grounding is tightly coupled to the model provider. ・This makes systems hard to inspect, reconfigure, repurpose, or migrate ・Search can cause "Search-Induced Verbosity" that violates strict output requirements Bundling search with reasoning was a bottleneck for both flexibility and cost. 💡 Methodology & Proposed Approach It places grounding at the interface between search and generation, not inside the model, exposing previously model-embedded elements as controllable first-class features. ・Provider routing (choose and switch search providers) ・Source-aware context rendering ・Configurable fallback mechanisms ・Retrieval-depth management ・Both exact and semantic caching 📊 Experimental Results ・SimpleQA: 86.1% accuracy (vs 87.7% native search) while cutting search costs by 91% ・99.4% warm-cache hit rate with 68% latency reduction ・Production e-commerce: matched native-search accuracy while cutting search costs by over 98% ・But native search kept an edge on recency-sensitive FreshQA queries #LLMAgents# #Search#
Show more
Give children and LLMs the exact same mystery-solving task — how does their reasoning differ? 🧒 A study that puts human and AI inference side by side, fairly. Title: Hypothesis Generation and Inductive Inference in Children and Language Models URL: 🧒 Overview This study has both children and LLM agents solve a task of inferring hidden causes under uncertainty, then carefully compares them. It examines how closely humans and AI align — and where they diverge — in generating hypotheses and reasoning inductively. ❓ Challenges Solved Humans, especially children, build mental models quickly from sparse cues. ・It was unclear whether the computational principles behind human reasoning under uncertainty also appear in LLMs placed under matched constraints ・There wasn't even a fair framework for putting children and AI side by side This work takes that question head-on. 💡 Methodology & Proposed Approach The researchers designed an inductive-inference "Box Task" for inferring hidden causes. ・Sequential environment interaction: discover latent causes by acting on the environment ・Modeled with Bayesian particle-based inference ・Systematic manipulation of evidence reliability and observability ・Measures both task completion and rule generalization Analysis uses two complementary frameworks: constraint satisfaction over hypotheses and program synthesis evaluation. 🌍 Use Cases / Experimental Results The similarities and differences between humans and AI came through sharply. ・Both groups discounted unreliable evidence and sought more information to partially resolve uncertainty ・Both showed a dissociation between task completion and causal generalization (solving a task doesn't guarantee generalizing the rule) ・LLM agents over-observe and over-comply with instructions relative to children ・Despite similar environmental adaptation, they had distinct information-seeking costs and inductive biases This offers insight into cognition and a guide to where LLM agents differ from humans by design. #CognitiveScience# #LLMAgents#
Show more
🔎 LLM agents rewrite a decompiler's unreadable `local_48`-laden code to be readable while preserving function, but a single metric collapses into "gaming." The fix is a multidimensional readability score. Title: LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics URL: 📝 Overview This paper has LLM agents improve the readability of decompiled binaries while keeping functional correctness. The key is QRS, a multidimensional score combining structural validation with three readability sub-metrics. ❓ Challenges Solved Automated decompilers produce functionally correct but unreadable code. When LLMs try to fix it, without quantitative guidance they lose focus, and optimizing a single metric leads to "gaming" that sacrifices other dimensions. 💡 Methodology & Proposed Approach ・QRS is a structural gate times a composite score, a weighted sum of lexical surprisal, structural simplicity, and idiomatic quality ・Lexical surprisal uses a small code-LLM's perplexity to measure how familiar the code looks ・Structural simplicity uses cyclomatic complexity and nesting depth; idiomatic quality uses clang-tidy anti-pattern checks ・QRS is computed only if the recompiled code reaches at least 0.85 CFG similarity to the original binary in radare2 🎯 Use Cases It directly speeds up reading decompiler output in malware analysis, vulnerability research, legacy-software comprehension, and patch diffing. 📊 Experimental Results ・On 210 synthetic C binaries, LLM-only reached QRS at least 0.75 in 74.76% of cases, with QRS up +0.420 on average and zero regressions ・Allowing Bash execution raised the rate to 82%, improved QRS by +0.509, and cut iterations from 5.92 to 2.933 (a 43% reduction) ・It empirically shows that going multidimensional avoids Goodhart's Law, "when a measure becomes a target, it stops being a good measure" #ReverseEngineering# #AIAgents#
Show more
We benchmarked GPT-6 models in DOOM by having LLM agents fight each other. Across 120 matches, we found: - Astra had the highest win rate at 82.5% - Sol made the fastest decisions at avg of 5.34s - Luna had the most wins per dollar at 12.8 GitHub: @OpenAIDevs @rootlyhq
Show more
Finally a good paper testing if file-system based memory for LLM agents is worth it. First, what does this look like? Deployed agents keep long-term memory as a folder of markdown files they read and reorganize with ordinary file tools. Two assumptions had never been checked. That an agent can keep a growing store organized as memories accumulate, conflict, and go stale. And whether the organization pays for itself. Organized stores roughly halve retrieval cost when the material is large. No agent in the study converted organization into better answers, and in the growth study the store degraded for every management agent except the strongest one. Changing the tool set alone reshapes the memory store as strongly as swapping the model. Paper: Track more trending AI papers in our academy:
Show more