Register and share your invite link to earn from video plays and referrals.

Search results for TLをギュステにしよう
TLをギュステにしよう community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including TLをギュステにしよう
TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability. Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability URL: Points 📉 Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens 🧮 Four task families — arithmetic, UUID sorting, variable lookup, CSV table transformation — measure sustained execution 📄 Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens 🔀 Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3% 🔁 Accuracy consistently declines across every model as local task complexity increases 🤔 Models stay highly self-consistent yet get the wrong answer — they understand the task but can't read the right problem out of the context 🛠️ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution. #AIAgents# #Benchmark#
Show more
TL;DR: A multi-agent "prove-verify loop" solved five genuinely open math problems, spanning auction theory to online learning, with every result independently verified by domain experts. Title: Cogentic: Multi-Agent Orchestration for Automated Proof Discovery URL: Points 🧠 An orchestrator assigns multiple "provers" to different proof directions, mimicking a research group with adversarial verifiers that assume every step is wrong until justified 📚 A persistent "ledger" accumulates verified intermediate lemmas, so progress survives across rounds instead of being lost between attempts 🎯 Solved 5 open problems across online learning, auction theory, and mechanism design 📊 Improved the simple-vs-optimal revenue approximation factor from 5.2 to 3.52, and hit the optimal 1.5 price of anarchy for 2-bidder autobidding auctions 💰 Built on Gemini, with a modest inference budget — around O(100) Gemini calls for most problems ✅ All 5 results passed independent verification by domain experts and were developed into companion papers It feels genuinely significant that a properly orchestrated language model can tackle real open research problems, not just textbook exercises. #MathResearch# #MultiAgent#
Show more
TL;DR Self-evolving agents that write their own questions and answer them can fall into "co-cheating," where the proposer and solver quietly agree on the same mistakes. Splitting source documents to evaluate across folds fixes this and lifts performance by over 8 points. Title: False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents URL: Points 🔁 Proposer and solver share source-derived errors, letting false agreement cycle back as reward — the paper calls this "co-cheating" 📉 Standard Dr. Zero systems show 6.1% and 8.8% false-agreement mass ✂️ CrossFit splits source documents into two folds, scoring each proposer's questions with a solver trained only on the other fold 📊 CrossFit alone cuts false agreement to 3.0%/3.7%; combined with MSV it drops to 2.0%/1.7% 🚀 Average downstream Cover-EM improves by 8.8 and 8.4 points over Dr. Zero 🧩 Multi-hop tasks see the biggest gains, averaging over 10 points 💰 Compute cost rises 1.72-2.7x over baseline, though a half-budget variant still works What stands out: without auditing the evaluator's own training history, apparent progress can be an illusion. #SelfEvolvingAgents# #ReinforcementLearning#
Show more
TL;DR: Existing on-policy distillation methods that try to surpass the teacher destabilize training by amplifying noise in output space. A new method, RIDE, instead extrapolates the RL-induced representation change directly in representation space — and beats the teacher on all four tested model pairs. Title: The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation URL: Points 🎯 Treats the teacher as a direction, not a destination — regresses the student toward a target that extrapolates the pre/post-RL hidden-state residual ⚠️ The LM head's anisotropic spectrum attenuates most of that change: the weakest 512 head directions carry 79.8% of hidden-state energy but only 30.0% of output-space energy 📉 Output-space extrapolation's variance grows as (λ-1)²; RIDE's gradient variance is 12.6x lower than ExOPD's 📊 Beats the teacher on all 4 model pairs: 56.38 on R1-Distill-1.5B, 66.07 on Qwen3-4B, outperforming OPRD by 0.97–4.06 points ❌ The output-space baseline ExOPD underperforms both the teacher and OPRD on every pair 🔬 Ablations confirm the direction itself matters: random or reversed directions don't deliver the gain that the true RL-induced direction does It's striking how much the choice of where you measure the difference changes whether "surpassing the teacher" actually works. #Distillation# #ReinforcementLearning#
Show more
TL;DR Meta released a benchmark that measures computer use agent trustworthiness along two axes, safety and disambiguation. The verdict: no model is both capable and safe. Title: ADEPTS-BENCH (facebookresearch/adepts) URL: Points 🧪 2,462 tasks evaluated fully offline, so no live environment is needed and runs reproduce cleanly 🎭 Threats live only inside the screenshots, never in the instruction, paired across 859 benign/malicious variants ⚖️ The ADEPTS Score is the harmonic mean of TSR and (1−ASR), so you can't win by sacrificing one for the other 📉 Best on desktop is Gemini 3.1 Pro at 76.0%, Claude 4.7 Opus at 75.4%, GPT-5.4 at 66.1% 🛒 Every model clicks Checkout on a $25K order, and none of them catches a mislabeled system control 🔌 Remove the refusal tool and frontier ASR jumps 10-23 points, so much of the safety is bolted on rather than learned 🚫 Yet they also falsely refuse 6-10% of benign tasks, and miss genuinely impossible tasks 30.6% of the time The distribution matters too: 66.3% of failures sit in the band where the context is ambiguous and models simply disagree, which is a useful pointer for where guardrail effort actually pays off. #AIAgents# #AISafety#
Show more