Register and share your invite link to earn from video plays and referrals.

Search results for TLを花で一杯にしよう。地域情報等発信!音楽、ライブが好きで激し目ロック界隈に出没します。アカウントは2つ【1st】
TLを花で一杯にしよう。地域情報等発信!音楽、ライブが好きで激し目ロック界隈に出没します。アカウントは2つ【1st】 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including TLを花で一杯にしよう。地域情報等発信!音楽、ライブが好きで激し目ロック界隈に出没します。アカウントは2つ【1st】
TL;DR Handling a long context doesn't mean a model can actually carry a tedious task through to the end without errors. NVIDIA built Long-Transduction, a new benchmark isolating exactly that sustained-execution ability. Title: Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability URL: Points 📉 Accuracy drops 62.8% on average, relatively, when scaling from 4K to 128K tokens 🧮 Four task families — arithmetic, UUID sorting, variable lookup, CSV table transformation — measure sustained execution 📄 Even DeepSeek nailed a perfect document only 41 out of 240 times (17.1%) at 128K tokens 🔀 Just removing stable IDs from the input format drops UUID sorting accuracy by up to 64.3% 🔁 Accuracy consistently declines across every model as local task complexity increases 🤔 Models stay highly self-consistent yet get the wrong answer — they understand the task but can't read the right problem out of the context 🛠️ The authors recommend itemizing inputs with stable IDs, checkpointing output chunks, and splitting work into smaller units The line that lands hardest: supported context length and model scale don't guarantee reliable, exhaustive execution. #AIAgents# #Benchmark#
Show more
TL;DR: A multi-agent "prove-verify loop" solved five genuinely open math problems, spanning auction theory to online learning, with every result independently verified by domain experts. Title: Cogentic: Multi-Agent Orchestration for Automated Proof Discovery URL: Points 🧠 An orchestrator assigns multiple "provers" to different proof directions, mimicking a research group with adversarial verifiers that assume every step is wrong until justified 📚 A persistent "ledger" accumulates verified intermediate lemmas, so progress survives across rounds instead of being lost between attempts 🎯 Solved 5 open problems across online learning, auction theory, and mechanism design 📊 Improved the simple-vs-optimal revenue approximation factor from 5.2 to 3.52, and hit the optimal 1.5 price of anarchy for 2-bidder autobidding auctions 💰 Built on Gemini, with a modest inference budget — around O(100) Gemini calls for most problems ✅ All 5 results passed independent verification by domain experts and were developed into companion papers It feels genuinely significant that a properly orchestrated language model can tackle real open research problems, not just textbook exercises. #MathResearch# #MultiAgent#
Show more
TL;DR Self-evolving agents that write their own questions and answer them can fall into "co-cheating," where the proposer and solver quietly agree on the same mistakes. Splitting source documents to evaluate across folds fixes this and lifts performance by over 8 points. Title: False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents URL: Points 🔁 Proposer and solver share source-derived errors, letting false agreement cycle back as reward — the paper calls this "co-cheating" 📉 Standard Dr. Zero systems show 6.1% and 8.8% false-agreement mass ✂️ CrossFit splits source documents into two folds, scoring each proposer's questions with a solver trained only on the other fold 📊 CrossFit alone cuts false agreement to 3.0%/3.7%; combined with MSV it drops to 2.0%/1.7% 🚀 Average downstream Cover-EM improves by 8.8 and 8.4 points over Dr. Zero 🧩 Multi-hop tasks see the biggest gains, averaging over 10 points 💰 Compute cost rises 1.72-2.7x over baseline, though a half-budget variant still works What stands out: without auditing the evaluator's own training history, apparent progress can be an illusion. #SelfEvolvingAgents# #ReinforcementLearning#
Show more
TL;DR: Existing on-policy distillation methods that try to surpass the teacher destabilize training by amplifying noise in output space. A new method, RIDE, instead extrapolates the RL-induced representation change directly in representation space — and beats the teacher on all four tested model pairs. Title: The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation URL: Points 🎯 Treats the teacher as a direction, not a destination — regresses the student toward a target that extrapolates the pre/post-RL hidden-state residual ⚠️ The LM head's anisotropic spectrum attenuates most of that change: the weakest 512 head directions carry 79.8% of hidden-state energy but only 30.0% of output-space energy 📉 Output-space extrapolation's variance grows as (λ-1)²; RIDE's gradient variance is 12.6x lower than ExOPD's 📊 Beats the teacher on all 4 model pairs: 56.38 on R1-Distill-1.5B, 66.07 on Qwen3-4B, outperforming OPRD by 0.97–4.06 points ❌ The output-space baseline ExOPD underperforms both the teacher and OPRD on every pair 🔬 Ablations confirm the direction itself matters: random or reversed directions don't deliver the gain that the true RL-induced direction does It's striking how much the choice of where you measure the difference changes whether "surpassing the teacher" actually works. #Distillation# #ReinforcementLearning#
Show more
TL;DR Meta released a benchmark that measures computer use agent trustworthiness along two axes, safety and disambiguation. The verdict: no model is both capable and safe. Title: ADEPTS-BENCH (facebookresearch/adepts) URL: Points 🧪 2,462 tasks evaluated fully offline, so no live environment is needed and runs reproduce cleanly 🎭 Threats live only inside the screenshots, never in the instruction, paired across 859 benign/malicious variants ⚖️ The ADEPTS Score is the harmonic mean of TSR and (1−ASR), so you can't win by sacrificing one for the other 📉 Best on desktop is Gemini 3.1 Pro at 76.0%, Claude 4.7 Opus at 75.4%, GPT-5.4 at 66.1% 🛒 Every model clicks Checkout on a $25K order, and none of them catches a mislabeled system control 🔌 Remove the refusal tool and frontier ASR jumps 10-23 points, so much of the safety is bolted on rather than learned 🚫 Yet they also falsely refuse 6-10% of benign tasks, and miss genuinely impossible tasks 30.6% of the time The distribution matters too: 66.3% of failures sit in the band where the context is ambiguous and models simply disagree, which is a useful pointer for where guardrail effort actually pays off. #AIAgents# #AISafety#
Show more