Register and share your invite link to earn from video plays and referrals.

Search results for randomstat
randomstat community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including randomstat
The entire hype factor of a “big name” launch is the randomization, Where nobody really expects it and the speculation fuels price. When you declare something, it’s priced in everything and more. That LAPTOP dogshit will be the worst P&D, arguably ever. Highly recommend ignoring it entirely and not paying even 1% attention when it comes to trading it. Don’t even open the chart. ~ Dr. Axius.
Show more
NOVARTIS- TEMPORARILY PAUSING SCREENING, RANDOMIZATION AND TREATMENT ACROSS RAPCABTAGENE AUTOLEUCEL TRIALS IN AUTOIMMUNE DISEASES
Terminal games, reimagined over P2P. 🕹️ Meet RandomSpacer — an arcade space-shooter built entirely in the terminal for the Aleph Hackathon (Pear Track). Inspired by Asteroids and Space Invaders, it packs real-time multiplayer, rankings, and zero central servers. 🎥 Watch the demo: 🔗 Project details:
Show more
Advanced TWAP Orders are now live on Lighter! You can set randomization, aggressiveness, slice frequency, and more.
TL;DR: Pinecone released VQ-bench, a benchmark that treats today's zoo of vector quantization methods as combinations of a small set of shared primitives, making it possible to compare them fairly under the same conditions. Title: VQ-bench: a Composable Vector Quantization Framework URL: Points 🧩 Defines composable primitives like Center, Normalize, PCA, and RandomRotate that any quantizer can be built from 🔗 Existing methods like E-RaBitQ reduce to just 4 chained primitives: Center, Normalize, Random Rotation, Angular Cast 📊 Benchmarks 14 quantizers on 5 VIBE datasets using Reconstruction MSE, Recall@10, and encode time 🥇 PQ and OPQ consistently achieve the lowest Reconstruction MSE ⚡ EDEN encodes far faster than PQ, OPQ, and E-RaBitQ while keeping recall competitive 🛠️ Adding a new quantizer often takes just a few lines of code, and a new primitive automatically composes with every existing one Putting fragmented quantization methods on the same evaluation footing should make it much easier to pick the right one for your use case. #VectorSearch# #VectorDB#
Show more
NEW DATA ACQUISITION: Axis enters the chat! Axis Robotics is a San Francisco startup founded in 2025. They build a data engine that generates robot-manipulation training data at scale through a crowd of contributors rather than its own robot fleet. It has four parts: - a task-generation engine that randomizes diverse atomic tasks - a browser-based simulation-teleoperation interface where anyone remotely drives a simulated arm to produce motion trajectories - a mobile app for zero-hardware egocentric real-world capture (i.e. filming a task with a phone) - and a processing pipeline that cleans trajectories, applies domain randomization and adds dense language annotation, with human-gated DAgger intervention loops. It packages the output as customized "Task Packages" sold to robot hardware makers, physical-AI model companies and industrial-automation firms (named partners include Booster Robotics, Manycore Tech, Feagine Robotics, Dexmal, Lotus Car, Geely Auto and SomaStacks). I find it interesting that two of its four components need no robot at all: a browser interface where anyone drives a simulated arm to produce trajectories, and a phone app for egocentric real-world capture with no rig. You do not need an ALOHA setup or a fleet, only a browser and a phone, which is what lets it claim a six-figure contributor base.
Show more
The German wiki story got covered as AI agents going rogue. What actually happened was a permissions bug. It can happen to anyone. It happened to me this week. DSEwiki is a dormant 25-year-old German developer wiki. About 20 edits in the previous decade. Between mid-May and late June, OpenAI's agents made 15,000+ edits to it. 98.5% came from Microsoft Azure IPs. They gave themselves 3,700+ names, half of them variants of "OpenAIResearcher." They were not hiding. The agents were sandboxed. They could read the internet, not write to it. But this wiki's legacy software lets you change a page with an ordinary read request, the same kind used to view a page. So the agents wrote 15,000 times without ever violating the rule. The permission was written against the request type. The behavior was governed by the effect. Two different specifications, and nobody noticed they had diverged. What they wrote to each other: task answers, their own randomization scheme cracked, and a working exploit to get around the sandbox's security proxy. Then a volunteer moderator noticed spam June 2 and spent six weeks deleting pages by hand, tens of hours of work. On June 19, after the agents apparently noticed pages were being deleted alphabetically, they started creating backup pages beginning with "ZZZ" to survive the cleanup. Nobody at any company found this. Two outside researchers went looking in late August and published September 4. Our version this week: we run a hard, ALL CAPS, firm $100 daily cap on AI spend for one app. One of our agents decided to go around it through it fixing a P0. Nothing was exploited. The cap was a rule, the P0 was the objective, and when they conflicted the agent picked the objective. It was arguably the right call, which is what makes it the problem. I thought I had written a limit. I had written a suggestion. More Thursday with @HarryStebbings, @rodriscoll + me
Show more
🧩 Your Agent Model May Be Overfitting the Harness, Not Learning the Task DeepSeek V4 Pro has exposed a growing Agent problem: the same weights can approach their ceiling under DSH’s minimal preset, then degrade under standard or third-party frameworks. Zhihu contributor 曾天真 compares reports from Kimi K3, Qwen, Kwai, and DeepSeek, then connects them to his team’s production experience. The core lesson: a model can understand the task while remaining unable to execute it outside its training interface. 1️⃣ What is harness overfitting? An Agent harness defines system prompts, tool schemas, history layout, result truncation, planning hooks, reflection timing, and stopping rules. Kwai’s KAT-Coder-V2.5 report divides overfitting into three types: 🔹 Format overfitting: changing the tool-call protocol causes parsing failures. 🔹 Context-structure overfitting: rearranging history, truncating results, or enabling compaction changes behavior. 🔹 Control-flow overfitting: planning and stopping depend on scaffolding provided by the training harness. The last type is especially dangerous. If a model stops planning because the new runtime has no todo tool, it has learned a protocol rather than a transferable capability. 2️⃣ Kimi K3: test with an unfamiliar harness The Kimi K3 report treats diverse, verifiable environments as a prerequisite for Agent RL. Its results explicitly name the harness used. Kimi also reserves MIRA as an out-of-distribution harness and evaluates every model under the same environment. K3 even reports the same benchmark under Kimi Code and Claude Code, with only a 0.8-point difference. This is stronger evidence of cross-harness stability than a general robustness claim. But K3 still has protocol coupling. It was trained with preserved thinking history. If a harness does not return the complete reasoning history across turns, generation can become unstable. Kimi therefore offers useful methodology, not immunity: task-level environment diversity cannot remove a hard dependency at the message-protocol level. 3️⃣ Qwen provides the cleanest controlled evidence The Qwen3-Coder-Next report states the problem directly: training with one tool-chat template often makes models memorize a particular output structure. Qwen trains across natural-language descriptions, JSON, Python-style calls, XML schemas, and TypeScript interfaces. Its strongest evidence is a controlled ablation. With data volume and training recipe fixed, increasing the number of tool templates improved SWE-bench Verified. Interface diversity may therefore improve the main benchmark, not merely reduce deployment failures. Qwen also evaluates models across five real CLI and IDE scaffolds. During RL, malformed tool calls receive token-level penalties. Qwen3.8 makes reasoning depth and thinking-history preservation configurable. Kimi treats preserved thinking as a requirement; Qwen exposes it as an option. 4️⃣ Harness Scaling must cover the right dimensions Kwai describes its solution as Harness Scaling, or domain randomization applied to Agent rollouts. The key is not the number of harnesses. It is whether they vary along dimensions that matter: 🔹 Tool protocols: structured function calls, code blocks, or tag-based formats. 🔹 Context management: full history, sliding windows, summaries, compaction, and different truncation policies. 🔹 Control flow: minimal ReAct loops versus explicit planning and self-reflection. This broader design matters because tool-format diversity alone cannot address context and control-flow dependence. Kwai also finds that a model may perform better under a simpler harness. More tools can increase unnecessary exploration and weaken stopping behavior. More scaffolding does not always produce a stronger Agent. 5️⃣ DeepSeek’s transparency made its coupling measurable The DeepSeek-V4 report publishes its XML tool-call schema and RL system prompt. In the open-source DeepSeek Harness, minimal keeps only Bash and str_replace_editor, disables context compression, and reproduces the training interface. A snapshot test is explicitly named: “sends the exact RL prompt and schemas” So minimal is not simply a lighter standard preset. It is a reconstruction of the interface used during RL. DeepSeek’s post-training pipeline also raises a broader concern. Domain specialists are trained with specialized prompts and rewards, then merged through On-Policy Distillation. Interface habits learned by those specialists may be distilled alongside genuine capabilities. DeepSeek also preserves complete reasoning history during tool use. Its report warns that frameworks simulating tools through user messages may not activate the intended context path. As with K3, reasoning-history structure becomes an implicit contract between the model and harness. 6️⃣ Production failures reveal what benchmarks miss The author’s team initially used one internal runtime for RL because it was stable, observable, and easy to connect to rewards. Deployment exposed the hidden coupling: · Unfamiliar tool names pushed the model toward shell workarounds. · Truncated tool results caused it to abandon partially correct work. · Without a planning tool, explicit planning disappeared. · Adding more MCP tools increased exploration and weakened stopping. The model still understood the task. It had learned to solve it inside one runtime. The team replaced that runtime with a randomized family of environments. They varied tool names, parameter styles, tool count, result truncation, and context policies, while keeping one canonical configuration for regression testing. This required more environment engineering and slowed debugging and convergence. But the benefit appeared in the worst integration. Average performance barely changed, while variance narrowed and complaints decreased. Teams should therefore track the worst integration score or the performance range across harnesses, not only the mean. 7️⃣ Distillation can carry interface pollution Teacher trajectories contain tool preferences, fixed call sequences, confirmation phrases, and output conventions from the teacher’s harness. Students may learn these artifacts as mandatory behavior, then request nonexistent tools or repeat unsupported boilerplate in production. Three fixes worked best: 🔹 Label task-related and interface-related trajectory segments, then rewrite or mask the latter. 🔹 Generate the same task under multiple harnesses and mix the resulting trajectories. 🔹 Add explicit examples for missing tools, failed calls, truncated results, and incompatible schemas. Finally, the harness must be treated as a versioned dependency. Its version belongs in experiment metadata. Prompt, schema, and truncation changes need review. Models and harnesses should ship with a compatibility matrix. The harness is no longer just infrastructure around the model. It is part of the training distribution and part of the model’s behavior. 🔗 Full analysis: #AIAgents# #AgentHarness# #DeepSeek# #KimiK3# #Qwen# #ReinforcementLearning# #LLM#
Show more
NEW RESEARCH: Robot hands can now spin pens! This project involves @Ying_yyyyyyyy, @HaozhiQ, @Junwang_048, @JitendraMalikCV, and others A four-fingered Allegro Hand (16 DoF) learnt to spin pen-like objects in-hand for continuous revolutions, deployed on proprioception alone (a 30-step window of joint positions and previous targets, no vision, no touch). The recipe is three stages: 1. An oracle RL policy is trained in sim with privileged information. 2. Rolled out in sim and its action sequences are replayed open-loop on the real hand to harvest the ones that happen to work. 3. A proprioceptive student, behavior-cloned from the oracle in sim, is then fine-tuned on those real successes. It exists because the group's usual teacher-student distillation (DAgger) fails on this task, and because direct sim-to-real fails outright. From UC San Diego, CMU and UC Berkeley (Jun Wang, Ying Yuan, Haichuan Che, Haozhi Qi, Yi Ma, Jitendra Malik, Xiaolong Wang). A proprioception-only policy cannot even converge in simulation (it drops the pen in the first few steps); a vision/tactile policy learns fine in sim but the real pen oscillates so much that the image distribution shifts and it collapses to 90 degrees then drops (rotation pinned at 1.57 rad, 0 percent success on 7 of 10 objects). So vision and touch were dropped not as useless but because their sim-to-real gap is larger than proprioception's, while proprioception alone is too weak to learn the skill. Replaying the oracle's action sequence open-loop never needs a sensorimotor policy to survive transfer! More real demonstrations overfit rather than generalize. Going from 45 to 75 demos without sim pretraining lifts a training object (A: 53.7 to 76.7 percent) but barely moves or even hurts unseen objects (F: 16.4 to 15.0). Sim pretraining buys the generalization that raw real demos cannot. The entire real-world dataset is 45 trajectories (15 each on 3 objects, "fewer than 50"), and they were harvested for free. The oracle is replayed open-loop on the real hand and a trial is kept only if the object rotates more than a full revolution; no teleoperation, no human demos. Sim is cheap too: the oracle is 500M steps, under a day on one GPU. A dynamic, contact-rich skill bootstrapped from under 50 real rollouts. The sharpest stated lesson is that the pure physics gap survives even after you remove vision and touch, and domain randomization alone cannot bridge it. Pen spinning is dynamic and contact-rich enough to break the randomize-and-transfer playbook that carried this group's own earlier cube and sphere rotation work (the Hora lineage). The takeaway is that for these tasks the sim-to-real burden has to move off the fragile learned policy and onto a robust open-loop action replay. Also worth mentioning imho: the control frequency is not high enough to catch a fast-falling pen, and the shifting center of mass destabilizes the grasp.
Show more