Register and share your invite link to earn from video plays and referrals.

Search results for ReasoningModels
ReasoningModels community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including ReasoningModels
🌱 Seed-2.1-Pro Review: Better Post-Training Can't Hide an Aging Base @ByteDanceSeed_ Seed-2.1-Pro 0915 earns higher reasoning scores with fewer tokens, and its agent ability has climbed from nearly unusable to passable. But in the three months since the last version, rivals iterated roughly a generation and a half — and the old base model is running out of tricks. That is the verdict from Zhihu contributor toyama nao, who runs a long-running monthly logic benchmark and put the 0915 build through his full evaluation suite. 1️⃣ Coding and agent work: from unusable to passable The jump over the predecessor is large. Deliverable completeness now beats DeepSeek V4.1 Flash, though first-pass success still trails top-tier models — a model caught in the middle. 🔹 Frontend: some aesthetic sense, but unstable. Constrained stacks like iOS system components look decent; open stacks like web or raw canvas expose visible flaws in proportion and color. 🔹 Task adaptability: every task now at least completes. In a HarmonyOS-flavored project it barely knew, the model read docs and iterated its way to usable — a good sign. 🔹 Delivery efficiency: roughly tied with DeepSeek V4.1 Flash and GLM-5.3-Flash, splitting steps evenly between writing and verifying. All three sit below the current Chinese SOTA. 2️⃣ Multi-step reasoning: the biggest gain This is where 0915 improved most — and where it leads its tier, with a small token-efficiency edge. On problems where GLM-5.3 falls into exhaustive enumeration, 0915 repeatedly finds higher-scoring answers with fewer tokens, showing what the author calls real "big-model intuition." The caveat: multi-turn reasoning, which demands in-context learning and reflection, remains mediocre — on par with Chinese peers. 3️⃣ Where 0915 still stumbles Hallucination runs high, and context confusion appears regardless of prompt length — especially when source details are tangled. The agent symptom is subtler than dropping requirements: 0915 keeps every requirement but misreads semi-ambiguous ones — arguably the more dangerous failure mode. 4️⃣ The contrarian bet: a true non-thinking mode Seed is one of the few teams still maintaining a genuinely non-thinking mode, and this version quietly got good: average output dropped from 8K tokens back to ~1K, with no measurable capability regression, a slight gain in complex reasoning, and readable prose intact. For latency-sensitive scenarios that still need some reasoning, the author considers it a legitimate option. 5️⃣ The long march His closing frames 0915 as a rest stop, not a destination: the predecessor's lukewarm market reception forced ByteDance's team into a forced march of biweekly iterations, and better post-training is now visibly paying off in cost per task. But the base is old, and the competition has moved. His last line is worth keeping: sometimes the long way around is the real shortcut. 🔗 Full Reading: 🔗 Key links: Author's monthly logic benchmark (Aug 2026): #ByteDance# #Seed# #LLM# #AIAgents# #LLMBenchmark# #ReasoningModels# #AI#
Show more
Reading an AI's "chain of thought" to predict its behavior? Turns out that's not very reliable 🔮 The fresh idea: make behavior prediction itself a learning task. Title: Forecasting Future Behavior as a Learning Task URL: 🔮 Overview A method to predict how large reasoning models (LRMs) will behave on new inputs. Instead of relying on explicit explanations, it introduces trainable "Behavior Forecasters" that analyze a single reasoning trajectory to predict outputs. ❓ Challenges Solved We want to understand and predict LRM behavior, but prior approaches have limits. ・Existing explanation methods don't scale to long reasoning trajectories ・Read as natural language, those trajectories are often unreliable A model's written "thoughts" don't necessarily reflect its actual behavior. 💡 Methodology & Proposed Approach ・It treats behavior prediction itself as a learnable task ・Training data comes directly from querying LRMs — no human annotation needed ・It runs in a single forward pass at inference ・Instantiated on two tasks: estimating answer consistency across reruns, and predicting how input modifications affect outputs ・End-to-end fine-tuning of the backbone, and initializing from the target LRM's weights, proved essential 📊 Experimental Results ・Behavior Forecasters outperform GPT-5.4 and Claude Opus-4.6 as "naive readers" ・And they achieve higher accuracy at a small fraction of the inference cost #LLMInterpretability# #ReasoningModels#
Show more
SERV Reasoning API is now live: specialized AI models to make agents smart and reliable. Pair SERV Reasoning models with @CoinbaseDev AgentKit to build enterprise-grade onchain agents for DeFi, trading, commerce, payments. Join the hackathon:
Show more
As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens actually buy is becoming increasingly important. In this episode, @Stanford professor and Big Spin co-founder @ChrisGPotts joins us to discuss AI tokenomics and his research into “tokenflation”—the possibility that token usage is growing faster than the measurable value those tokens produce. We explore how to measure the return on AI spending, why benchmarks alone provide an incomplete picture of model progress, and what inference-time scaling means for the economics of increasingly capable models. Chris also explains why expert AI users tend to get better results by challenging and iterating with models, how AI fluency affects outcomes, and why more efficient architectures could change the underlying economics. We also discuss DSPy, interpretability, the limits of today’s transformer architectures, and where Chris sees opportunities for more fundamental innovation in AI. 🗒️ Full show notes: 📖 CHAPTERS =============================== 00:00 - Introduction 05:38 - Linguistics in the Age of Language Models 09:26 - Scale Limitations in NLP Research 12:54 - Challenging the Bitter Lesson Mindset 15:12 - Relationship Between Data, Mechanistic Interpretability, and Efficiency 17:03 - DSPy 21:32 - Prompt Optimization and Model Variability 24:35 - Tokenomics and the Rising Cost of AI 28:13 - Measuring Token Purchasing Power with a CPI 32:18 - Inference-Time Scaling 35:36 - Defining Value Across Different AI Tasks 38:33 - AI Value Creation 40:38 - Predicting AI Costs 42:25 - Tokenflation 46:22 - AI Fluency 50:12 - Key Lessons of AI Fluency Work 54:44 - Future Directions
Show more
A very interesting new approach to AI reasoning: Flow Reasoning Models. Instead of generating a solution sequentially and being stuck with earlier decisions, FRMs repeatedly refine the entire solution, allowing the model to reconsider and correct its own decisions until it converges. The results are impressive: 99.5% on Sudoku-Extreme, 100% on Zebra and 99.9% on Maze-Unique. Even more interesting, FRMs match the 98.7% solve rate of the next-best Sudoku method with 44× fewer inference FLOPs. Iterative self-refinement may be a very powerful way to scale reasoning without simply throwing enormous amounts of compute at inference.
Show more
Excited to share the latest expansion of the @nvidia #Alpamayo# open platform for reasoning-based autonomous vehicles. Since its launch earlier this year, Alpamayo has seen rapid adoption across industry and academia, with its reasoning models surpassing 400,000 downloads and earning a #COMPUTEX# 2026 Best Choice Award. As announced by Jensen Huang during his #COMPUTEX# keynote, we are now introducing several major additions designed to accelerate the development of next-generation AV systems (more details here: 🚗 Alpamayo 2 Super — a new 32B-parameter driving foundation model with: • Full 360° surround-view perception • Advanced reasoning capabilities and chain-of-causation outputs • Meta-actions such as lane changes, yielding, and stopping • Reasoning auto-labeling and visual grounding for scalable data annotation • State-of-the-art performance across reasoning, prediction, and alignment tasks 🔄 AlpaGym — an open-source framework for closed-loop reinforcement learning, enabling AV models to learn from the consequences of their actions in simulation and helping bridge the gap between training and real-world deployment. 📊 New Open Benchmarks — including challenges for closed-loop driving and long-tail reasoning to help the community measure progress and drive innovation. 🛠️ Alpamayo Recipes — a centralized repository of end-to-end workflows covering supervised fine-tuning, reinforcement learning, quantization, and model customization. Reasoning models and closed-loop training are becoming foundational technologies for autonomous systems. Our goal is to provide the open tools, models, infrastructure, and benchmarks needed to accelerate progress across the entire AV ecosystem. A huge thank you to the many researchers, engineers, and community members whose feedback helped shape this release. Resources: • Overview of the latest Alpamayo release (note: some components will be released over the coming weeks): • @nvidia announcement: #AutonomousVehicles# #PhysicalAI# #Robotics# #AI# #MachineLearning# #ReinforcementLearning# #OpenSource# #NVIDIA# #Alpamayo# @NVIDIADRIVE @NVIDIAAI
Show more
LLMs (up to GPT-4) were System 1 models. Jev is also a System 1 model. Reasoning models are System 2 models. But what are System 2 models for Jev-like models?!
$MU $SKHY $DRAM Inference era: here now. Agentic AI: barely off zero. And agentic already needs 10 to 40x more memory than standard inference. Reasoning models generate 3 to 10x more tokens. Position accordingly.
Show more
We are hiring! The Autonomous Systems and Physical AI Research (ASPIRE: group at @nvidia is looking for talented PhD Research Interns to join us in advancing the frontiers of #autonomous# #systems# and #Physical# #AI#. We work across a broad range of research areas, including #reasoning# models, generative simulation, #agentic# AI workflows, and Physical AI #safety#, with applications spanning autonomous vehicles and a broad range of Physical AI systems. Interested in pushing the limits of what’s possible? Apply now:
Show more
OpenAI @yong_zhengxin reveals GPT-4 couldn't do current cyber breakthroughs even with infinite compute: "I don't think GPT-4 has really good reasoning. One of the key parts is you want to make sure the explanations you propose are consistent with the prior knowledge." "You do want very strong reasoning models. Before the current model, there was some test about, do you want to take the car to the car wash, certain weird cookie questions you ask the models, and models weren't able to figure out the right answer." "You want the capability of reasoning what you see right now and how it fits to the big prior body of knowledge. GPT-4 wouldn't be able to do that even if you give infinite compute, because I don't think it has that good of a reasoning capability." @OpenAI
Show more