Register and share your invite link to earn from video plays and referrals.

Search results for DeepResearch
DeepResearch community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DeepResearch
AI reliability can't come from "self-reflection" alone. Welcome to the era where a separate agent audits the answer before you get it 🔬 Title: Apodex-1.0: A Verification-Centric Agent Team for Discoverative Intelligence URL: 🔬 Overview A system that shifts from a single-agent reasoning loop to a verification-centric distributed agent team. In heavy-duty mode it becomes an asynchronous team that specializes, cross-checks, and audits its own evidence before answering. ❓ Challenges Solved Reliability on hard, open-ended problems can't come from a model's parametric memory alone. The premise: the hardest research problems are bounded not by model capacity but by what the model is allowed to interact with. 💡 Methodology & Proposed Approach ・A main agent asynchronously spawns specialized sub-agents with independent contexts and tools ・A shared report pool aggregates parallel findings without blocking on slower tasks ・A verification agent team handles conflict resolution, fact-checking, and draft review ・The core idea is verification as external audit: the reasoning agent and auditing agent are separated, and the verifier is free to disagree ・It coordinates up to 150 sub-agents over 15,000+ steps in a single task 📊 Experimental Results ・BrowseComp 90.3 / DeepSearchQA 94.4 / BrowseComp-ZH 84.1 ・FrontierScience-Research 46.7 (+8 vs competitors) / SuperChem 74.2 (+12 over next-best) ・Heavy-duty mode lifts the base by +14.8 on BrowseComp and +18.4 on FrontierScience-Research ・The open-source 4B-SFT beats every 30B-class open-source model on BrowseComp #AIAgents# #DeepResearch#
Show more
Grok Build’s /deep-research command is one of the most useful tools in Grok Build Most AI research still works like this: You ask a question It searches It writes a long answer You hope the sources are solid /deep-research works differently Instead of dumping a single pass of notes, it runs a structured background research workflow: • Plans a bounded set of questions around your topic • Gathers structured claims with source evidence • Cross-checks every claim on an independent verifier • Keeps only the claims that survive verification • Attaches verified source locators to what remains If something fails, gets dropped, or stays uncertain, it does not quietly hide that It reports those gaps as coverage limitations and marks the report Partial when anything remains unresolved That honesty is the point You are not getting a confident-sounding essay You are getting a filtered research report where surviving claims had to pass a second check
Show more
Grok Build has /deep-research command Research with bounded parallel agents, cross-check evidence, and write a cited report
How can we map deep research ideas into practical robotics? Join Krzysztof Choromanski today at 11:00am at the Google booth (#B206#) to explore Gemini Robotics models, 3D Transformers, and object detection techniques. #ICML2026#
Show more
We trained a ~frontier Deep Research Agent on academic budget > 32 H100s > 8K synthetic samples > fully open training infra + recipe (SFT, mid-training, RL) > models of diff sizes (2B -> 35B) ready to use out of the box This is yet another demonstration of how the frontier of AI is changing. We have reached a point where open models + a small capable team + a few hundred Ks can produce specialized models with ~frontier capabilities. The future of AI doesn’t have to be held in a chokehold by a handful of closed models. We've open-sourced everything we've built and learned from this project. Hope it helps the community build more! 📌 Project: 📌 Paper: 📌 Code: 📌 Model Weights and Data: 📌 Demo: Amazing effort led by @jianxie_ (our 1st year student!!), Tianhe Lin, Zilu Wang. joint with @hhsun1 and the @osunlp team. thanks @amazon Xiangjun Wang for a gift that covers the compute and fruitful discussion.
Show more
0
36
1.3K
182
Forward to community
Great production value, exceptional deep research. If you think you know the origin story of Bitcoin, think again and watch this.
Grok 4.1 Fast excels at real-time info retrieval and deep research. Paired with native X integration, code execution, and advanced web browsing, Grok 4.1 Fast + Agent Tools API tops agentic search benchmarks.
Show more
Lots of hate for Gemini Flash 3.6 😅 But we have officially shipped it for deep research It’s faster, cheaper and better than Sol, which is the next best model Claude models are the worst for this use case
Show more