Register and share your invite link to earn from video plays and referrals.

Yao Fu
@Francis_YAO_
Prev. Grok 4.5/4.6 pretraining scaling @xAI; Gemini 3 perception and project Astra @GoogleDeepMind
2.2K Following    23.6K Followers
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:
Show more
0
375
8.8K
866
Forward to community
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
Show more
0
419
9.9K
976
Forward to community
Introducing ApprenticeBench: computer use + continual learning on a real job. We show Fable 5.1 and GPT-6 Astra can now continually learn on a job and surpass human professionals. A decisive step change in AI's job readiness. No FDEs. Agents deploy themselves into the job. 🧵
Show more
0
34
1.2K
170
Forward to community
At Fireworks, we believe open models and ecosystem let intelligence compound where value is created - inside every company serving a special purpose. We signed the open letter to support open weights.
Show more
The most important word here is *ecosystem*. It's not just about having an open-weight model. Open-weight models are a means to an end. To have a truly strong, open ecosystem, we need four critical frontier-level ingredients: open-weight models, open training datasets, open software stacks, and open process knowledge. Few people realize that NVIDIA actually has been pushing beyond open weights by releasing code and datasets for their Nemotron models, which is something open-weight model developers don't do. Marin further opens up the process knowledge - not just how to train one model, but how to iteratively improve and shape a model given particular goals, custom data, and hardware, e.g., how to design scaling laws and evals to guide architecture and data ablations. Open weights, datasets, software, process knowledge: these are the four critical ingredients (renewable resources) that give everyone the ability to most efficiently turn their compute (consumable resources) into the best models according to their needs and values.
Show more
Strong support for that. The worst situation that could happen to human kinds is that a few elites control the best model (and their APIs), while other people treat them as "god" and pray for access. That would be the real dystopia...
Show more
Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available from @GoogleDeepMind @demishassabis . Onwards!
Show more
0
749
22.7K
1.9K
Forward to community
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
0
16.1K
172.2K
29.5K
Forward to community
@garctrob @ronnygjunkins @EyubogluSabri @HazyResearch We’re excited about this simpler view of MLPs! Check out our paper for the formal details. Next: can we extract and directly rewrite facts inside pretrained LMs? 📝 Blog: 📄 Paper:
Show more
Fantastic slides from the ICML tutorial of @MarkSchmidtUBC (“is opt theory relevant in 2026”)
Who tf else want to question linear attention now?
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: 🔗 Tech blog:
Show more
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: 🔗 Tech blog:
Show more
0
1.7K
57.1K
7.6K
Forward to community
It was quite a journey working on the pretraining scaling for Grok 4.5 with all amazing colleagues! It feels like a distant life in retrospect. All mist erased and I see only wonders. Look forward to the next chapter!
Show more
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.
Show more
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.
Show more
0
1.5K
26.9K
3.6K
Forward to community
We are hiring for @Harvey’s model training team. This team will help Harvey expand from the application layer into the model layer and from legal into high end knowledge work more broadly. We are hiring AI researchers of all seniority, particularly those with experience post-training frontier or open source models. Our program is centered around large-scale model training, synthetic data generation, long horizon reinforcement learning, and rigorous evaluation in real world deployments. We are scaling-pilled and believe that nothing beats the combination of larger models and better training data. We’ve been able to generate incredibly realistic legal environments and validated that this allows us to post-train open source models to achieve frontier performance with agents. We plan to scale up these data generation and training efforts significantly across legal to start, and eventually other verticals. As a researcher, you will have access to thousands of GPUs and unique training data from our product and customer relationships. Your research will inform Harvey’s product strategy and power AI used for some of the most economically and societally impactful work in the world.
Show more
We taught a brand-new mini-series this year at @SCSatCMU on Modern GPU Programming for ML Systems, as part of the ML Systems course, touching on fun questions like what data layout swizzling is, how to use 3D TMA, and state-of-the-art Blackwell programming. We released a curated online book based on the materials: check it out
Show more
0
24
1.8K
291
Forward to community
I'm amazed by this line from @deepseek_ai 's official announcement "Not lured by praise, not frightened by slander, follow the righteous path and discipline oneself with integrity" DeepSeek's real edge isn't just the tech, it's the ethos behind the work. And that's what will carry them further than any benchmark ❤️
Show more
0
29
1.6K
189
Forward to community
We keep striving to build things that bring long-term value to everyone. We hope you enjoy our latest model — try it now on web, app, and API 🚀 「不诱于誉,不恐于诽,率道而行,端然正己」
Show more
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at via Expert Mode / Instant Mode. API is updated & available today! 📄 Tech Report: 🤗 Open Weights: 1/n
Show more
0
1.7K
45.6K
7.5K
Forward to community
What makes this slide even worse is the little trick trying to hedge the language: the speaker first make an offensive quote specifically targeting a particular nationality, then make the note that tries to exempt herself from the responsibility, or equivalently “I didn’t say it, you said it yourself”. It is rather unfortunate to see language technique is used in this way 🙃
Show more