Register and share your invite link to earn from video plays and referrals.

apolinario (poli)
@multimodalart
ML Engineer for Art and Creativity @HuggingFace
707 Following    16.4K Followers
updated Decision Index 0.1 → 0.2 🎯 better formula, +29 jev-like models, +21 benchmarks AutoJev-27B by @perplexity_ai CTO @denisyarats took the open lead, trailing jev by 0.8 points 🏆 come find the best model at every size, speed and use-case
Show more
Introducing the Decision Index 0.1 ⚖️ a rigorous leaderboard comparing jev with 30+ open weights decision models 35+ benchmarks. asking 130K questions to each model testing knowledge 🧠, automation ⚙️, understanding 🤔and even creativity 🎨
Show more
too many open jev claims and reproductions which one works? which one can you run on your laptop? i created this tracker that categorizes 38 artefacts (github repos, hugging face models) and separates signal from noise
Show more
Meridian by @ViggleAI just dropped on @huggingface an H3 based video model that allows you change the camera angle and timing of any existing video or image 🎥↔️ freeze it in time, or reimagine the shot 🏀 try on spaces
Show more
I've added 3 new models to the H3 Acceleration Arena for evaluation VDN-H3 (8 steps), Lightx2v 1.2 (8 steps), Alibaba TaoMate H3 (3 steps)
A global challenge to make antibodies for breast cancer. 💖💖💖 Check it out on @huggingface
Tencent just dropped AuK on @huggingface. It is the nano banana for audio: edit content, voice, emotional tone of any audio ▶️ on Spaces
zoom 256× into a photo, one 4× step at a time 🔍 OracleZoom approaches super-resolution by writing its own prompt at every level and re-rendering ▶️
Show more
This is my favorite kind of open-source result. We released H3 at 28 steps. Then the community started optimizing it. Now, after 11,850 votes in the H3 Acceleration Arena, several accelerated versions are statistically in the top group — with the current highest point estimate coming from @larryvrh’s 6-step H3 Turbo v4. 28 → 6 steps. Not from one team, but from an ecosystem. This is exactly why we open-weight H3. ❤️
Show more
New 1T+ parameter class model dropped: Nex-N2.5 Competitive with GLM 5.3, Kimi K3 and DeepSeek v4 (also 397B Pro and 35B mini variants!)
Meet Nex-N2.5 family, our latest open source agentic models. Mini (35B) and Pro (397B) bring multimodal understanding and computer use. Max (1.6T) is a text-only MoE model built for complex reasoning, coding, and agent workflows. Strong benchmark results in workflow automation and computer use: ⚡ Max scores 50.2 on AutomationBench v1.0.6—just 0.1 points behind Claude Opus 5 ⚡ Pro scores 56.4 on OSWorld-2, ahead of Qwen3.8-Max at 46.7 With continuous action, visual feedback, and self-correction, NEX-N2.5 can work across Blender and CAD tools, handle expense reimbursements, and even play PC games—turning visual understanding into sustained, real-world execution. 🔗@huggingface 🔗@modelscope 🔗 Github 🔗Website
Show more
Meet Nex-N2.5 family, our latest open source agentic models. Mini (35B) and Pro (397B) bring multimodal understanding and computer use. Max (1.6T) is a text-only MoE model built for complex reasoning, coding, and agent workflows. Strong benchmark results in workflow automation and computer use: ⚡ Max scores 50.2 on AutomationBench v1.0.6—just 0.1 points behind Claude Opus 5 ⚡ Pro scores 56.4 on OSWorld-2, ahead of Qwen3.8-Max at 46.7 With continuous action, visual feedback, and self-correction, NEX-N2.5 can work across Blender and CAD tools, handle expense reimbursements, and even play PC games—turning visual understanding into sustained, real-world execution. 🔗@huggingface 🔗@modelscope 🔗 Github 🔗Website
Show more
0
247
864
255
Forward to community
in just 4 days, @ArtificialAnlys Intelligence Index did substantial updates 4.1 → 4.2 → 4.3 it's been a bit hard to keep track for me, so i made this visualization
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
Show more
the folks @ViggleAI have been one of the most creative teams in machine learning. it's super exciting to see them open sourcing
Viggle Animate is now open source! 🔥 driving video + frame of what to animate = animated video. in 4 steps, H3 base ▶️ on Spaces
StreamTalk is here: audio in, body gesture out, and you get the pose skeleton (SMPL-X), so the 3D rig is up to you 🗣 ▶️ on Spaces
super cool to have this surprise guest at our all hands! excited for this next chapter for @huggingface + @nvidia
Super happy to share our intention to join forces with NVIDIA in a $12,930,300,000 acquisition 💛💚 10 years after starting Hugging Face, open-source AI is at an inflection point. Thanks to the community, we’ve shown that it can be a complement, and even an alternative, to closed-source APIs. But for it to happen at larger scale, it needs more compute, more support, more collaboration and more visibility. That’s why we went to talk to Jensen, who offered to do exactly that with us. In addition to doubling down on NVIDIA’s massive contributions to open-source AI (I called them the “King of American open-source AI” earlier this year), they’ve committed to strongly supporting Hugging Face and our mission while keeping the platform open, independent and compute agnostic. The founders and the team are all staying to keep pushing this mission forward. Together, we think we can make open source the default way to build AI, with the goal of empowering 100 million AI builders to own their intelligence rather than rent it. Excited about the next 10 years! 🤗🤗🤗
Show more
full write-up is out with everything open > the env, the hand-rated pool, the 3 runs and every painting they made
Genuinely surprised this Bbox IC LoRA for LTX 2.5 is not getting more attention, I'm really impressed 🔥 > Animated spatial control - objects follow bounding-box positions & sizes > Independent regional prompts - every object has its own description Try▶️
Show more
reference to video working super well!
Reference-to-Video is now live for MiniMax H3 Max on fal! This is an early preview that can achieve up to RTF=1 (and with references around RTF=1.5) on 768p, and we aim to improve the speed by 2x while slashing costs during this week. Quality wise it delivers the same performance you come to expect from H3 Max.
Show more
LTX Ripple is a new efficient IC LoRA approach for video editing ✏️🎞️ edit just the first frame and have the edits ripple into the rest of the video. lands perfect edits, incredibly fast! LTX 2.5 based ▶️
Show more