Microsoft just released Mage-VL on Hugging Face
A codec-native, proactive-streaming multimodal foundation model for image and video understanding.
JarvisHub: Open Canvas-Native Multimodal Creative Agents
Unlike one-shot prompt-based tools or chatbot agents, JarvisHub turns an editable canvas into a shared project state where users and agents collaborate. Natively supports images, videos, websites, and presentations.
Show more
NVIDIA just released JEPA-DNA on Hugging Face
A new genomic foundation model that combines generative pretraining with a joint-embedding predictive architecture for better functional understanding of DNA.
Show more
Most upvoted papers on
@huggingface this week
Harness Handbook - Making evolving agent harnesses readable and editable
LongStraw - Long-context RL beyond 2M tokens on fixed GPU budget
Weak-to-Strong Generalization via Direct On-Policy Distillation
Boogu-Image-0.1 - Open-source unified multimodal understanding and generation
VideoChat3 - Fully open video MLLM for efficient generalist understanding
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agents
ABot-N1 - Toward a general visual language navigation foundation model
Find them below:
Show more
DeepSeek released DSpark
A speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification to accelerate LLM inference
NVIDIA just released Audex on Hugging Face
A unified audio-text MoE that handles speech, sound, and reasoning
without regressing on text intelligence.
30B params. 3B active. 1M context.
Show more
Tencent just released Hy3 on Hugging Face
295B MoE,
21B active.
Rivals 2-5x larger flagships on reasoning,
coding, and agents.
256K context,
with anti-hallucination improvements
and stable tool calling.
Show more
Paper:
Project page:
Hugging Face:
Code:
PhotoQuilt makes training-free photomosaics at any resolution
Bootstraps a global layout at low res, then upscales and re-noises tiles via Black Forest Labs FLUX. Each tile denoises into its own image while the full scene stays coherent. Scales past 14K without quadratic attention cost.
Show more
This week on Hugging Face Daily Papers: world models, agentic abstention, and code without containers
- Orca: The World is in Your Mind — a general world foundation model using Next-State-Prediction to unify text, image, and embodied action generation
- Agentic Abstention — do LLM agents know when to stop acting? A new method improves timely abstention without any fine-tuning
- Dockerless — environment-free program verifier for coding agents, scoring 62% on SWE-bench Verified without containers
- Scaling the Horizon, Not the Parameters — how a 35B agent reaches trillion-parameter performance through long-horizon scaling
- LiveEdit — real-time diffusion-based streaming video editing running at 12.66 FPS for interactive AR apps
- Program-as-Weights — compile fuzzy functions into tiny neural artifacts that match 32B model quality with 50x less memory
- BlockPilot — instance-adaptive speculative decoding for diffusion models, achieving 4.2x speedups
- DOPD — Dual On-policy Distillation that fixes privilege illusion during student-teacher knowledge transfer
- Does VLA Even Know the Basics? — measuring how much commonsense and world knowledge VLMs lose when becoming embodied agents
- Formalizing Latent Thoughts — four axioms revealing that LLM latent representations may encode far less reasoning than we assume
Show more
GEAR: 10× faster autoregressive image generation
Tencent Hunyuan's new method jointly trains VQ tokenizers and AR generators end-to-end, beating LlamaGen-REPA with a novel dual read-out. All tokenizers are on Hugging Face.
Show more
DART: one-shot VLA adaptation under environmental shifts
Seoul National University researchers show that weight space arithmetic can isolate domain shifts from task knowledge, letting you adapt a robot policy to new cameras or embodiments with a single demonstration.
Show more
ByteDance Seed just released PAR on Hugging Face
A new model checkpoint.
Apache 2.0 license.
Ready to explore.
Google just released TabFM on Hugging Face
A tabular foundation model
for structured data.
Do LLMs actually think in silence?
A new paper formalizes four axioms — Causality, Minimality, Separability, and Stability — to audit latent thought representations. No open-weight LLM satisfies all four. Most representations encode nothing beyond the input embedding itself.
Show more
Top papers on Hugging Face this week: Agentic AI, world models, and real-time streaming
- MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision
- Qwen-AgentWorld by Alibaba: Language World Models for General Agents
- Wan-Streamer v0.1 by Alibaba: End-to-end Real-time Interactive Foundation Models
- Are We Ready For An Agent-Native Memory System?
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
- OpenRath: Session-Centered Runtime State for Agent Systems
- Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- DanceOPD: On-Policy Generative Field Distillation
Discover them on Hugging Face Daily Papers
Show more
V-Zero: answer-label-free visual reasoning
It uses on-policy distillation with contrastive evidence gating.
It trains 5x faster than SFT and 10x faster than RL.
4B model is on Hugging Face.
Show more
EvoEmbedding
An evolvable embedding model that maintains a latent memory queue to generate dynamic representations for long-context retrieval, outperforming static specialists 3× its size.
Show more
CLI-Universe
A principled engine that synthesizes verifiable terminal-agent tasks grounded in real-world materials.
Qwen3-32B fine-tuned on just 6K trajectories hits 33.4% on Terminal-Bench 2.0, beating models 10x larger.
Show more
Just added Lilian Weng's blog as a recommended read to the Scaling Laws method on Papers with Code
Find the original paper + all papers which cite it, here:
Show more