New in @berkeley_ai blog. Cao et al. create new kernel optimizer K-search and use it to auto adapt CUDA to MLX. Their attention translation measures 0.97x speed to Apple’s native implementation. Translate your code with their OSS tool. Available now.
🐻📄
New on @berkeley_ai blog. What actually affects the output of LLMs? New work by Butler et al. (NeurIPS 2025) introduces a model-agnostic method of measuring input, data, and component attribution, giving insight on what really steers generation.
🐻📄
Why and how do diffusion models memorize vs generalize? Can we have scaling laws for memorization? This is increasingly relevant scientifically and pragmatically (e.g. Sora 2).
🚨 Our new preprint "On the Edge of Memorization in Diffusion Models" addresses this timely question!
Now we can have self-refinement for VLAs in the real world (with the aid of a big VLM)! VLM critiques VLA rollouts and iteratively refines the commands to make it perform better.