Fast H3 implementation for Metal. Enjoy, modify, and so forth: Contains code from
@liuliu which is welcomed in taking back whatever parts he likes for
@drawthingsapp in case there are H3 plans there.
Show more
Fable - Sol loops found another 23% memory reduction in the already very optimized sorted set PR. Now, let's explore radically different approaches compared to the zkiplist+hash approach.
I believe having to split large model files is an anti-pattern, and could be definitely abstracted away by Hugging Face. I look forward to changes in the way we upload and download large models, as models are headed to be ... larger :D
Show more
I reviewed a Redis me optimization and was about to submit the PR. Then I said myself: "wait, we are in 2026" and wrote a script that calls Sol and Fable at each turn, challenging to improve the work of the previous one... Soon to be continued with the results.
Show more
That's a bit slow. Streaming directly the K3 official hugging face 1.6TB of weights in mxfp4 in an m5 max 128gb.
The real AI risk is inside the labs (a reply to Amodei's post on open weight models):
Hint: min-p 0.05 (never accept tokens < 5% score of top one) recovers a lot of quality in Laguna S2.1 without impacting output richness in any significant way. min-p should be the default for most inference engines.
Show more
Two days of testing of Opus 5. My take based on insufficient observations, but this is the best I can do: 1. Anthropic no longer has the incredible Opus 4.8 / GPT 5.6 Sol gap. 2. The model is not clearly superior to Sol. 3. In chat, the model vibe is not great.
Show more
I tried to adapt the transcript of my video "Being Linus Torvalds" into an english language blog post here:
I bet that we will discover that this new capability does not proxy to a generally stronger model for other activities, but emerges from the fact the model learned a better representation for problems that otherwise could be already approachable by Fable / GPT 5.6 Sol.
Show more
Claude Opus 5 from
@AnthropicAI is the new SOTA on ARC-AGI-3: 30.2%
The previous high score (7.8%) was set by GPT-5.6 Sol (Max)
Throughout our analysis, we observed novel behavior that allows Opus 5 to solve previously unbeaten environments, outperforming Fable
Show more
Warning my friends: for the sake of saying that distillation is a good thing, and not a bad one, you are falling in the trap of admitting that frontier Chinese models are *mainly* the result of distillation (which is not just not true, but also not possible).
Show more
All right. The world also needs inference hardware to run and train such models at a fair price.
For my first post, I’m sharing a letter
@NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
Show more
Quality of Opus 5 is crucial for Anthropic. Anthropic bad times didn't start with Fable, but with Opus 4.7-8 terrible quality. Opus 4.6 was a great general purpose model, but it too struggled at coding VS GPT o the same time. Opus 5 will be "fix or break" in pre-IPO times.
Show more
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
Show more
Update on the DwarfStar laguna-s2.1 branch, the GGUF files got updated, now the support for the file with Q8 projections and smaller routed experts (68GB total IIRC) is implemented as well. Speed improved significantly to 60 t/s generation, 550 t/s prefill.
Show more
I have a modest proposal: let mathematicians be the ones to crack the open problems that can be addressed with LLMs. They worked for a long time in this field. It should be their prerogative to do so.
Btw if you think the argument against Chinese models is too weak to ban them internally, in the US, you should check the status of Heath Care and costs for American citizens of things that cost 1/10 elsewhere in the world. They can do it, especially with this administration.
Show more
Laguna S2.1 is, among the other things, completely unable to write correct Italian. Something that even much smaller models can do very well. I understand specialization but this is a red flag. I saw this with Q4 quants, verified with the official API (via openrouter).
Show more
This "they distilled our models!" is starting to resemble the moment when Redis was winning in the database space really hard, and certain actors started with "but it is not linearizable!", and also attacked me personally. Eventually, a few started selling Redis, before closing.
Show more
Still didn't attempt any quality testing, but let's say that GPT 5.6 Sol was able almost unassisted to write this Laguna S2.1 inference implementation following the other two models in DwarfStar, and it is very fast, 50 t/s generation, very fast prefill as well.
Show more