Wow, two years anniversary on a $1,000 finetune that beat Claude 3.5 Sonnet (by using their API and not telling people lmao)
Same guy that's baiting you all about how Local AI isn't good btw 😂
Show more
I'm excited to announce Reflection 70B, the world’s top open-source model.
Trained using Reflection-Tuning, a technique developed to enable LLMs to fix their own mistakes.
405B coming next week - we expect it to be the best model in the world.
Built w/
@GlaiveAI.
Read on ⬇️:
Show more
Bait used to be believable
We’re now in the copy the HF quantization repo and put your name on it as an expert phase of Local AI
That’s not how Opensource as a community flourishes
You might have good intentions, but please do right by others & lift them up with you rather than falsely crediting yourself
Show more
Normalize shaming people that steal others work and put their name on it
Almost passing
@BrandonMusicKy in downloads.... and over half the likes for just copying the model....
Data center hardware prices will go down
Consumer hardware prices will go up
This GPU-based class system won’t last forever thanks to Opensource btw
Get used to a Local AI where everyday your tokens per second gets better and better
The optimizations in MLX are horrid, DHH is right Apple doesn’t have a hardware problem they have a software problem
Faster chips, new bottleneck 🤯
Two clustered M5 Ultra Macs ran SLOWER than one: 35 tok/s vs 53. Which is strange, considering my experiments with M3 Ultra were the opposite.
Keep the GPUs awake and they jump to 58 tok/s.
Show more
Gentle reminder that this site can turn you cruel if you’re not careful
Remember, on the other side there is a person, you gotta do right by them from wherever you're standing, however big or small you are and the people you're talking to are
Stay kind, stay human
Show more
Most notable thing
M5 Ultra with 1.2 TB/s bandwidth is only 6.4% faster in token generation than the DGX Spark (273 GB/s)
M5 Ultra has ~5x the Bandwidth of the DGX Spark
Mac Studio M5 Ultra vs 2x DGX Spark
for DeepSeek V4 Flash
- DGX Sparks: 2.41x faster on prefill (compute-bound)
- Mac Studio M5 Ultra: 6.4% faster on generation (bandwidth-bound)
Time to first token
- Spark 1.3s
- Mac 2.9s
Of course, the Mac M5 doesn't have optimized kernels
Show more
M5 Ultra vs dual sparks on deepseek-v4-flash. Disag still on the table
Still a week left to show up and participate in our Opensource Contributors program for September 💪❤️
There’s a difference between having a Kofi link to support your Opensource contributions and begging btw
We even have an Opensource Contributors program at
@OsmanticAI where we spend thousands of dollars a month to support our Opensource Contributors ❤️
Show more
Wait, why is Zach not monetized? This guy actually posts real alpha, gives free GPUs credits to people doing interesting work / doesn’t beg for money, and actually does his homework before he talks so it’s not nonsense or stolen
Go follow this guy and give him some views
Show more
Every day I try and write carefully crafted tweets and then one random post gets like 150k views.
Not sure what to do with this information but almost to the point I can get monetized so that’s neat
Show more
Open systems always end up winning because openness is freedom and we all are meant to be free
Opensource AI Must Win
It’s one of the great ironies of our time that the people building artificial intelligence are often the ones who understand the human condition the least.
They can teach a machine to write a symphony, but they forget they can’t teach it what it feels like to sit alone at 3 am wondering if things will ever work out or how to find the strength to move forward after losing something you loved with all your heart despite all the logic in your head telling you to give up.
Being human isn’t something you can compute. It’s something you survive. You have to live, love, and lose. AI might be able to reason faster, remember more, and never get tired… but that’s not what the essence of who I am or who anyone is. AI isn’t replacing me or anyone else because we are more than just some line item in a spreadsheet that justifies our economic value. We are people with lived experiences, everyone is someone’s child, someone’s friend, someone’s love, and these are all bonds that form organically.
So although I’m deeply optimistic about the potential positive impact AI will have on the economy… I’m still going to be massively long humans.
Our resilience is unmatched and whenever you would have counted us out and bet against you would have lost. We’ve survived ice ages, plagues, wars, and our own worst instincts, and every single time we came back stronger. Our success as a species in such a short time on this earth is something we should all be proud of. Humans were always going to eventually build an AI, but no AI was ever going to build a human, no matter how much resources you gave it.
So rest easy guys, we still got this. ✊
Show more
i don't think nearly anyone, myself included, has truly internalized there'll soon be a machine better than us in every single intellectual & physical capacity
deep down we all share a feeling that our 'entrepreneurship' or 'taste' or... is special & safe. it isn't. it's over
Show more
Anthropic is the first SaaS provider with a twist (Sabotage as a Service)
@TheAhmadOsman Ain't no way my kernel generator project was willingly sabotaged by this dogshit?
Since I switched to GPT models for it, it's been going MUCH better
How does a decision model like Jev work
Think of it as a tree of probabilistic decisions - a bunch of answers, each with scored confidence
You run Kernels not models
The model is just a graph
The Inference Engine is scheduler / optimizer / executor
But the actual work? That happens in the Kernels
- MatMul Kernels
- Attention Kernels
- RMSNorm Kernels
- KV cache Kernels
- Quantized linear Kernels
- Sampling Kernels
- Fused “please don’t write this back to memory 9 times” Kernels
Same model, same GPU, same VRAM
Wildly different performance
Because one stack is using optimized fused Kernels that understand your hardware
And the other stack is playing hot potato with tensors through 47 tiny launches and pretending the GPU is the problem
Bad Kernels make people say:
“this model is slow”
Good Kernels make people say:
“wait how is this running locally?”
This is why Inference Engines and the Kernels implemented within them matter
The model is the recipe
The hardware is the kitchen
The Kernels are the knives, pans, burners, and the chef not cutting onions with a spoon
Most people benchmark models
The real ones benchmark the Kernels underneath
Show more
I have previously said this: You don't pick an inference engine first. You pick a hardware strategy, a workload shape, and a serving model. The engine follows.
But there is one more layer under it.
You also do not pick a model and a GPU and call it done.
You pick a file encoding, and a kernel path.
The GPU follows those.
Start with the loop.
Text becomes tokens.
Tokens move through a Transformer.
Attention decides which earlier tokens matter.
The runtime keeps a KV cache so the model does not recompute the whole conversation every time.
Then it picks the next token and does it again.
The model is not writing a whole answer in one shot.
It is generating one token at a time.
That loop has two phases, and they are not the same job.
Prefill reads the prompt and builds the first KV cache.
It is compute-heavy.
That is the pause before the first word.
Decode writes the answer one token at a time.
It keeps rereading weights and the cache.
It is bandwidth-heavy. So an RTX PRO 6000 with 1.8TB/s beats the DGX Spark with 273GB/s.
That is the typing effect you feel.
Long prompts punish prefill.
Long answers punish decode.
Long chats punish both, because the working memory grows.
The inference engine is not the model.
It is the traffic cop, the memory manager, the kernel dispatcher, the scheduler, the cache accountant, and the API surface.
It loads the weights.
It tokenizes the input.
It runs the forward pass.
It samples the next token.
It keeps the KV cache.
It streams the result.
Serious engines also pick kernels.
A kernel is not "the model."
A kernel is a specific tensor program: shapes, layouts, datatypes, and what the silicon is actually allowed to multiply.
Same math on paper.
Different contract.
Different kernel.
The file format matters.
It decides what can load, what can quantize, and how fast it runs.
Quantization is not one switch.
Storing weights in 4-bit is not the same as doing 4-bit math.
Weight quantization shrinks the model.
The live context is a different thing.
A Q4 sticker is not universal.
The right format is the one your engine has optimized kernels for.
Assuming every quantization label is portable is how people buy a 5090, download "NVFP4," and still miss out on performance.
Here is the worked example.
I put Qwen 3.8 27B on an RTX 5090.
Two downloads.
Same model.
Same GPU.
Both folders said NVFP4.
One of them does 4-bit math for real.
Weights in 4-bit.
Activations in 4-bit.
The matrix unit can multiply them as 4-bit times 4-bit.
The other stores 4-bit weights.
Then unpacks them in the kernel.
Then does 16-bit math.
Same sticker.
Different kernel path.
The 5090 did not choose that.
The checkpoint did.
Especially whether the activations are 4-bit too.
If they aren't, there is no legal 4-bit times 4-bit multiply to run.
The engine falls back.
Quietly.
The file still loads.
The logs are easy to miss.
That is the difference between an encoding you can load, an operation a backend implements, and arithmetic the hardware actually executes.
They sound interchangeable.
They are not.
This is also why prefill and decode do not get the same gift.
Prefill has enough token rows to keep the matrix units busy.
Native 4-bit math can matter there.
Decode is still walking the weights and the cache, one token at a time.
If the live state never went 4-bit, you do not get a 4-bit win on that part.
The bottleneck moves.
Do not benchmark "the model."
Benchmark the stack you will actually run.
Separate prefill from decode.
Pin the exact file, the exact engine, the exact GPU.
One more thing people get wrong with the word Blackwell.
It is a marketing name on four chips that cannot run each other's kernels.
A data-center B200 is not a 5090.
A 5090 is not a Spark.
A Spark is not a Thor.
Instruction support is necessary.
It is not sufficient.
A checkmark on the spec sheet is not the kernel that ran.
In the inference engines article under my profile, I asked a question I still want on the wall:
What quantization format has optimized kernels on my target engine?
This Qwen run is that question in a box.
Same GPU.
Same NVFP4 label.
Different kernels.
The engine followed.
The checkpoint decided which kernel it was allowed to launch.
The file loaded but that is NOT the same as the kernel you wanted for the GPU you bought.
Show more
Why do you think Rust works so well for LLMs? Same idea
The biggest impact of mathematics won’t be the millennium problems but proofs in general software at the scale.
I think all mainstream software will be verified by 2030. You likely won’t touch an unverified library.
Some of the biggest hurdles such as formalization of specs and doing this at 1B LoC scale are the ripe targets for auto research/self-improvement loop.
The reason we can say this confidently is because either AI will get paused or the world will end if this doesn’t happen.
Show more