Look at the speed on GLM-5.3-Flash
240 tok/s one stream, controlling it though litter which will have a lot of UI improvements
Keeping the name and cat but improving the fonts and performance
Show more
"Because you have so little faith.
Truly I tell you, if you have faith as small as a mustard seed, you can say to this mountain, ‘Move from here to there,’ and it will move.
Nothing will be impossible for you."
Show more
Opus-5.5 is getting 50-100% improvements on literally anything inference I have thrown at it.
I've never seen such fast progress
The frontier models are so obviously superhuman in capability and raw IQ compared to 99.9999% of human beings across such a wide range of cognitive tasks that you'd have to be either ignorant of what they can do or willfully obtuse about what's happening here to deny the reality.
Show more
2 sparks and you’re set. Grab em before there’s none left
We just hit under 100 Sparks left at Microcenters across the nation only 9 more stores have them in stock... Look at this freaking sell of!
Absolutely insane.
It looks amazing (: I finally got it to look good.
Getting so close to cracking this.
All exl3, based quant. This is the best timeline. Our computers can help us ascend
At this point OpenAI should work with me to get a GPT-OSS-2 model lineup bundled with Codex/ChatGPT app with license.
It's going to happen anyway, you can really reach a billion people this way, there's more memory to go around than we think.
Show more
Going to make local AI so easy and so good you won't want to use anything else. 🤗
The Omacom Foundation is becoming a premier sponsor of
@0xSero’s work on local AI for the next three years, through his company Sybil Solutions! We’re going to collaborate on making local models work beautifully out of the box on Omarchy.
Show more
Just a few more prompts bro it's going to be perfect
Intel Arc B70
- Qwen3.8-27B
- 256K context
- exl3 - 4 bits
- 16 sequences
- 60 tok/s up to 330 tok/s at 16 sequences
This costs 1600$
Very cool, yandex made a new base model.
At this point Claude/GLM/Kimi can write kernels making any quant run on any hardware
They can systematically discover and measure why kernels are slow.
They can manage fleets to fix every single thing physically possible.
Any hardware which can run something, will and fast.
Show more
20 billion tokens locally.
1.6B tokens a day reached!
50B tokens a month! at the rate i'm going that's 15,000$ in api credits blended of all the models I use at production cache hit rates..
4x 6000 is a no brainer if you have the budget, it'll recoup it's cost tenfold.
Show more
MiMo V2.6 Pro RL on 8 DGX Sparks: 17.8 → 68.3 tokens/s with DFlash.
3.8× end-to-end speedup across four long structured-output tests, versus the same runtime without speculation.
We’ve released the patches, setup and raw results:
Show more
👁️ you can preserve capabilities despite pruning to get the cost per task down.
Artisan pruning, retained 92% of BF16's recalls at just 21.4% of the original size (:
Damned if you do, damned if you don't.
The reality is for Huggingface to have larger ambitions they need funding, otherwise some companies can threaten them, or say hurt them.
Nvidia has done more for Open Source than most companies, and for open weights more than you know.
Show more
I interviewed DHH, the creator of Omarchy - The Agent Operating system + Ruby on Rails, one of the most important frameworks of our time.
He's a father, builder, race-car driver, and such an inspiring person.
One of the most positive people I've had the honor of meeting.
Show more