DeepSeek-V4-Flash-0731 (all 284B parameters) running on a single DGX Spark.
⚡ ~17 tok/s generation, 40 tok/s prefill
📦 One 80GB GGUF file, mixed IQ2_XXS/Q8 imatrix quant
🧠 256-expert MoE, 6 routed per token
🆓 MIT licensed
frontier model (kinda) on your desk 😎
🤗
Show more
Kimi K3 vs DeepSeek v4 Flash (0731) to design a landing page for a $10,000 luxury phone.
Same prompt, same harness.
it's actually impressive to see DS4 close to K3's results considering it's weight. it's also faster and cheaper to run.
K3 won, of course, but the difference is not huge.
Show more
DeepSeek v4 Flash (0731) vs Kimi K3 (10x more weights)
i gave both the exact same prompt, in the exact same harness: one-shot a luxury product landing page ($10,000 "ARC-01" concept object).
results in the video. my take:
Kimi K3: wins on polish. Better typography hierarchy, more confident use of negative space, and the hero product shot integration is genuinely beautiful
DeepSeek v4 Flash: very close. The structure is right, the copy is right, the layout logic is right. It loses on finish, not on fundamentals.
a 10x smaller model getting ~90% of the way there. a year ago this gap would have been embarrassing.
mind you it cost less than $0.10 for 6 build iterations and 8 QA passes.
small models aren't "good for their size" anymore. they're just good.
Show more
DeepSeek v4 Flash (0731) vs Kimi K3 (10x more weights)
i gave both the exact same prompt, in the exact same harness: one-shot a luxury product landing page ($10,000 "ARC-01" concept object).
results in the video. my take:
Kimi K3: wins on polish. Better typography hierarchy, more confident use of negative space, and the hero product shot integration is genuinely beautiful
DeepSeek v4 Flash: very close. The structure is right, the copy is right, the layout logic is right. It loses on finish, not on fundamentals.
a 10x smaller model getting ~90% of the way there. a year ago this gap would have been embarrassing.
mind you it cost less than $0.10 for 6 build iterations and 8 QA passes.
small models aren't "good for their size" anymore. they're just good.
Show more
deepseek released the v4-flash-0731 weights!
spin up your GPUs, let's get testing
you're supposed to announce a reset after that
y'all have lessons to learn from
@thsottiaux
Two incidents over the last 24 hours resulted in elevated errors or reduced availability on Claude for some users. Multiple separate network failures cut into our capacity to serve Claude, and some requests failed while we rerouted traffic.
Show more
they are following
@NousResearch's portal by cutting prices
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
Show more
interesting experiment using the same model on different harnesses:
Claude Code:
→ 5.664M average tokens
→ 7.23% precision
→ 28.9% recall
→ 13m06s
Alibaba's Open Code Review:
→ 385K average tokens
→ 33.9% precision
→ 20.0% recall
→ 1m23s
Open Code Review ends up producing fewer false positives while consuming a fraction of the tokens
Show more
We ran Kimi K3 through 3 agent harnesses (Claude Code, Hermes, Kimi Code) on 28 identical tasks.
All 3 harnesses completed the tasks at similar success rates, but the interesting story is token efficiency: the same task cost up to 30x more tokens depending on the harness. 🧵🧵
Show more
quick DeepSeek V4 Flash benchmark on one DGX Spark:
→ Code: 16.79 tok/s
→ Prose: 16.77 tok/s
setup: 86.34 GiB Q2 target, ds4 runtime, greedy decode, 1,024-token context, three measured runs.
increasing concurrency failed to improve aggregate tok/s in a separate run. instead, latency increased a lot
i'm working on improving those numbers: profile the active bytes for every generated token, identify the tensor families dominating memory traffic, and attack the bandwidth floor
every candidate gets a strict control, output checks, and an intelligence eval before it earns a speedup claim.
Show more
you can collect robot training data without robots now
HiFi-UMI records humans performing tasks with a handheld gripper, then uses that data to post-train policies that deploy directly on real robots
Across 3 model families, it matched robot teleoperation within 3.1 percentage points
Best result: 85% success on precision insertion.
it's just a matter of collecting accurate human demonstrations now.
Show more
opus 5 one-shotted this LiDAR horror game
you can only perceive the world through LiDAR signals, the rest is pitch black.
honestly impressed, the world emerging point by point is ridiculously satisfying
i didn't manage to finish the game tho i'm too scared
Show more
what a time to be alive
Catching skin cancer early is a home robotics problem.
Melanoma is highly treatable when detected early, yet today’s screening process depends heavily on patients noticing tiny changes across their entire skin surface. This requires patients to solve a near-impossible visual-memory and registration problem.
I built OpenDerm, an open-source 4-DOF robot that captures high-resolution images of the skin and uses them to reconstruct and track the skin surface in 3D over time.
The best way to make skin screening truly routine is to bring it into the home. OpenDerm shows that inexpensive robotic skin imaging is possible, but the path to scale is not a dedicated screening robot in every household—it is to make skin screening one of the many useful things a general-purpose home robot can do.
Read more about why I built OpenDerm and how it works here:
Blog:
Project:
Show more
someone check on claude
Try sending “see the below —“ to Opus 5
It appears to generate a user completion rather than respond 🤨
We saved $260M in cash for cos. like Meta, Unilever, and J&J by finding baseless charges across 19M+ invoices.
Now, we've raised $75M to extend our guarantee: if we can't find $500K in overpaid invoices, we'll pay you $10K.
Book a demo:
------------------------
How it works ⬇️
PROBLEM:
A Fortune 500 gets a million invoices a year. Most of them are for a couple hundred dollars.
Checking one properly means opening the contract, pulling the rate card, matching the PO, and hunting down the bill of lading that proves the shipment moved the way the carrier says it did.
All of that work to defend $200. So nobody does it.
The pile goes to a human. The human audits only a sample (12%) and pays the rest (88%) blindly. That's where the money goes.
At a pharma co, we found an uncontracted surcharge billed across global carriers for two years. $26M nobody questioned, because every individual line looked ordinary.
Each surcharge is small enough that escalating it costs more than paying it. And there are a million of them.
------------------------
SOLUTION:
Freehand’s AI reads everything: every line of every invoice, against every contract, rate card, transaction data, bill of lading, warehouse record, time sheet, and email exchange between you and the supplier. It also analyses every past invoice, and transaction data with that supplier,
ever.
It holds all of this knowledge in a Context Graph. When an invoice arrives, it knows what you should pay, what actually shipped, if the service was delivered or not, what the SLA was, and what this supplier billed you for the same service last quarter.
Then it proactively:
> writes up the dispute with the evidence attached
> emails the supplier
> calls her when the email goes quiet
> Slacks your purchases team for more information
> Follows up until the invoice comes back corrected.
> Approves the correct invoice
> Pays it across different currencies and tax structures
> And accrues the right amount to your general ledger
Freehand’s will do all of that to recover $20, because it is doing it across all millions of invoices at the same time, saving 5-10% of company spend with 100% SOX compliance and auditability.
-------------------------
🚨 RT + reply "FREEHAND" and we'll send you the AI upskilling guide that's already helped 750+ displaced workers move into more secure, higher-paying AI-native roles through Freehand's Transition Bootcamp.
Show more
i don't blame models for not scoring on this benchmark, even I would probably fail miserably
GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games?
We investigated. The harness was not letting it remember what it had learned.
We found that enabling two API settings tripled our scores with 6x fewer output tokens.
Show more
i gave Opus 5 the same prompt (one-shot)
for some reason it decided to restrict itself by creating a self-contained HTML file.
even with that, the results are convincing. the landing page layout seems cleaner and the 3D is working just fine
i'd say it's on par with Kimi K3 on this exercice, but it would probably have been better if Opus hadn't restricted itself
Show more
Kimi K3 vs GPT-5.6 Sol (one-shot)
GPT-5.6 gets mogged pretty hard, Kimi K3 outperforms by far.
Again, Kimi's taste is better and its 3D capabilities are way cleaner.
One thing that I notice is that Kimi seems to be less "lazy". It keeps on iterating and improving its design and does not content itself with the bare minimum.
Show more
opus 5 built a tiny paris inside the browser
the eiffel tower, champ de mars, a chunk of the trocadéro, the lighting: all generated as a browser-native 3D scene you can explore.
one-shot. i just added a follow-up prompt to add the cinematic path
i gave the same prompt to Kimi K3, i'm waiting for it to finish
Show more
i gave Kimi K3 the same prompt
the eiffel tower is less impressive and the surroundings do not really good either.
it generated a better park than Opus 5 tho so prompts to K3
it's still an impressive output don't get me wrong. it just feels a bit underwhelming comapred to opus
also one-shot.
Show more
not sure it would fit better than GLM5.2 in Sparks but nice progress
Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio + 128GB RAM device.
Kimi K3 is the strongest open model to date.
Guide:
GGUF:
Show more
i gave Kimi K3 the same prompt
the eiffel tower is less impressive and the surroundings do not really good either.
it generated a better park than Opus 5 tho so prompts to K3
it's still an impressive output don't get me wrong. it just feels a bit underwhelming comapred to opus
also one-shot.
Show more
opus 5 built a tiny paris inside the browser
the eiffel tower, champ de mars, a chunk of the trocadéro, the lighting: all generated as a browser-native 3D scene you can explore.
one-shot. i just added a follow-up prompt to add the cinematic path
i gave the same prompt to Kimi K3, i'm waiting for it to finish
Show more