Register and share your invite link to earn from video plays and referrals.

stevibe
@stevibe
LLM. Local AI addict. Building @BenchLocalAI DGX Spark workbench Builds things nobody asked for. Benchmarks things for fun.
Joined July 2009
1.3K Following    27.7K Followers
I gave 4 "Flash" models the same job: redact the PII in an HR letter using an agentic mask tool. One of them cost 630ร— more than another. And it wasn't 630ร— better. Results ๐Ÿงต ๐Ÿฅ‡ GLM 5.3 Flash $0.005, 15 tool calls Clean sweep. Overshot a bit (also redacted employee ID + last-4 of the bank account, which the test didn't require), but zero leaks. Half a cent. ๐Ÿฅˆ Qwen 3.8 Flash $0.013, 28 tool calls Clean pass, cleared its markers, and finished the job. ๐Ÿฅ‰ Gemini 3.8 Flash $3.10, 155 tool calls Passed... eventually. 73 reasoning turns, 3.9M input tokens, an add-mask / remove-mask loop that went on for minutes. Correct result, brutal bill. โŒ DeepSeek V4 Flash Vision $0.053, 31 tool calls Failed. Its own reasoning says "the mask needs to be wider" and "the text is still visible", then it stopped calling tools. Name, DOB, SSN, address, email all partly readable in the final output. Takeaways: > Cheapest model won. Not "won on value", won outright. > Tool-call count predicted cost better than model tier. 155 calls is where the money went. > Seeing the problem โ‰  fixing it. DeepSeek diagnosed its own failure correctly and still shipped it. > Total spend for all 4 runs: $3.18. Gemini was $3.10 of that.
Show more