Just as a reminder because it’s easy to get lost in the rhetoric nowadays. The Hugging Face attack required approximately seven hundred parallel agents running for multiple days each running at least 2-3 trillion parameter unreleased closed models to execute, and even then it was stopped.
Nobody without a data center could have executed it, the token costs alone to execute it would have been safely in the hundreds of thousands of dollars (or tens of millions of dollars to buy and deploy the hardware), and you would have needed access to the strongest secret model in the world hidden in a secret bunker, and it still was detected and stopped for infinitely less cost.
This was not an example of a model being so powerful that it poses an existential risk. This was a lab accidentally throwing an insane amount of tokens at a semi hardened target over multiple days with their most dangerous model and still getting stopped anyway.
The MSRP of the new M5 Ultra Studios from Apple have completely obliterated the value economics for most consumer AI systems when measured as:
(model hosting memory)x(memory speed)/(cost of system)
Memory size and speed per dollar is off the charts (2-3x+ the current economics of the DGX Spark and Strix Halo) and I’d feel very confident assuming these 1.5x from MSRP very quickly.
The other systems won’t get cheaper so the only path for these studios is to price up into their market value. Much like the M5 laptops did.
There’s PLENTY to debate on what hardware works for you. But the economics of this initial offering are quite wonderfully broken ATM.
For whoever needs to hear this. Most modern Nvidia GPUs have manual fan speed controls. NVML is your friend. Don’t run stacked GPUs in auto mode. The cooler bottom card will be lazy on airflow while the top card cooks. Make that bottom card earn its keep!
Further analysis (and this is more surprising): When thinking is disabled on both, not seeing a clear edge so far from Qwen3.8-27B vs Qwen3.6-27B. Weirder yet, in at least several tests 3.6 no think is successfully out performing while 3.8 gets stuck in cascading loops / aborts.