prism cooked so hard with Ternary-Bonsai-2-27B, utterly insane how good it is for the size
have so far converted the mlx to vLLM, added MTP + some custom kernels 🍿
aiming to create the best local LLM for those with ~10-32GB VRAM NVIDIA cards
image/weights coming soon... 👀
Self-replicating, multi-agent AI systems are an overlooked existential risk.
I want to lay out why I believe this, via example.
Each dot is a simple creature with a randomly evolved NN brain.
There is no training; yet complex, unpredictable survival drives always emerge.
1/14
DeepSeek 4.1 Flash full precision + DSpark @ 200TPS on 4 Max-Qs- with just 64GB system RAM!
Offloading the 200GB Engram (i.e. a hash table) to NVMe seems fine rather than keeping it all in RAM, my wallet is thankful 🙏
More juice to squeeze though, vLLM recipe to come... 👀