Register and share your invite link to earn from video plays and referrals.

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
Joined January 2013
637 Following    2.4K Followers
GLM-5.3-Flash, 320B params, one DGX Spark, 256K context, 33.8 tok/s peak. Released 2 days ago Serving recipe with MTP speculative decode already up, incl. the branch pick that makes it 2x faster:
Show more
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: GGUF:
Show more