If LLM-as-a-verifier works as well as my last post suggests, this is going to be massive for local AI.
Like the dev pointed out, pairing a large model with a small model lets you run verification on the cheap.
This means if you combine the GLM-5.3 API with Deepseek-V4-Flash running on a DGX Spark, you get a huge jump in performance without driving up costs.
If you have a 4xDGX Spark, you could multi-batch both GLM and Deepseek on a single node and beat Fable entirely offline without touching an API.
Welcome to the era of local AI.