Model-specific inference engines are basically what everyone autoresearches and hill-climbs the hell out of for TPS after each Qwen release. In the early days I worked with MLX too, but these days I just trust oMLX to do the job.
Meet Husky: a Model-Specific Inference (MSI) engine up to 4.5× faster than Apple's MLX
Woof, Underdog's Pareto frontier model, now runs up to 730 tokens/sec on a MacBook
Finally local models are as fast & capable. Try it now in - your personal private AI