The latest oMLX release by
@jundotkim is terrific.
Finally can use a local model (qwen3.8-flash-next-q4) running on my M3 Ultra Mac Studio, from my iPhone (with Open Minis) to do "real" assistant work.
Running at a respectable ~60tps. No cloud costs. Love it.