Thoughts on mlx-lm: top-priority is making it the central registry of MLX model implementations, with tools for evaluation and profiling, and we should add vision models too.
Inference engines can have their own schedulers and custom kernels, and do whatever hack to make inference ultra fast, while importing mlx-lm as a library of models.
We can rely on community to contribute model implementations, but there would be a fixed procedure to verify correctness of the implementation, ideally automatically.
Everything else except for critical bugs, should be irrelevant at the moment and I'm closing PRs and issues aggressively. Many people will be mad at this, and certainly I would be making mistakes closing legitimate things, but for the project to survive, and for the community to grow healthy, I don't see another way.