The last update already beat what we promised for some setups, and more is coming!
I'm testing the next one now: swap during model load will likely be lowered by over 90%, and long-prompt cache reuse will work better too.
Results as soon as validation ends ๐