I sent Cursor Grok 4.6 on a mission to make me a Qwen 3.8 46B A13B MoE by merging two 27B models together. It has currently been running for 2 days and I only have 1 spark so it’s been a little slow and a few hiccups here and there but so far it’s going well, i can’t wait to see what the model is actually like when it’s done (it’s my first time doing something like this so it’s genuinely interesting to read and see what its doing)
The architecture is labeled as “Backbone-Anchored Layerwise Expert Selection Hybrid” or BALESH (name is iykyk)