i forced myself to use a few non-mainstream models today. sharing my experience with everyone in case you're wondering
1. muse spark 1.3
i was genuinely surprised by how good this model is. if you silently swapped opus 4.8 with this without telling me, it'll probably take me quite a while to figure it out, and it would likely be from the communication style rather than capability
at its contributor pricing, the ROI is quite incredible. i find very little motivation to even try deepseek v4 flash when this exists
i don't like the phrase they keep using though - "intelligence that's too cheap to meter". if i do all my work with this model, even at its extremely good pricing, it'll still cost me thousands of dollars a month - that's not "too cheap to meter"
reduce the cost by another 100x then let's talk about "too cheap"
2. glm 5.3 flash
i thought it'd be better than muse spark 1.3, but actually immediately after i switched to this as my firstmate, it made quite a few mistakes
it could be a bit anecdotal but now it lost trust with me and i'm a bit nervous about letting it manage my work. i'm going to let it do some more straightforward implementation rather than acting as my primary model
3. gemini 3.8 flash
maybe it's because i never spent much time with gemini before, but this model is SO GOOD at communicating. it's such a breath of fresh air. i can understand every word without using much brain power at all
but this is significantly more expensive than the two above. i think i'd have to get a google ai subscription if i want to use this more, but not being able to use it in 3rd party harnesses gives me a pause
it's also powerful enough as an opus replacement for most things i tried. my general sense after trying these models is - everyone has an opus now, and they are all cheaper than the real opus
fable and astra is probably the only moat the labs have now - the game has fundamentally shifted from what it was 6 months ago