Register and share your invite link to earn from video plays and referrals.

Kyle Hessling
@KyleHessling1
Father | Local AI Infra Engineer | Striving to be like Christ
730 Following    7.7K Followers
Qwen 3.8 27B at 56tps; on 9 year old GPU btw Nvidia V100 32GB ~$650 on EBay right now! Using Dflash 2; disabling the ECC adds some more speed too! Thinking and prose is a bit slower, but 56-63 tps in code gen! MTP runs faster for prose vs DFlash2 but slower sustained code generation speed. MTP also runs much faster power limited than DFlash does. Working on a repo so you can get up and going quickly. Fun fact, the Nvidia v100 was $11,500 per card when they first launched. Price you pay for future proofing I guess; they’re still great cards. Pcie 3.0 and the older software/architecture are the only drawbacks, but also those aren’t as much of an issue as you’d think. Especially when you consider the price today!
Show more
Alright guys, wow. Qwen 3.8 27B is basically Qwen 4. What we have here is one of the best one-shot shark survival games I've made with any model. I ran this same prompt through fable, and the biggest thing all models struggle with in this test is the top-down view of the sharks. Fable even had all of the fins reversed and details jumbled on the first turn. Before now, it was really a 2 to 8 turn test to get something final regardless of model. Qwen 3.8 27B just nailed this in one output, no harness at all, just purely the model thoroughly thinking before outputting the final result. On every asset used, different shark variants, everything, it all looks incredibly tight. The boat movement, shark speed and engagement, everything is so much better than I expected, and I was expecting a lot. Imagine what 8 turns would do in a harness? This confirms to me that while Qwen 3.8 27B thinks a lot, for the first time in a local model on the base weights, IT MIGHT ACTUALLY BE WORTH IT This is a Frontier Lab model result, and it ran on my 5090 at 60tps in Q5_K_M. I am just in complete astoundment at what @Alibaba_Qwen has accomplished here, and I am incredibly grateful that this model is open source. The speed will also increase drastically with optimisations. This is just an incredible day. What a blessing. More to come!
Show more
Hello again, everyone! We've got another really fun 9b, this one specifically trained for tool calling and agentic coding workflows in @NousResearch Hermes agent. Happy to report that it crushes, and as a 9b it runs on super affordable hardware. We also hit this one with some coding domain-specific training, and it scored a 53.33% on SWE bench on a slice of 200 samples! To me, I was really shocked to see this high of a score on a 9B model in swe, correct me if I'm wrong, but I think that's nipping at the heels of the Gemma 4 series, much larger models on this particular benchmark, which is really incredible to see! It also crushes the HermesAgent-20 benchmark, scoring an 85 vs the base model's 71! Make sure to run it hot, --temp around 1, that seems to be the sweet spot for running these particular fine tunes in harnesses. If you have trouble, you can work your way down, but it does a much better job departing from base models, overthinking when you run it, high temp ~1. Please spin it up in Hermes and let us know your thoughts! Looking forward to hearing your feedback as always! Also, those of you waiting for Qwopus 3.6 27B, I have put together a preliminary evaluation for you in my HF repo, go check it out; we will be releasing the full model very soon! I will put the preliminary repo in the comments!
Show more
0
67
1.4K
130
Forward to community