Register and share your invite link to earn from video plays and referrals.

kaios
@kaiostephens
founder @nipuxx, @neutralityorg | data science @uwaterloo | ambassador @Alibaba_qwen
404 Following    25K Followers
Qwen3.8 Max is getting open-sourced tomorrow and before it does I want to state my opinion on the model after rigorous use. Although I am a Qwen ambassador this is my full real opinion. The model from my experiences acts more like Opus than it does Sol, it's less programmatic and has more a expressive personality. On lower thinking modes (low, medium) this expressive-ness is very clear, it's more opinionated, moody, and definitely more likely to make irreversible mistakes. Qwen3.8 has been my daily driver for a couple weeks now and I want to say the output in omp is some of the best I've ever got, Qwen3.8 + OMP > GPT-5.6 Sol + Codex, this is a hard truth to straddle as I send hundreds of dollars/month to OpenAI for codex, but it's ability to hold this harness by the neck and work like a dog is truly unbelievable. it orchestrates sub-agents really well, and although a (very large) token burner it never disappoints. However, it still falls far behind in real-world developer experiences and definitely behind Sol/Fable in "out of the box" thinking. This thing is the biggest work horse model I've ever tested and can be set on a task to run for days without any interruptions (they really didn't lie there), to the point where it can spend hours doing extremely deep research on a indirect task. Also, dont use xhigh thinking, waste of tokens.
Show more
while testing this I saw there is a new model in codex CLI "Daybreak Blue" It seems to be currently unavailable for the codex app yet I have access through the CLI. Seems like a similar model to Mythos relative to Fable. Fully unrestricted? anyone else have any info?
Show more
has anyone extensively used paseo? its a codex-looking UI which just uses codex, opencode, omp, claude.. etc as a back-end. Very similar to the codex app but doesn't drain resources nearly as much. Also has relay + IPhone app which allows connection like codex. The main reason I still pay $200/month for codex is because of the app and integrations, I would switch in a heartbeat to this if it treats me well. along with if anyone has any other recommendations of apps similar to this please lmk
Show more
has anyone extensively used paseo? its a codex-looking UI which just uses codex, opencode, omp, claude.. etc as a back-end. Very similar to the codex app but doesn't drain resources nearly as much. Also has relay + IPhone app which allows connection like codex. The main reason I still pay $200/month for codex is because of the app and integrations, I would switch in a heartbeat to this if it treats me well. along with if anyone has any other recommendations of apps similar to this please lmk
Show more
The goal was never to remove humans from the picture. It was to get them out of spreadsheets and onto climbing walls
0
141
10.6K
855
Forward to community
If you want to run local AI for inference, buy one of these. This is an idea I had long ago but never had the balls to bring it to fruition. Really exited with what lucebox has been cooking, really want to get my hands on one of these machines.
Show more
Today, we're really excited to announce that @luceboxai is partnering with @AMD to bring heterogenous consumer hardware to life. We worked really hard on this, and were able to have Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) to beat one NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed. R9700 takes the dense path, the hot experts, the cache and the draft model. Strix Halo holds the other experts in 128gb and computes them at the same time, not after. 51.1 tok/s on the full 284B, 3.63x one DGX Spark, and $2,899 less than having two of them. More details on the machine at and full technical breakdown in the article below ⬇️
Show more
if your co-founders aren't dreaming about the startup you aren't accelerating
benchmarks of a 50% pruned Qwen3.6-35b-a3b and expert-specific quantization technique (made by me) 7.3gb model preforming => 51gb model, exiting to see where I can bring this technique to. I have some more things lined up too. I need a DGX spark😭
Show more
instagram comment section on video about LLM’s btw
0
85
2.7K
34
Forward to community
Timelapse #14#; 7 hours (9:00am-4:00pm) What I did: - Started experimenting with quantization aware REAP - Studied online post-training techniques (DPO, PPO, GRPO) - Tested qwen3.8 through ambassador program - benched qwen3.8 @neutralityproject - Made 2 synthetic datasets for left/right output of LLM’s - Stretched every hour + posture excises and rice bucket training - Watched rekraps newest video - Watched Martin Shkreli stream - Scrolled X Song: Memory limitations in artificial intelligence - Infinity Frequencies
Show more
really exited to see what qwen3.8 has to offer, all the posts saying "its not even close to fable" than showing it merely being bad at 3d design is not an accurate representation of the model. currently waiting on fixes through the dev ambassador program, once fixed I'll benchmark it through @neutralityorg + run it personally and quote with my actual thoughts.
Show more
Kimi K3 is bigger than the Deepseek moment. this one release is going to change the world as we know it. At some point models get so good they can do full training runs e2e, 5.6 Sol already post trained Luna, getting Kimi K3 alliterated in BF16 will be a super weapon. full OSINT capabilites, research new drug discovories without being blocked, the list goes on. On top of all of that Kimi K3 will be the perfect distillation model for fine-tunes of smaller models, it will be able to generate amazing synthetic data, all that can be run on rented hardware on 8xB200
Show more
hey @bryan_johnson what do we do about >400 AQI if we have to do manual labor outside?
We've just merged its first community contribution! 🚀 6 new models benchmarked and submitted via open PR: DeepSeek, Kimi K2, Hunyuan, Nova Pro, Command A, Jamba and two new countries on the map. 24 models total, from 6 countries. And because trust is the whole point: community runs are displayed flagged as "pending re-verification" until we reproduce them ourselves. Open methodology means anyone can contribute, and nobody, including us, gets taken on faith. Thank you bricepirard-spec. Who's next?
Show more
guys hear me out: pay-per-prompt + ads to advance thinking levels
When AI embeds itself in decision-making, who oversees its influence? Right now: nobody. We vows to build and open-source the methodologies, datasets and benchmarks that change that. Nothing stays behind closed doors. Helps us fund the research:
Show more
after many years on this platform I have finally hit 25k followers. many brainrot posts, ragebait, and genuine public releases have led up to this milestone so far. Thank you everyone, this means a lot.❤️
Show more
ok so i am waking up to @elonmusk retweeting our discovery 3 times. thanks! for anyone curious about how we found that Grok is the most neutral model out there, please give a follow to our new initiative @neutralityorg and feel free to support us. we aim at doing more.
Show more
and yes, i would say that @elonmusk mission with @SpaceXAI is going well. the results definitively shows how Grok 4.5 training pushed even further to become as neutral as possible and it does feel good to know there is at least 1 model out there that is within that neutral zone.
Show more
0
262
3.1K
799
Forward to community