Register and share your invite link to earn from video plays and referrals.

xjdr
@_xjdr
building AI that wont embarrass me in front of my own standards
725 Following    29.3K Followers
its not perfect but AA 52 feels like my functional canonical threshold even though i disagree with everything about the measurement and collection
i have been a classic not qwen fan but 3.8 27b is the first model that i will wholeheartedly say is legitimately good no qualifications. its benchmaxed but it probably doesn't matter anymore
mute dont block anyone who was wrong about the ox model
its promising but woefully incomplete (and kind of naive). i still fuck with it tho, and when the priors on the causality claims are fixed (even thought they punted on some of the core causality complexity and temporality complexity) its kind of nice.
Show more
Cordis is very solid work . Hoping to spend some real time with the paper and code this weekend
i love cooking . in another life, I'd love to make a twitch stream about cooking Michelin quality meals at home (even though its mostly about ingredients and sourcing) . (this is also a post about AI)
No one is releasing base models anymore. Feels noteworthy. The ones that do should be celebrated.
My desire to vague post has never been higher and it's impact has never been lower
my review of cordis is incumbered by how bad dsh is at many trivial things . this was an unexpected development
Cordis is very solid work . Hoping to spend some real time with the paper and code this weekend
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now!
Show more
Probably nothing (they should do the same for opus)
i have decided that people throwing tokens at vibe slopping math is a net positive. i've turned the corner and am now entertained . So go on all you truth seekers, download lean4 and Mathlib and sling them tokens!
Show more
no ... please dont do that
shoutout to the OG olmo team (and specifically the olmoe team) . i needed a testbed for a specific thing and the only model that fit the bill was the olmoe model(s) . excellent work and i just want you to know its still appreciated
Show more
this is fable and the 'remote team' is luna medium (prompted carefully)
at this point id start from scratch in a new repo using the old one as a reference. it sounds crazy but its probably a 5 day distraction that will solve 80% of your issues (at a cost of tokes and time obviously) . when i have gotten to this place in the past this is the only thing that allowed me to get back on track
Show more
Bend2's state is a bit sad right now, the last prompt failed The situation is: - theory is formalized in Lean and proven consistent. BUT the formalized statement has some silly, non critical errors that require update - the core implementation is very stable. I audited each line obsessively, several times. AI models can't find any exploit (inconsistency / proof of Empty). this is the most solid part and I'll place a bounty in its consistency - the compiler and runtime are working. Bend outperforms C in most benchmarks, it parallelizes with near ideal speedup up to 1000's cores, 10x-100x faster than Bend1. but the code is a mess. layers upon layers of AI slop, a bunch of r-word stuff I didn't have the time to purge yet. "it works" is the best I can say about it I'm now torn between: 1. launch all as is, clean up AI slop and fix bugs later 2. launch just the core, leave the compiler + runtime to later 3. not launch anything at all and wait 1-2 months til each line is pristine I'm super stressed because I wish the AI would just write good code so this would be done already. launch is strategically relevant because it is when many people outside of this small bubble will try it, and they won't be as kind. if things break they might just give up and not come back
Show more
sure i could do this with sol ultra but am i smart enough to do this with luna max?
if this is accurate ... oh boy
0
67
1.8K
73
Forward to community
"our apps let your agents work like a full fortune 500 company with our proprietary and customized personas" - brought to you by the team that has never worked in a F500 company "our ai coding harness and ai operating system lets you write faang quality code at scale that is secure and production ready " - brought to you by the team that has never worked in faang and has 400 active P0s and SEV1s
Show more