Register and share your invite link to earn from video plays and referrals.

xjdr
@_xjdr
building AI that wont embarrass me in front of my own standards
711 Following    29K Followers
"our apps let your agents work like a full fortune 500 company with our proprietary and customized personas" - brought to you by the team that has never worked in a F500 company "our ai coding harness and ai operating system lets you write faang quality code at scale that is secure and production ready " - brought to you by the team that has never worked in faang and has 400 active P0s and SEV1s
Show more
ive spent a lot of time testing new models over the last few weeks and im genuinely curious what anthropic is going to do . there is absolutely no reason at this point to use sonnet or opus (with the gpt 5.6 family as the latest releases there are now _many_ much better and much cheaper alternatives) and fable, while seemingly a very good model and at the frontier in certain places, is hostile to most professional use and is far too expensive to be either practical or desirable (outside of potential exceptional cases) . i have genuinely grown to like the very opus3 flavored fable (for non professional work) so i hope they use this as motivation to speed up and not just make a mediocre opus 5 like they made a mediocre sonnet 5. and they can say what they wish about glm5.2 but it is very clearly not _just_ an opus distillation, it has distinct qualities and character and grok4.5 is close enough to opus for most code related use cases and a such a lower price that there has to be some pressure felt at some point (at least one would think) .
Show more
spent yesterday on both grok 4.5 evals and using it in practice. it is a very good model for its price and intended use case. its smarter than glm5.2 (the model i would most immediately compare it to) in a lot of ways for ncode harness use and shows the beginning of real frontier RL / post training polish . it is very good at operating in ncode and makes excellent use of the tools at its disposal all while being extremely token efficient (i'd say its most distinguishing quality) . my 2 major complaints are no 1m ctx (which is basically standard now and i find myself missing often in real use) and its price point is just a _touch_ too high for where it fits (in my world at least) . ideally, i would use it as the subagent execution arm to replace glm5.2 or gpt5.5 med or to replace GLM 5.2 as the planner and delegate to dsv4-flash or gemma 4 31b . that said, i do plan to use it (unless gpt5.6 replaces it today) quite a bit from now on. if they can apply (or even improve) this post training polish to their next 2T base model (that is currently training as i understand it) , then they could have a very interesting next release
Show more
at this point, i wonder why zuck and elon wouldn't just OSS their models? if progress isn't where you want it to be and you are selling hardware because you can't find enough use for it, open sourcing seems like the the logical and pragmatic next step.
Show more