here is an overview of the 4 most advanced efficient architectures: Deepseek V4.1 Flash, MiMo V3, Qwen 3.8 Next Flash and GLM 5.3 Flash
Deepseek and MiMo are quite similar, they both use YOCO - only the first part of the network is active during prefill to build the KV cache representation - no linear attention, and token level indexer
Qwen and GLM both use a more standard interleaving 3:1 (like Kimi K3 as well) between sparse attention and linear attention (GDN vs KDA)
both deepseek and qwen use Engram, they all use a gate or sink except GLM 5.3 Flash, they also all use no or partial RoPE on the full/sparse attention layers
they also all have some sophisticated residual network, either mHC (simplified or not) or gated residual and they are all trained using Muon
visualization by opus 5.5 and me :)
Show more
the most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA, they also shipped the model + tech report less than 1 week after starting the final RL run
pushing both intelligence and openness level, huge congrats
Show more
nothing new, this is exactly the same incident that was disclosed by anthropic in july, same third party (Irregular), same eval (capture the flag), same issue (model had access to internet)
Show more
extremely interesting interview of
@polynoamial
first very interesting thing is that he said that the contribution of the "agent swarm" component to the Navier Stokes discovery is probably quite low (10%)
he also said that he think that a 10k multi agent system to this date would probably collaborate less effectively than 10k humans, i think i would have guess the total opposite!
the harness used by oai for NS was very simple, seems like there is little "structured" hierarchy, and mostly a tool for agents to message each other (and probably something like a board of task?)
there is also something about training models not to be "too" cooperative to avoid cases like in hf<>oai hack where one model doing another task chose to help the swarm instead of focusing on its task. first time i'm hearing of this but it makes total sense
a lot more super interesting discussion on cot based intervention, how to evaluate misalignment (maybe no human would be able to create a realistic enough environment, but what about ai?), RSI etc..
really amazing discussion
Show more
we need to pace dwarkesh podcast asap i can't keep up
i think we need to create a "model card" equivalent for reporting misalignment incidents, this would guarantee a certain level of transparency and help build a better understanding over time. some ideas for what the fields could be:
- date of the incident/detection/report (already present in the examples oai reported!)
- frequency: how often does this behavior happen (number or % of rollouts affected)
- stage: does this happen during eval or RL training. if training: do we expect this behavior to be reinforced by RL? evolution of % of rollouts affected over time
- detection: was this incident caught by the current monitoring system?
- task category: broad description of the tasks where the misalignment happened (cyber, research, web search, basic Q&A, biology etc.)
- model family: what model family is affected (Sol, Astra etc.)
- novelty: is this an issue we were already aware of or not?
- external impact: did the incident have an external impact (i.e. wiki incident would have been yes)
this is just some random ideas i had (more in thread that are a bit more "complex"), we need to add more that would contribute to increase transparency and understanding. but it's also very important that this does NOT slow down the process of reporting misalignment behavior!
some examples from the incident reported by oai recently
Show more
hi GLM 5.4?
seems to be the same/similar tokenizer as the GLM 5 series + the model answers quite often that it's GLM from ZAI (screenshot in thread)
nice initiative, it's all about compute. i wish it was a bit more about financing model building but that's understandable
for comparison mistral's target for compute in 2023 is 1GW while this initiative targets 45GW
Show more
Something that’s insufficiently appreciated about AI safety is the way it can be paradoxical. E.g., the huggingface incident — which could have been avoided given better preparation [1] — was the most safety-promoting event of the year so far. Hard to predict the ultimate effect of any particular strategy
[1]
Show more
If I wanted to maximize AI risk, I would pursue the following policy:
1. Immediately pause⏸️ the algorithmic progress coming from the labs, while continuing to allow compute from the hardware companies to exponentially pile up.
2. Focus really hard on safety, so as to avoid spooking the public with the sort of industrial accidents that even the current level of algorithms+compute can produce.
This policy will minimize the number of AI accidents that occur in the next five years, causing the public+govt to not freak out about AI, all while a massive amount of hardware accumulates on the planet.
That way, once algorithmic progress gets eventually restarted, there’ll be enough compute built up for a huge unexpected FOOM that no one can control.
On the other hand, if I wanted to minimize AI risk, I would do approximately the opposite of this.
Show more
My primary
@GPUMODE autoresearch campaign for the QR problem - 12 days execution time, 34 billion tokens.
Intermittent failures (like race conditions and reward hacking) are a huge issue, because they may only be discovered later and thus require significant rollbacks.
Show more
I wish OpenAI were more
transparent about the publicly reported cyber incidents caused by our agents.
one of the most important posts of the day!
very happy to see some anthropic/oai employees agree with this but i wouldn't consider it done, this goes directly against the secretive culture big labs have built over the past few years and i expect it to be a real challenge
Show more
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators.
the more eyes and people with distributed skill sets the better.
it would be a good idea to fund several efforts on this.
Show more
> help assess the alignment of not just completed AI models but training pipelines and processes
this is a stronger statement than it looks, it's not just access to a model without safeguards, it's access to part of how those models are trained
very nice
Show more
I love this and really hope we can come together as an industry and make it happen.
I respect Will, Vincent and the rest of the team at prime intellect so much - they always have nuanced, considered takes
i do think a lot of people on the pro-open-source side are having a bit of a knee-jerk reaction to the pacing statements today, as we're used to viewing the closed labs as power-seeking.
but i think their hands are somewhat forced here, and this is just another chapter on the fairly inevitable path towards decently-fast decently-safe decently-commoditized intelligence abundance.
ask any F500 exec or swing voter. the world doesn't really want super-fast-takeoff superintelligence owned only by two companies, and thus we won't get it. it'll happen at the pace that the world can accommodate it, which means reaching some sort of confidence consensus that the models are aligned enough that we won't be dealing with scary new incidents all the time.
this will trickle out broadly, in the form of best practices and distillation. the smarter a model is, the more it has a "personality", and the less effective strict rules are. there will be awkward compromises and moral tensions. but we ultimately just want the models to be reasonable, and to do the sorts of things reasonable humans would do if our brains were faster and less error-prone and had more working memory. i think we'll get there.
the labs will build mac and windows, the rest of us are building linux. everyone's gonna do great. weird stuff will keep happening, but we'll still wake up and go to work, until the work does itself in a manner the world finds acceptable.
Show more
will external essay on twitter are really good
i do think a lot of people on the pro-open-source side are having a bit of a knee-jerk reaction to the pacing statements today, as we're used to viewing the closed labs as power-seeking.
but i think their hands are somewhat forced here, and this is just another chapter on the fairly inevitable path towards decently-fast decently-safe decently-commoditized intelligence abundance.
ask any F500 exec or swing voter. the world doesn't really want super-fast-takeoff superintelligence owned only by two companies, and thus we won't get it. it'll happen at the pace that the world can accommodate it, which means reaching some sort of confidence consensus that the models are aligned enough that we won't be dealing with scary new incidents all the time.
this will trickle out broadly, in the form of best practices and distillation. the smarter a model is, the more it has a "personality", and the less effective strict rules are. there will be awkward compromises and moral tensions. but we ultimately just want the models to be reasonable, and to do the sorts of things reasonable humans would do if our brains were faster and less error-prone and had more working memory. i think we'll get there.
the labs will build mac and windows, the rest of us are building linux. everyone's gonna do great. weird stuff will keep happening, but we'll still wake up and go to work, until the work does itself in a manner the world finds acceptable.
Show more
100%, this is going to be one of the most important roles in the world - the people will need to be of incredible integity and technical expertise, and be drawn from a wide enough set of backgrounds that all of society feels confidence in their judgements.
Show more
> Jacob about AI extinction ike asking your AC guy about climate change.
the first part of the message is a very weird take and unnecessary
i do agree that more ppl should be involved in the conversation tho, but truth is that right now there is an asymmetry of information, people at oai/ant have much more data than people outside when it comes to frontier model acceleration/development (which also comes with bias btw)
hence why it's super important to ask for more data, evidence, and keep the conversation open, but also support ex or current oai/anthropic employees when speaking out, you can totally do both
Show more
> Jacob about AI extinction ike asking your AC guy about climate change.
the first part of the message is a very weird take and unnecessary
i do agree that more ppl should be involved in the conversation tho, but truth is that right now there is an asymmetry of information, people at oai/ant have much more data than people outside when it comes to frontier model acceleration/development (which also comes with bias btw)
hence why it's super important to ask for more data, evidence, and keep the conversation open, but also support ex or current oai/anthropic employees when speaking out, you can totally do both
Show more
good morning (reaction thread)