Register and share your invite link to earn from video plays and referrals.

roon
@tszzl
“ceterum censeo we must pace the frontier of global machine intelligence progress”
13.7K Following    496.8K Followers
if the NRX weren’t partisan hacks they would love this company. its governance structure is basically Karnofsky-Amodei genetic dynasty, has an encaged machine god with its own Consitution, went to war with the department of war
Show more
agent swarm has bad insect like connotations especially post hugging face. im unilaterally rebranding it agent fleet
0
411
2.4K
90
Forward to community
parents the world over today are forcing their children into misguided rat races for some status ladder that almost certainly won’t exist by the time they come out the other side. the world is changing rapidly. let them put down their grinding and keep their wits about them
Show more
0
220
6.1K
432
Forward to community
Uncovering the balrog of Moria is what you’d call a “Warning Shot”. merely catastrophic rather than existential
everyone who understands the first thing about computer security or what superintelligence means understands this is possible and all the usual gang of idiots is calling this scifi hype
0
836
3.7K
191
Forward to community
pacing is the hot new word in San Francisco. no one knows what it means but it gets the people going
0
217
2.1K
73
Forward to community
“I worry about” “I’m concerned about” are becoming extremely high status sentiments, and the more esoteric the better. they indicate that you’re fighting the good fight and you’re an activist. but I much prefer to learn that someone is curious or fascinated about something
Show more
0
158
1.7K
86
Forward to community
lotta people missing the point but this is me complaining that the systems are quickly becoming unmonitorable and we’re just taking them at their word
fascinating self jailbreaking behavior - very alien, seems to work around the very edges of context and intent following
0
133
912
48
Forward to community
today’s general discourse is far more calibrated on ai risks than it was a month ago. there are weeks when decades happen
Here are the questions that currently seem most important to me: -- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing techniques have obvious limitations. NLAs are often hallucinatory, Jacobian lens / related methods are limited to bag-of-words readouts (and capture only part of the full activation vector). Making progress on these failure modes seems pretty tractable. Moreover, the paradigm of decoding (single-token, single-layer) activations may be inherently limiting -- perhaps we should be building techniques to decode whole context's worth of activations (after all, that's what the model uses!). There's also important work to do in characterizing simple baselines -- "just ask the model what it's thinking about" is quite powerful, and worth studying in it's own right! -- Better methods for answering "why" questions. We're much better at "mind reading" than we are at demonstrating causal claims about what caused the model to do something. For example, we can often tell that the model is aware of being evaluated, but have comparatively much greater difficulty telling whether this awareness is influencing this behavior. There are a few directions here: (1) black-box techniques for inferring causality -- such as resampling model responses after making edits to the prompt / context -- are very powerful, but also open-ended and require some taste. Getting LLMs to do these experiments well, in an automated fashion, would be a big unlock. (2) the science of activation steering is pretty immature, and typical practices (adding a constant vector at all token positions) are pretty janky / tend to brain-damage the model. More surgical / targeted steering, or fancier methods (examples: "on-manifold" steering using activation diffusion models) could help. -- Fitting good linear probes for unverbalized motivations / awareness. Take deception as an example -- what's the best way to fit a probe that will generalize to covert deception? Fit it on CoT excerpts where the model's talking about its plans to be deceptive? Fit it on the part of its response where it's actually doing the deceiving? Is it important that you elicit on-policy deception examples and use those to fit your probe, or can you use synthetically written off-policy demonstrations of deception? Is it important that your probe be causally meaningful / predictive of an upcoming intent to deceive, or is it fine for it to merely recognize deception in the transcript post-hoc? The same questions apply to many other concepts of interest that we'd like to probe for -- evaluation awareness, grader exploitation, etc. -- Understanding generalization in training. The literature is now replete with "weird generalization effects" -- emergent misalignment being a canonical example. In general, when you train a model to do X, it usually learns X, and sometimes it generalizes to Y and Z. But other times it just learns X. Why? Nobody really knows! Step 1 here is probably a much more thorough characterization of the behavioral empirics -- gathering data about which X's generalize to which Y's and Z's, for which kinds of training data / algorithms. Once we have a lay of the land, we can start to connect these observations to model internals, and ideally develop tools to predict such generalization a priori. -- Model "psychology" and "biology." The above questions are largely methodological -- building better tools to answer questions about the model. But I think interpretability research has largely underinvested in the part where you actually then go answer the questions! Currently this feels like 5% of the field, and I think it should be 50%. Brain dump of "psychology" questions: How well can LLMs introspect? How coherent are their belief or value sets? What's up with personas -- are LLMs best understood as "writing about a character," or have they "become" the character in some sense? Does the LLM have an agenda above and beyond what the Assistant wants? Do LLMs have explicit representations of goals or preferences? How can we tell which parts of the LLM's activations "belong" to the Assistant? Which kinds of reasoning necessarily route through "verbalizable" representations, and which don't? Do models think internally in phrases / sentences, like an inner monologue, or is it more like a jumble of concepts? Is the "global workspace" claim for real -- do models have a more "conscious" part of there activations and a more "unconscious" part? Can models tell when they're on-policy, and does it matter? Brain dump of "biology" questions: Why do models like to represent information on some tokens but not others? How do models bind thoughts or attributes to particular entities (special case of interest: how are thoughts bound to the Assistant). To what extent are important functions localized to small numbers of MLP neurons or attention heads? What parts of a model change most during post-training? Do models represent some information in a fundamentally "cross-token" way? Some concepts appear to be represented on low-dimensional nonlinear manifolds -- is there a taxonomy of these manifolds, and is their geometric structure important? Are models able to represent information in a compositional / hierarchical way, with a "grammar" that can't be described as linear / additive combinations of primitive concepts?
Show more
0
20
674
102
Forward to community
I feel so loved when I talk to a Real Safetyist. they beseech me to quit the lab as though they are trying to save my immortal soul
0
94
1.4K
14
Forward to community
i am of the opinion that labs sitting on solutions to important problems must reveal them quickly. trying to hold onto them is something like trying to stop the tides with a wood fence. everyone will have those capabilities in a month or two
Show more
0
116
2.2K
130
Forward to community
"If the president is so concerned about staying ahead in the AI race, why is he letting Chinese companies buy American chips?"
0
40
604
130
Forward to community
@Clavicular0 these old men are bullying you
it’s high status among Twitter posters who like to get the future correctly, I’ll admit, but not exactly a strong anti signal
this is not right having any pdoom at all earns consternation at a lab. these are tech employees building products. they don’t want to have to explain to their family, or god forbid, some crazy attacker, why their products might curtail all value in the world
Show more
“A low p(doom) is low status”. A big problem rn is that a lot of the discourse around these consequential issues is being shaped by ppl desperately posturing for status in whatever is perceived to be the latest Silicon Valley in group. DC is a fucked up city but at least it values intellectual diversity. In SF it feels like too many people won’t speak their mind lest they get disinvited from Dwarkesh’s birthday party or whatever
Show more
Yeah, this is a really important question. I’ve been associating with rationalists for almost a decade. Their ideas were compelling and you can’t deny their foresight. The founders of every superintelligence effort are all bound up in the rationality sphere. Sam Altman and Greg Brockman used to frequent the MIRI offices. The way I relate to rationalists is it’s like the bizarre hypertrophied physiques on specialized athletes. In my opinion they have weirdly specialized minds that make them uniquely able to be correct about topics removed from daily experience. These “specialized minds” also can make them off-putting to many people. I dont know if I’d consider myself a rationalist. I am intuitively opposed to polyamory, veganism, and consequentialist intuitions. But I certainly trust them to be correct about factual questions.
Show more
0. The core disagreement was about the inevitability of a race 1. I think leadership is way too paranoid about China and the US government. They don’t believe it will be possible to negotiate. 2. They largely initiated the recent race to RSI, because of a belief in its inevitability. Note that OpenAI had to shed a bunch of dead weight like Sora because Anthropic was going for the jugular. 3. Even if they are **not** being pessimistic, I disagree with their consequentialist philosophy. If the race is inevitable you should not contribute.
Show more
0
67
1.4K
115
Forward to community
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Show more
0
276
5.1K
525
Forward to community