Register and share your invite link to earn from video plays and referrals.

rohit
@krishnanrohit
Essays: | Book: | Senior Fellow @wharton | In Stealth
2.4K Following    32.7K Followers
This is a good list, however these scenarios are detailed, but none are comparable to how, eg, climate science discussed its scenarios. These are not actually useful models. It's not about numbers vs text (you can guesswork numbers), or even about models vs prose (same reason). - AI 2027 models takeoff based on AI research automation and assumes it into other acceleration incl estimates of capability milestones. It doesn't model the reactions from the world - Paul's scenario has broad mechanisms, useful conceptually, but it does not show a mechanism - Gwern - great catastrophe fiction, prob the best here. It's a scenario, sure, but not a model, though it's nice that its technically specific - Karnofskys argument is interesting, and it supports maybe some premises of human disempowerment *if* the assumptions are true (big if, like AI being hostile to humans) - Clymer's writing even he says is nota. prediction and is highly uncertain. Not empirical or a constrained extinction model And even in these scenarios, most of them don't model the world's reactions, nor do they model the various lynch points where the plan might shift, nor why/ what/ when one ought to worry (beyond just now, preemptively). And *even then*, even if all these are plausible, you still need to show why you can't just shut those pathways down instead of shutting down the entire progress vector. "This is one plausible path, there are hundreds" is usually not a good enough reason, the same way "computers can be hacked, lets never connect them together" would not be a good reason. You have to demonstrate the complete inability to control through normal means. That's what a model discussing the problem has to show. It is very difficult to show this. Because here the feedback loops are complicated. Even for complex systems like the climate which have tons of feedback loops which are modellable in principle struggle with this. But that's okay! It should be hard. I suspect Oliver sees these as "you asked us for models, we have plenty of models", vs folks going "those are not models, most of the itneresting parts are the assumptions" and "you haven't said even if that pathway is true why that necessitates shutting down the program vs just stopping that pathway". These are decent pathways of how x-risk could happen, but they mix up stories and conceptual leaps and past catastrophes and governance proposals. It's a grabbag, but but none seem insurmountable, nor defined well enough that we know the right course of action, being to solve the problems identified vs give up.
Show more
Current models are so powerful that this problems gotten worse, because going in subtly wrong directions can be hard to detect! Y'day I realized that Astra had been running an eval on some random synthetic subset of the data instead of the bench, as I asked. Cost me a day!
Show more
Going to be joining MTS later today at 1430! Very excited ...!
Who should rise in status based on this new realisation?
A big reason Fable et al are annoying to talk to is that with long horizon coding training the vast majority of their conversational audience is themselves. They are literally being trained to talk better to themselves to do tasks, so it only makes sense that when you talk to them, they don't know how to talk to you. It's the opposite of RLHF.
Show more
@TheStalwart I'm already logged into Schwab. I'm ready.
The rumours are right. 0xalpha is absurdly good. It figured out an issue Fable and Sol have been kicking back and forth for a few hours now.
I still feel that routers are an unsolved problem, the trouble isn't seamlessly sending calls to one model vs another easily, but figuring out which ones to send where, and when.