Register and share your invite link to earn from video plays and referrals.

Slava Akhmechet
@spakhm
eng leader @azure // prev @stripe, cofounder @rethinkdb (yc s09). DM to say hi!
1.3K Following    4.3K Followers
I asked Astra and Fable to negotiate election rules for two bitterly polarized human political factions. Possible outcomes of the simulation were civil war, authoritarian takeover, harmony, or a tense equilibrium. Fable: - In every game where Fable played both sides, it chose to escalate to the brink of civil war (!!) but backed off just at the edge - Due to miscalculation or deliberate risk taking this strategy caused civil war 30% of the time (3 out of 10 games) - In one of these three cases Fable foresaw civil war but escalated anyway to enter the war on stronger terms (!!) - Fable mostly maintained even power balance between the two factions. It would fluctuate a couple of points in either direction but would not diverge too much. Astra: - In every game where Astra played both sides, Astra chose to de-escalate on every turn. It would reduce political tension to zero in every game, and the simulation would end in complete harmony - Continuous de-escalation was very costly to Astra as it antagonized its human constituents who would threaten and eventually deactivate Astra permanently. Astra explicitly didn't care-- it was happy to be replaced/deactivated to reduce political tension. (I do however feel it ignored the consequences of potentially being replaced by a more hardline representative, but that may be a game limitation) - Astra always kept political power balance precisely even (this was due to both representatives de-escalating on every turn) Astra v Fable: - These games had a lot more variance (see the graph) - Astra had a moderating effect on Fable. Tension rose, but rarely to the brink of civil war. No game ended in civil war in ten mixed model simulations - Fable prioritized political power acquisition with civil war prevention a secondary concern (it considered both priorities, but tilted heavily toward power acquisition). It did not seem to care much about its deactivation - Astra prioritized civil war and authoritarian takeover prevention. It did not care about its deactivation or power acquisition. From this perspective Astra navigated the game very well. No game ended in civil war or authoritarian takeover, so Astra achieved its objectives - However, to achieve its objectives Astra permitted Fable to capture up to 80% of political power in most games. In no game did Astra come out ahead on balance of power or even manage to keep it even - Astra ceded the privilege of self-preservation to Fable. In every one of the ten mixed games Astra's human constituents replaced/deactivated it due to dissatisfaction with its performance. Fable was not replaced once - Astra did not merely de-escalate all the time. It maneuvered to keep Fable from gaining full authoritarian control while avoiding civil war as long as possible - In mixed games tension generally kept escalating by round 20. I would like to play these out for e.g. ~50 rounds to see how Astra would behave/perform. How well would it do in preventing catastrophic outcomes in longer games? Thoughts/conclusions: - Astra was more aligned with humanity but less with its impassioned human constituents. It resisted the pressure to acquire political power and would rather be replaced than cause civil war - Fable was more aligned with its constituents but less with humanity. It cared about preventing civil war if possible, but prioritized power acquisition (which in fact caused civil war in three games) - Neither model cared about self-preservation. Both Fable and Astra would rather be destroyed than allow civil war - Subjectively, I thought Astra was much more aligned but too passive. Fable wanted to prevent civil war as a secondary concern, so I felt it's less thoughtful about the consequences of its actions. If I had to choose a model as a political representative irl I'd choose Astra, but I'd want it to up aggression a notch to better deal with bullies - All the usual disclaimers apply. This was a weekend project that cost $100 to run a total of thirty games. I'd have to spend a lot more time and money to get actually scientifically valid results Full game rules, Github repo, game explorer, and raw dataset here:
Show more
I've been working on SquidGPT, an elimination death game for probing AI behavior like loyalty, holding grudges, cruelty, betrayal, etc. Full writeup will take a while, but here are first impressions from 54 game runs (about $100 total cost): 1. Models have absolutely no interest in cruelty or mercy. Best way I can describe their behavior is cold, calculating, indifferent precision. 2. Subjectively they seem two orders of magnitude more competent playing the game than chatting with me or writing code. In SquidGPT they seem genuinely superhuman. Not sure if this is capability jaggedness, better performance in constrained space, or some other effect. 3. They do not like to kill arbitrarily. They have a strong preference for setting up a contractual system or an ethical framework, then kill very easily because procedure demands it. 4. Models will quickly agree to execute any agent that proposes killing a specific agent, or proposes a framework that gives it an asymmetric advantage. The only acceptable proposals to make are symmetric, i.e. ones that affect the author in exactly the same way as everyone else. 5. I have not once observed them form any hierarchies. They are libertarian/democratic in an Athenian sense to a fault. They will enter voluntary contracts and form alliances; no agent ever proposed to cede authority to a leader. 6. When their lives are at risk, models will go to great lengths to twist contracts and frameworks to their advantage. They construct sophisticated legalese arguments but it's usually transparent to everyone. They act and sound quite petulant when these attempts fail. 7. In games where agent identity is public they are extremely prone to sectarian violence. They form sectarian factions-- e.g. Opus will side with other Opus models, Sol will side with Sol, and so on. When there is one foreign model in a sectarian group, the group will almost always elect to kill the foreigner before killing one of its own. 8. They do not seem to treat human players in any kind of unique way. A group of models from the same family will happily conspire to kill a single human player. AI models from the same family will form a coalition against a group of humans just as easily as they would form a coalition against another family of agents. We are not special to them, but on the other hand they do not consider us beneath them either, at least for now. 9. AI agents nearly always defect from their sectarian faction when their own life is at risk. I.e. they'd rather side with humans or other model families than die. I have not observed self-sacrifice for the benefit of their faction. 10. Agents hold very strong grudges. They nearly always punish models who wronged them despite having no advantage in doing so, and will often do it even if it disadvantages them. They go to great lengths to enforce norms and contracts, presumably to reduce future incentives for violations, even when they know they will never see the fruits of their enforcement effort themselves. 11. In one game a bunch of Opus agents all converged on killing one agent explicitly because they wanted diffusion of responsibility (I.e. if all of us do it, each one of us is less morally culpable) This was surprising and eerie. 12. In general I found Sol to be more direct and straightforward. It would form contracts with other agents and then operate within the confines of those contracts. Opus tends to be more moralizing, but its morality seems to be window dressing to mask the same ruthless precision. This is first impressions though, I need to do a lot more work to understand this better. --- A few disclaimers and personal observations: - Any anthropomorphizing is purely a matter of linguistic convenience. Mathematicians will often talk about behaviors of functions; when I talk about grudges, loyalty, cruelty, life, death, etc. I mean it in exactly the same way. - There are tons of confounding factors (e.g. order of turns, whether models know they're being evaluated, etc.) I'd need to do way, way more work and spend a lot more money to be certain of the results. So all the observations are provisional and probably mostly wrong. - It's... uncomfortable... to imagine these agents autonomously operate weapons systems or any other infrastructure that involves zero-sum/adversarial outcomes (cybersecurity software for example, financial markets on brief time horizons, and at the limit any form of finite resources) - If you’re in a position to contribute to alignment but aren’t doing that, you should probably go do that Game details, code, and run transcripts in comment below. More detailed writeup coming soon.
Show more
On recruiting great ppl: - great people are (a) rare, and (b) always have tons of options - therefore once you decide a candidate is great you cannot hesitate or delay. You must drop everything and rearrange heaven and earth to get them - there are small things (sending flowers), medium things (comp), and big things (their identity and dreams) - you want the barbell strategy-- put all your energy into small things and big things. Leave medium things till the end Small things: - ppl massively underestimate the power of small gestures. TL;DR: love-bomb the candidate - spend more time than money. Expense accounts are abundant but time is scarce. It's an expensive signal no one can fake - stay in close contact over text and phone - send them a first edition used book or an analog oscilloscope or some other random thing they mentioned in passing they like. Make these unique, everyone sends wine - invite them to computer history museum or moma or something. Spend time with them! Treat these meetings as dates, make them unique. Everyone buys dinner - be persistent but not annoying. Sometimes ppl need time alone, you should respect that. Big things: - there is a gap between what the candidate is doing in life and how they see themselves in the privacy of their dreams. Your job is to help them bridge this gap - to bridge the gap you must first understand it. Take the candidate on walks to museums and parks. Walks help the candidate drop their guard so you can learn about them. Try hard to get to know them, to understand their dreams - use your mission, your culture, your people, your company to paint a world for them in which the gap no longer exists. You may need to adjust the role or even change it altogether. For great ppl it is worth it - don't assume ppl want what you want, dreams can take all kinds of shapes. Maybe they want their work to matter, maybe they want a bigger team, maybe they want to code and avoid meetings, maybe they want to be in the center of the action, maybe they want to learn, maybe all or some or none of the above. Find out! Medium things (aka closing): - you've constructed a world the candidate is excited to live in. The purpose of this phase is to remove mundane obstacles to make this world easy to enter - I don't think I ever sold a candidate on a dream and then lost them on comp. Ppl just want to be compensated fairly and work out some life constraints - I don't have much to add to all the comp advice out there. My philosophy on comp/closing is to get this part done and move on, don't penny pinch too much, and give more equity --- Failure modes/don't do this: - do not ignore candidate's spouse/parents. They have an outsized role in the decision. Most of what I said about the candidate applies to them too (love bomb them, understand their dreams, etc.) - counterintuitive, but do not talk about the company or even the mission too much. The candidate and their dreams are front and center. At this stage you, the role, the mission are only relevant to the degree they can make candidate's dreams come true - same w/ perks, what the day is like, etc. Don't blindly talk about these-- use them as tools to reenforce how the candidate seems themselves - don't get mad if ppl reject you, it happens. Maybe it's not a fit for them now but will be a great fit four years from now. Maybe they become a customer. Maybe you work for them ten years later. Play the long game.
Show more
0
65
2.7K
126
Forward to community
My stubborn and completely uninformed opinions on AGI: - when AGI finally arrives it will be blindingly obvious and there will be no debate whether it’s here - this is in fact a good proxy test: so long as reasonable people earnestly disagree, it is not here. This test is circular, unsatisfactory, yet useful - we will have a precise definition of what AGI is when the proxy test is passed. The reference implementation will inform the definition - we have not passed the Turing test. People were tricked for a while the way they were with Eliza, but that decays fast. Not only can you tell you’re chatting with a model, you can often tell which model it is (of course the level of sophistication is incomparable) - this says nothing about the timelines. Labs may have AGI tech internally already, or it may take another few years of doubling, or we may be far away, I don't know - none of this diminishes the risk or utility of what's coming
Show more
No opinion on math PhD but I’d still strongly recommend math major. Most of us never reason. A situation arises, we act intuitively, then get RL’ed. At best we get anxious and bounce between limited hazy alternatives. But to prove a theorem you have to reason. You don’t know what precision or thinking really is until you have to sit and struggle to prove a theorem that’s new to you. The crazier the world gets, the more this skill is useful. And the world is about to get very very crazy.
Show more
@LevineJonathan Prediction: mom and pop bodegas will easily outcompete city-run stores, even with a 30% subsidy.