Very impressive as usual from @Meta and @alexandr_wang!
Muse Spark 1.3 debuts as the 3rd top performing model in Stata Benchmark.
It beats EVERY model except Claude Fable
Putting LLMs in a game theory set up where they need to coordinate and reason about each other's beliefs. I show a few things: first, that LLMs can play a 'global game' with close to optimal strategy.
Second, that there is a downstream "agitating" effect to communication: when agents communicate, they are more likely to revolt against their government.
Third, that agents are more likely to revolt exactly when they get evidence that others are willing to act.
And finally, that surveillance that is perceived as adversarial reduces participation, as agents omit mentions of direct action and willingness to participate.