Long-horizon RL in the multi-agent context seems to converge to the sort of swarm intelligence seen in ants, not humans. Swarm intelligences are inherently harder to control and monitor.
Looking for smart, timely updates on business, markets and the stories shaping India and the world? Follow Bloomberg on WhatsApp and get them delivered straight to your phone.