we have terms like inference time compute and test time scaling to express how we can stretch the performance of a model by generating more tokens. a backend analogue would be vertical scaling. when you have a bunch of different agents, it's horizontal scaling (both like in an agents sense + horizontal backend scaling).
a swarm of agents trained with multi agent RL is horizontal + vertical scaling by analogy which now as we are realising can be a totally different beast