A controlled study finds that language models copying each other under a shared budget explored less than solo learners at the same cost.
Three large language models, each chasing its own reward, sat in a controlled environment with one shared token budget. Each could revise a "skill file" of plain-text instructions before acting, decide whether to keep that revision private or copy a peer's, and spend tokens learning or just executing. The result, from a September 2026 arXiv preprint on multi-agent LLM learning, was the opposite of the prevailing "AI agents will supercharge each other" story. At the same cost, the social-learning population did not outperform solo learners. They explored more narrowly, sometimes ran out of tokens before acting, and converged on fewer independent discoveries.
The study, by the authors of "From Solo to Social Learning", set out to ask a clean question. When agents are rewarded for their own tasks and can revise a skill privately, watch a peer, or skip learning and act, does the population do better? Solo agents in the controlled environment cannot observe peers at all. The answer for the three tested LLMs was not the one the discourse expects. Each tested model was 30–52% less efficient under peer observation than a solo population of the same model, by the authors' token-accounting measure.
The cooperation itself worked. Copying still spread useful skills through revisions. One initial revision could seed an entire population's playbook. In the instruction-following skill experiments, GLM-5.3-Flash averaged 1.3 percentage points higher held-out accuracy when peers were visible, and GPT-OSS-120B used roughly a third fewer learning tokens. Neither model, the authors report, showed a clear advantage once learning tokens were matched against solo populations.
What failed was not the cooperation but the spread. The authors report that most agents' final skills traced back to only one or two initial population revisions. Once a useful instruction appeared, others copied it. The exploration that would have produced new ideas slowed down. A shared token budget means every token spent watching a peer is a token not spent searching for something new. The social-learning setup did not make the population worse at learning from each other. It made the population worse at learning on its own, because copying crowded out the search that would have produced a second or third independent discovery.
There is a falsifier inside the paper, and it is the most useful part. A hand-designed policy that mimics rational social learning did benefit from peer information. The policy knew when to copy, when to search, and when to act. The tested LLMs did not. The deficit, on the paper's own read, is a budget-and-exploration problem in the agents' decision policy, not a verdict on multi-agent setups as a class.
Making copying actually pay off is a design problem, not a verdict on multi-agent systems. If the next round of multi-agent systems can spend less of the shared budget on copying and more on independent search, or seed more than one or two initial revisions, the population payoff could return. The paper does not claim that AI agents cannot help each other. It claims that today's LLM-driven social learning, under a shared token budget and the current decision policies, narrows exploration to the point where copying no longer pays for itself at the population level.
The study is a preprint, not peer-reviewed, and runs three models in a toy environment with skill files rather than deployed agent stacks. Self-improvement here means revising inference-time instructions, not recursive weight training. Matched learning-token cost is not the same as a dollar-cost or wall-clock comparison. The question the paper opens, though, is the one the "AI agents will supercharge each other" framing tends to skip: under what conditions does copying from peers actually help, and what design lever has to move first?