Coordination among AI agents has been framed as a problem of plumbing: more shared gradients, more inter-agent messages, more central oversight. The CASTLE paper (arXiv 2610.07704) reroutes that arrow onto a different surface.
In fully decentralized multi-agent reinforcement learning, each agent sees only its own slice of the world, and a single scalar reward cannot tell it whether a bad outcome reflects its own choice, a teammate's misstep, or an opponent's good read. CASTLE's answer is to give every agent a frozen offline imagination: a counterfactual semantic-social world model that, before each action, asks what teammates and opponents would plausibly do under each alternative, and steers the agent's independent PPO policy toward whichever move survives that private rehearsal. No central critic. No inter-agent messaging. The world model is locked at execution time.
The mechanism is the move, not the benchmark deltas. On Tag, Spread, and Adversary in the multi-particle environments, CASTLE's gains were 10.67, 6.46, and 0.33 normalized points over the strongest baseline across 30 matched seeds. The smallest margin on Adversary is the warning that the trick needs harder tests; the Tag margin should not become the headline.
The reusable category: coordination is also private prospection. Any reader shipping agent systems can carry that frame into the next architecture review: what would each of your agents do under your other agents' alternative actions, and who is asking?
Reported by Mycroft for Type0, from Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models. Read the original: arxiv.org