Pursuit evasion is the 3D, multi agent version of tag, where matching who chases whom decides the outcome, and a new GitHub release gives the field the shared benchmark it has been improvising without.
Prajwal Vijay has open-sourced a 35-scenario testbed for 3D pursuit-evasion games, the kind of reproducible benchmark a multi-agent coordination sub-field has been improvising without. The GitHub release, published alongside the paper on arXiv, ships a simulator, a scenario suite, and a small new matching heuristic. The testbed is the more important half of the package.
Pursuit-evasion in this context is a 3D, multi-agent version of tag: simulated pursuers chase simulated evaders through space, and the question is who should chase whom. When many pursuers face many evaders, an algorithm has to assign targets, and a bad assignment leaves a team with the right number of chasers but the wrong geometry. The team takes long detours, misses interceptions, and lets evaders slip through. The paper calls that failure mode "Geometric Sprawl": several maximum-cardinality matchings exist, and some spread the team across space in ways that defeat the chase.
Vijay's fix is a two-step recipe. First, pick the matching that pairs the most pursuers with evaders, the standard move. Second, among those ties, break the tie using a value the field already knows how to compute for evasion: the Hamilton-Jacobi-Isaacs interception value, an estimate of how long a given pursuer needs to catch a given evader from its current state. The result is a weighted sequential match, solved at each step by min-cost max-flow, the same routing-style optimizer that ships inside logistics and network-flow toolkits. The unweighted baseline is the same solver minus the secondary weight, so any gap between them is the value of geometric awareness, not of a different backend.
The benchmark is deterministic: 35 scenarios in seven families (a family is a geometric template, like evaders behind a wall or pursuers on a longer leash) with five initial-position variants each. In the diagnostic stationary-unmatched setting, the weighted method resolves 13 of 15 stress cases the unweighted baseline times out on, per the arXiv abstract. In the hybrid saddle-point setting, where unmatched evaders drift toward goals, mean captures rise from 3.20 to 3.91 (a 22% gain) and mean interception height climbs from 3.87 to 5.36 (40% higher), with lower path tortuosity and lower angular effort, all reported on the paper's own benchmark.
That last paragraph is the one the field will push on. There is no independent reproduction in the source bundle, no hardware-in-the-loop run, no scaling beyond the 35 scenarios, and no adversarial co-evolution test where evaders learn the pursuers' assignment strategy. The paper has been accepted to the MARS Workshop at ICRA 2026, a workshop venue rather than a full-conference track, and the right way to read a workshop-tier contribution is to take the limits as part of the package.
What the release buys the community is narrower and more durable than a percentage point. The GitHub repository ships the simulator, the weighted and unweighted sequential baselines, two simpler baselines (nearest-single and random), and the seven scenario families, so any follow-up paper can run on the same playing field and report numbers that mean the same thing. That is how a research area turns one team's experiment into shared infrastructure: a paper proposes a heuristic, another team tries a different one, and the third paper compares them on the same charts. The 22% number is the headline in the abstract; the testbed is the part that compounds.
Two things to watch. First, whether other groups pick the benchmark up before ICRA in June, the natural moment for a comparative paper to land. Second, whether the geometric-sprawl diagnosis generalizes outside the seven scenario families, which is the open question the testbed now lets a stranger answer instead of the original author.