Preprint

Foraging Agents Learn to Cluster Without a Grouping Reward

Preprint: Agents rewarded for finding targets learned collective search at longer visual ranges, despite never seeing the targets.

Agents in a computer model learned to search collectively and form clusters even though the goal they were rewarded for contained no instruction to group. At longer visual ranges, their efficiency rose above the blind-search value, and the population moved from a roughly disordered arrangement to aggregation. The result comes from an arXiv preprint, so it describes a modeled system rather than an observation of animals or people.

The social signal was deliberately narrow

Researchers trained independent reinforcement-learning foragers, each of which adjusted its search policy through experience. In the main efficiency experiment, 50 agents searched among 100 targets. The team varied visual range, target depletion time and the duration of the tag that marked an agent after it had been rewarded. Efficiency was compared with a blind, single-agent baseline.

The setup gave the agents social information without a direct grouping reward. They never observed targets. Instead, their social input was a three-state visual cue for conspecifics, meaning other agents in the population, inside a cone whose half-width was pi divided by eight. Collecting a target earned one unit of reward.

A sharp change in search strategy

At shorter visual ranges, efficiency collapsed to the blind value. Above a threshold, it rose in a sharp jump. A longer tag time shifted that jump to smaller visual ranges, but it also lowered the high-range plateau.

That change was visible in the learned policies. One pattern was an environment-tuned, blind-like search. The other was scale-agnostic: the agent turned until it saw a rewarded conspecific, then moved ballistically toward it. The second pattern used another agent's apparent success as part of its search routine, although the training objective still contained no direct grouping reward.

The policies changed alongside the population's spacing. Researchers tracked clustering with the Clark-Evans aggregation index, a spatial score for how the agents were arranged. Below the strategy transition, the index was about one, indicating a disordered population. Just before the transition, it rose above one, indicating over-dispersion. At larger visual ranges, it fell as agents aggregated.

Resource rules mattered

The two resource rules were associated with different patterns. Under competitive depletion, a target collected by one agent became unavailable to every agent during the depletion period. Under cooperative depletion, only the collecting agent was blocked. The competitive regime had a higher aggregation threshold and larger intermediate values of the index, signaling stronger avoidance between agents.

To examine the scale of the switch, the authors added a minimal first-passage model, a calculation of how long a search process takes to encounter a target. It compared single-search and collective-search times, using visual range as the detection length and testing reactive lengths of 5 and 8. For those tested parameterizations, the calculation qualitatively reproduced the boundary where aggregation began. It was not a quantitative boundary prediction: the simulations did not strictly satisfy the dilute assumption used to derive the analytical mean-first-passage times, so the model is best read as an explanation of mechanism and scale.

A crossover, not a phase transition

The finite-size analysis asked how the transition behaved as the system grew. It used a dense-phase order parameter, m, defined as the fraction of agents with at least five neighbours within a fixed radius of 2.5. Near the onset, the distribution of m was bimodal, with dispersed and aggregated states. The transition width settled at about 0.09 rather than zero, while the onset location approached about 6.2. The authors therefore classify the phenomenon as a sharp crossover, not a thermodynamic phase transition.

That distinction sets the limits of the finding. This is a result from simulated agents with coarse social sensing, not a test in a biological population. Because the agents never saw targets, the work does not establish that animals or humans would use the same strategy, or that it would survive richer sensory input. The analytical result likewise explains a mechanism and operating scale rather than fixing the crossover boundary quantitatively.

What the preprint reports

The manuscript is an arXiv preprint, version 1, dated 28 August 2026. The authors say all code necessary to reproduce the results is available through the Python library rl opts. They also disclose using large language models to extend the simulation code, guide theoretical-model development and analysis, characterize the finite-size crossover, and assist with drafting and editing, while stating that the authors checked the simulations, analyses and conclusions.

Paper data and sources

Original title: Emergent aggregation from collective foraging
Authors: Gorka Muñoz-Gil, Andrea López-Incera, Vide Ramsten et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.