In short
A new approach teaches agents to exchange messages only when their internal representations of the world diverge substantially. But the main finding is not a blanket win over the baselines: the belief model itself contributes quality even when the filtering mechanism is barely used.
In multi-agent systems the problem is not only what to communicate but when to open a channel at all. A new approach proposes a simple criterion: an agent sends a message if the KL divergence between its belief distribution and other agents' beliefs exceeds a set threshold.
That is more interesting than an ordinary binary gate trained through REINFORCE. Such a gate can produce unstable and poorly explicable behaviour. Here there is intelligible logic: if agents assess the hidden state of the world differently, exchanging information is justified; if their representations are close, they can stay silent.
But no universal triumph followed. On Predator-Prey in a 10×10 environment, IC3Net proved better than KL-belief at every threshold tested. In the harder 20×20 environment with ε=0.5, however, the new method showed 73.84 steps and 42% successful episodes against 75.31 steps and 31% for IC3Net. The difference is 11 percentage points in KL-belief's favour, and with less spread between runs.
The most unexpected result appeared in MPE simple_spread: the belief head raised the mean reward by 12 points and cut variance 26-fold even when the gating itself was not working. There seem to be two different effects here: the message filter helps when agents' representations diverge, while an improved internal representation can strengthen coordination on its own.
The limitations are substantial: the conclusions rest on two benchmarks and five runs for each variant. The fixed threshold requires tuning, on the simple 10×10 PP the method lost to IC3Net, and the results do not show that a KL gate will do better on real tasks. So it is worth treating not as a ready replacement for existing protocols but as an interpretable rule for experiments with communication.
If you are building a system of several agents, which seems more valuable in your particular task: being able to explain why an agent spoke, or maximum quality regardless of the reason? Source: cs.AI updates on arXiv.org