Back
Collaboration & competition / ACL 2025

MultiAgentBench

Evaluating the collaboration and competition of LLM agents.

MARBLE’s interactive environments evaluate task outcomes and coordination milestones across different communication structures and planning strategies.

MultiAgentBenchCONCEPT / RESEARCH SYSTEM
ONE TASK · COMPLEMENTARY WORKSHARED DELIVERABLEContributions → integration → shared outcomeINTERFACES · ROUTING · MESSAGE EXCHANGEProtocolRouterA2AACPANPAgoraRequests ↔ responses · protocol-dependent paths
Complementary contributions · a shared outcome
CONCEPTUAL RESEARCH COVER / MultiAgentBench
THE RESEARCH QUESTION

How do language-model agents collaborate and compete?

Collaboration & competitionMilestone evaluationMARBLE
FROM THE PAPER / Publication figure · ACL 2025

Evaluating collaboration and competition

MultiAgentBench overview of interactive research, Werewolf, coding, social, database, and Minecraft environments with task-performance and coordination evaluation.View original
Original MultiAgentBench overview, retained from the existing publication page. Interactive environments are evaluated through task performance and multi-agent coordination.
Read the paper

Collaboration and competition

MultiAgentBench evaluates LLM teams through MARBLE’s interactive environments, including research, coding, database work, social interaction, Minecraft, and Werewolf. It measures task outcomes alongside milestone-based indicators of collaboration and competition, making intermediate coordination visible rather than reducing a run to its final answer.

Structure matters

The benchmark compares star, chain, tree, and graph communication structures, alongside group discussion and cognitive planning. This makes both the connections between agents and their coordination strategy explicit experimental choices.

Its findings are scenario-dependent. For example, the paper reports that graph coordination performs best in the research setting; that is not a claim that one topology is always best or that adding agents necessarily improves performance.

Paper and contribution

Hongyi Du is a core contributor and co-first author, including work on Werewolf Arena and theory-of-mind evaluation. This brief follows MultiAgentBench: Evaluating the Collaboration and Competition of LLM Agents, published at ACL 2025, and its contribution appendix. MARBLE contains the code and datasets.

FURTHER THOUGHTS

Beyond the paper.

MultiAgentBenchCONCEPT / RESEARCH SYSTEM
ONE TASK · COMPLEMENTARY WORKSHARED DELIVERABLEContributions → integration → shared outcomeINTERFACES · ROUTING · MESSAGE EXCHANGEProtocolRouterA2AACPANPAgoraRequests ↔ responses · protocol-dependent paths
Complementary contributions · a shared outcome

Why agent cooperation fails.

When more agents create more coordination work—not better work. From MultiAgentBench to persistent organizational capability.

Research note · 4 min read
OPEN CONVERSATION / MultiAgentBench

Continue the conversation.

Questions and perspectives on this page are welcome.

Prefer a private conversation?

Loading comments…

Leave a public comment

This conversation belongs to MultiAgentBench. Comments appear only after review. For contact details or personal matters, use the private message form.

Public · reviewed before appearing