Skip to content
BUPT-GAMMAPublic

About

Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

MASBench

Benchmarking LLM-based Multi-Agent Collaboration
under Partial Observability

Reasoning · Scheduling · Game

Python 3.10 or newer 3 task categories 600 released instances API-based inference

Overview · Quick Start · Reasoning · Scheduling · Game

MASBench overview: agents exchange private evidence, negotiate meeting times, and coordinate resource collection under partial observability

Three tasks with progressively richer collaboration: protocol → memory → routing.

Overview

MASBench evaluates how LLM agents collaborate when each agent observes only part of the information. It separates three collaboration mechanisms—protocol, memory, and routing—and measures task performance together with communication cost.

Task Collaboration challenge Mechanisms Included data
Reasoning Combine evidence distributed across two agents Protocol 300 questions: 200 simple + 100 complex
Scheduling Find a shared 20-minute meeting using five private calendars Memory + Protocol 200 scenarios: 100 simple + 100 complex
Game Explore, collect, and deliver resources with limited local visibility Memory + Routing + Protocol 100 maps: 50 at 15×15 + 50 at 18×18
  • Private observations: agents receive their own evidence, calendar, or local map view.
  • Mechanism comparisons: compare free-form communication with structured protocols, persistent memory, and message routing.
  • Task-grounded scores: evaluate answers, calendar feasibility, and resource delivery against task references.
  • Ready-to-run data: all three datasets are included; no data generation or local GPU is required to start.

Quick Start

1. Install

We recommend Python 3.10, the version used for validation on Linux. From the repository root:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
Conda installation
conda create -n llm-mas-bench python=3.10 -y
conda activate llm-mas-bench
python -m pip install -r requirements.txt

The shared requirements.txt also includes the analysis, Excel export, and replay dependencies.

The shell batch launchers use Bash; some require tmux or at. Direct Python commands below require neither.

2. Configure an endpoint

The three tasks support OpenAI-compatible chat-completion endpoints. Set the following variables in your terminal, replacing the placeholders with your provider's credentials:

export OPENAI_API_KEY="your_api_key"
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint/v1"
export QWEN_OPENAI_API_KEY="$OPENAI_API_KEY"
export QWEN_OPENAI_BASE="$OPENAI_BASE_URL"

Reasoning reads OPENAI_*; Scheduling and Game use the provider prefix selected on the command line (QWEN_* in these examples). QWEN is a configuration label and can point to any compatible endpoint serving the requested model.

Alternatively, copy each task's .env.example to .env in that same directory. Environment variables take priority. Use a model ID available from your endpoint; the examples use qwen-plus. Credentials and generated logs are ignored by Git.

3. Try each task

Run these commands from the repository root, one at a time.

Reasoning — one question, up to six speaking turns

(cd Reasoning && python -m src.run_asymmetric_agents \
  --model-key qwen_plus --chat-style protocol --mode context \
  --num-samples 1 --max-rounds 6 --max-tokens 1024 --seed 42)

Scheduling — one scenario, up to two negotiation rounds

(cd Scheduling && python simulate_meeting_qwen.py \
  --scenario-dir meeting/minutes_10/5days/4/1 \
  --model qwen-plus --provider qwen --t 20 --max-rounds 2 \
  --memory false --protocol false --thinking false --max-tokens 1024)

Game — one map, two environment rounds

(cd Game && CUSTOM_MAP_PATH=custom_map/15x15_v1.json MECHANISM_PRESET=base \
  python eval.py --provider QWEN --model qwen-plus --max_round 2)

These small runs check API calls, agent interaction, and output generation. A short round budget can end without solving the task. For full runs and mechanism settings, use the three task guides above.

Results and Evaluation

Task Where to look Main outputs
Reasoning Reasoning/results/ and Reasoning/logs/ Answer, F1, transcript, communication counts
Scheduling Scheduling/logs/meeting/<run>/ Negotiation log, meta.json, calendar-feasibility evaluation
Game Game/.preset_logs/<run>/ Agent actions/messages, environment trajectory, scores, usage

Each task guide includes result-analysis commands. Communication cost measures inter-agent messages and differs from total billable API usage.

Metric conventions and reproducibility notes

Reasoning reports F1 on a 0–1 scale. Its current comm_tokens field counts normalized words, and f1_per_1k_tokens divides mean F1 by total communication across samples. Unanswered samples are omitted from mean F1. Keep these legacy conventions in mind before comparing different sample counts or reproducing paper tables.

Reasoning result filenames contain the model alias and chat style. Archive results before changing settings. Scheduling and Game create run directories. Model versions and provider behavior can affect outputs even with a fixed dataset seed.

Repository Structure

MASBench/
├── Reasoning/                 # Asymmetric QA, dialogue, and scoring
├── Scheduling/                # Calendars, negotiation, and data generation
├── Game/                      # Grid world, mechanisms, and replay
├── docs/assets/               # README figures exported from the manuscript
└── requirements.txt           # All Python dependencies

Task illustrations are from the MASBench manuscript. Reasoning follows its HotpotQA-based construction; Scheduling uses synthetic calendars; Game includes authored resource-coordination maps.

Data Sources and Licensing

The Reasoning dataset is adapted from HotpotQA (Yang et al., 2018) and is distributed under CC BY-SA 4.0. We rewrite the original questions and context passages and distribute evidence across two agents for collaborative reasoning. See the Reasoning attribution for details.

Original code, documentation, Scheduling data, Game maps, and manuscript figures are licensed under MIT. Copyright (c) 2026 Qizhi Chu. The HotpotQA-derived dataset in Reasoning/reasoning.json is excluded from this MIT grant and remains under CC BY-SA 4.0 as described above. Third-party materials retain their respective licenses.

About

Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages