Three tasks with progressively richer collaboration: protocol → memory → routing.
MASBench evaluates how LLM agents collaborate when each agent observes only part of the information. It separates three collaboration mechanisms—protocol, memory, and routing—and measures task performance together with communication cost.
| Task | Collaboration challenge | Mechanisms | Included data |
|---|---|---|---|
| Reasoning | Combine evidence distributed across two agents | Protocol | 300 questions: 200 simple + 100 complex |
| Scheduling | Find a shared 20-minute meeting using five private calendars | Memory + Protocol | 200 scenarios: 100 simple + 100 complex |
| Game | Explore, collect, and deliver resources with limited local visibility | Memory + Routing + Protocol | 100 maps: 50 at 15×15 + 50 at 18×18 |
- Private observations: agents receive their own evidence, calendar, or local map view.
- Mechanism comparisons: compare free-form communication with structured protocols, persistent memory, and message routing.
- Task-grounded scores: evaluate answers, calendar feasibility, and resource delivery against task references.
- Ready-to-run data: all three datasets are included; no data generation or local GPU is required to start.
We recommend Python 3.10, the version used for validation on Linux. From the repository root:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtConda installation
conda create -n llm-mas-bench python=3.10 -y
conda activate llm-mas-bench
python -m pip install -r requirements.txtThe shared requirements.txt also includes the analysis, Excel export, and replay dependencies.
The shell batch launchers use Bash; some require tmux or at. Direct Python commands below require neither.
The three tasks support OpenAI-compatible chat-completion endpoints. Set the following variables in your terminal, replacing the placeholders with your provider's credentials:
export OPENAI_API_KEY="your_api_key"
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint/v1"
export QWEN_OPENAI_API_KEY="$OPENAI_API_KEY"
export QWEN_OPENAI_BASE="$OPENAI_BASE_URL"Reasoning reads OPENAI_*; Scheduling and Game use the provider prefix selected on the command line (QWEN_* in these examples). QWEN is a configuration label and can point to any compatible endpoint serving the requested model.
Alternatively, copy each task's .env.example to .env in that same directory. Environment variables take priority. Use a model ID available from your endpoint; the examples use qwen-plus. Credentials and generated logs are ignored by Git.
Run these commands from the repository root, one at a time.
Reasoning — one question, up to six speaking turns
(cd Reasoning && python -m src.run_asymmetric_agents \
--model-key qwen_plus --chat-style protocol --mode context \
--num-samples 1 --max-rounds 6 --max-tokens 1024 --seed 42)Scheduling — one scenario, up to two negotiation rounds
(cd Scheduling && python simulate_meeting_qwen.py \
--scenario-dir meeting/minutes_10/5days/4/1 \
--model qwen-plus --provider qwen --t 20 --max-rounds 2 \
--memory false --protocol false --thinking false --max-tokens 1024)Game — one map, two environment rounds
(cd Game && CUSTOM_MAP_PATH=custom_map/15x15_v1.json MECHANISM_PRESET=base \
python eval.py --provider QWEN --model qwen-plus --max_round 2)These small runs check API calls, agent interaction, and output generation. A short round budget can end without solving the task. For full runs and mechanism settings, use the three task guides above.
| Task | Where to look | Main outputs |
|---|---|---|
| Reasoning | Reasoning/results/ and Reasoning/logs/ |
Answer, F1, transcript, communication counts |
| Scheduling | Scheduling/logs/meeting/<run>/ |
Negotiation log, meta.json, calendar-feasibility evaluation |
| Game | Game/.preset_logs/<run>/ |
Agent actions/messages, environment trajectory, scores, usage |
Each task guide includes result-analysis commands. Communication cost measures inter-agent messages and differs from total billable API usage.
Metric conventions and reproducibility notes
Reasoning reports F1 on a 0–1 scale. Its current comm_tokens field counts normalized words, and f1_per_1k_tokens divides mean F1 by total communication across samples. Unanswered samples are omitted from mean F1. Keep these legacy conventions in mind before comparing different sample counts or reproducing paper tables.
Reasoning result filenames contain the model alias and chat style. Archive results before changing settings. Scheduling and Game create run directories. Model versions and provider behavior can affect outputs even with a fixed dataset seed.
MASBench/
├── Reasoning/ # Asymmetric QA, dialogue, and scoring
├── Scheduling/ # Calendars, negotiation, and data generation
├── Game/ # Grid world, mechanisms, and replay
├── docs/assets/ # README figures exported from the manuscript
└── requirements.txt # All Python dependencies
Task illustrations are from the MASBench manuscript. Reasoning follows its HotpotQA-based construction; Scheduling uses synthetic calendars; Game includes authored resource-coordination maps.
The Reasoning dataset is adapted from HotpotQA (Yang et al., 2018) and is distributed under CC BY-SA 4.0. We rewrite the original questions and context passages and distribute evidence across two agents for collaborative reasoning. See the Reasoning attribution for details.
Original code, documentation, Scheduling data, Game maps, and manuscript figures are licensed under MIT. Copyright (c) 2026 Qizhi Chu. The HotpotQA-derived dataset in Reasoning/reasoning.json is excluded from this MIT grant and remains under CC BY-SA 4.0 as described above. Third-party materials retain their respective licenses.
