This guide covers how to use EvalBench for evaluating Antigravity CLI (agy)
agent workflows using MCP Servers and Skills.
- Overview
- Architecture
- Prerequisites
- Quick Start
- Configuration Reference
- Authentication & Supported Models
- Tool Paradigms
- Scorers & Reporting
- Comparison Reference (vs. Gemini CLI)
- Troubleshooting
EvalBench's agy CLI integration enables automated, multi-turn evaluation of agentic AI workflows running on the Antigravity CLI binary. EvalBench drives the agent with simulated users, captures tool invocations and token usage via agy's structured event stream, and evaluates correctness with configurable scorers.
- Multi-turn evaluation driven by LLM-powered simulated users following natural language conversation plans.
- Two tool paradigms: Model Context Protocol (MCP) servers (HTTP and stdio) and Skills (packaged as plugins).
- Fake MCP server support for deterministic, offline testing without cloud resources.
- Automated scoring and reporting across trajectory matching, goal completion, CSV export, and Google BigQuery.
The evaluation pipeline coordinates test scenarios, user simulation, and the sandboxed Antigravity CLI process:
Run Config -> AgentOrchestrator -> AgentEvaluator -> AgyCliGenerator -> agy
|
v
MCP servers / skills
- AgentOrchestrator (
orchestrator: agent) loads the dataset and dispatches evaluation scenarios. - AgentEvaluator manages the multi-turn interaction loop between the generator and the SimulatedUser LLM.
- AgyCliGenerator (
generator: agy_cli) manages the sandboxed agy CLI process, stages credentials and tool configurations, executes non-interactively via--output-format stream-json, and extracts tool calls, responses, and token stats.
-
Python 3.10+ and project dependencies installed using
uv:cd evalbench uv sync -
GCP Authentication (Application Default Credentials):
gcloud auth application-default login
-
Environment Variables:
export EVAL_GCP_PROJECT_ID=your_project_id export EVAL_GCP_PROJECT_REGION=global
Note
You do not need agy installed globally on the host. EvalBench installs its
own isolated copy into <fake_home>/.local/bin/ per session and always runs
that sandboxed binary.
# For MCP Server evaluation:
export EVAL_CONFIG=datasets/agy-cli-tools/example_run_config.yaml
# For Skills evaluation:
export EVAL_CONFIG=datasets/agy-cli-tools/example_run_skills_config.yaml
# For Fake MCP (offline testing):
export EVAL_CONFIG=datasets/agy-cli-tools/example_run_fake_config.yaml./evalbench/run.shSet orchestrator: agent and dataset_format: agent-format in your top-level
YAML configuration:
| Key | Required | Description |
|---|---|---|
dataset_config |
Yes | Path to the evalset JSON file |
dataset_format |
Yes | Must be agent-format (or legacy gemini-cli-format) |
orchestrator |
Yes | Must be agent (or legacy geminicli) |
model_config |
Yes | Path to the agy CLI model config YAML |
simulated_user_model_config |
Yes | Path to the model config for the simulated user LLM |
scorers |
Yes | Dictionary of scorer configurations |
reporting |
Yes | CSV and/or BigQuery output options |
Example (example_run_config.yaml):
dataset_config: datasets/agy-cli-tools/agy-cli.evalset.json
dataset_format: agent-format
orchestrator: agent
model_config: datasets/model_configs/agy_cli_model.yaml
simulated_user_model_config: datasets/model_configs/gemini_3.1_pro_model.yaml
scorers:
trajectory_matcher: {}
goal_completion:
model_config: datasets/model_configs/gemini_3.1_pro_model.yaml
turn_count: {}
reporting:
csv:
output_directory: 'results'Specifies the generator, model label, execution timeouts, and environment:
| Key | Required | Description |
|---|---|---|
generator |
Yes | Must be agy_cli |
model |
Optional | Model label (e.g. "Gemini 3.1 Pro (Low)" or "Gemini 3.5 Flash (Medium)"). Omit to use agy's default. |
timeout |
Optional | CLI turn timeout string (e.g. "20m", passed to --print-timeout). Defaults to 5m. |
env |
Optional | Environment block passed to agy. Set GOOGLE_CLOUD_PROJECT (see below). |
setup |
Optional | Tool setup block for mcp_servers, skills, or fake_mcp_servers. |
Important
Quota project: With a credential file, agy reads its project only from
the file's quota_project_id. If the file has none, EvalBench adds
env.GOOGLE_CLOUD_PROJECT, or else the key's project_id. Without a
credential file, agy uses the metadata server's project.
Evalsets define the test scenarios. See the agentic dataset format
for field definitions. Expected tool trajectories use the canonical
<server>__<tool> format.
Example scenario:
{
"scenarios": [
{
"id": "list-instances-01",
"starting_prompt": "List all Cloud SQL instances in project cloud-db-nl2sql",
"conversation_plan": "Ensure the agent accurately calls list_instances. Verify the output is returned correctly.",
"expected_trajectory": ["cloud-sql__list_instances"],
"max_turns": 4
}
]
}agy authenticates non-interactively using Google Application Default
Credentials (ADC) (AGY_ADC_AUTH=true), which also supplies outbound
credentials for authProviderType: google_credentials MCP servers.
agy resolves ADC in the standard order:
- The file in
GOOGLE_APPLICATION_CREDENTIALS, or the gcloud ADC file (gcloud auth application-default login). - The metadata server (Cloud Build, GCE, GKE Workload Identity). agy takes the quota project from the metadata project.
If a credential file has no quota_project_id, EvalBench adds one from
env.GOOGLE_CLOUD_PROJECT or the key's project_id. Without it, the model
registry of agy is empty. If GOOGLE_APPLICATION_CREDENTIALS names a missing
file, setup fails, because agy does not fall back to the metadata server in
that case.
- Supported under ADC: Flash models (e.g.
"Gemini 3.5 Flash (Medium)","Gemini 3.6 Flash (Medium)") and"Gemini 3.1 Pro (Low)". - Not supported under ADC:
"Gemini 3.1 Pro (High)"requires interactive OAuth user login and will fail if configured for automated ADC runs.
Configured under setup.mcp_servers. EvalBench translates httpUrl to agy's
native serverUrl and writes <fake_home>/.gemini/config/mcp_config.json:
setup:
mcp_servers:
cloud-sql:
httpUrl: "https://sqladmin.googleapis.com/mcp"
authProviderType: google_credentialsEvalBench pre-verifies MCP attachment during generator initialization by probing
the CLI and confirming that tool schemas were materialized under
<fake_home>/.gemini/antigravity-cli/mcp/<server>/. A server that attaches zero
tools halts the run, because agy otherwise degrades silently to shell commands
and the eval scores that as poor model behaviour.
Skills are delivered as plugin bundles under setup.skills. EvalBench invokes
agy plugin install <target> to install local plugin directories or remote git
repositories:
setup:
skills:
# Local directory
- "/path/to/local-plugin"
# Git repository with branch/tag pinning
- action: install_from_repo
url: "https://github.com/gemini-cli-extensions/cloud-sql-postgresql.git#v1.2.3"For offline testing without live network or cloud endpoints, define a local
stdio MCP server in setup.fake_mcp_servers with mock tools in
fake_mcp_tools. See
datasets/model_configs/agy_cli_fake_model.yaml.
- Scorers: See the scorer reference for full configuration options on trajectory matching, goal completion, and turn metrics.
- Reporting: Exports to local CSV files (
reporting.csv.output_directory) and Google BigQuery (reporting.bigquery.gcp_project_id).
For teams familiar with the Gemini CLI harness, this table summarizes key operational differences:
| Area | Gemini CLI | Antigravity (agy) CLI |
|---|---|---|
| Installation | npm install -g @google/gemini-cli@<ver> |
Auto-staged into <fake_home>/.local/bin/ |
| Invocation | npm exec @google/gemini-cli -- ... |
agy -p <prompt> --dangerously-skip-permissions |
| Output Format | --output-format stream-json |
--output-format stream-json |
| Session Resume | --resume <id> |
--conversation <id> |
| Settings Path | ~/.gemini/settings.json |
~/.gemini/antigravity-cli/settings.json |
| MCP Config | mcpServers in settings.json |
mcpServers in ~/.gemini/config/mcp_config.json |
| MCP Tool Naming | mcp_<server>_<tool> |
call_mcp_tool wrapper (canonicalized to <server>__<tool>) |
| Skill Registration | gemini skills <cmd> |
agy plugin install <target> |
| Model Parameter | GEMINI_API_MODEL env var |
--model flag (UI label or slug) |
| Authentication | Token + ADC | Non-interactive ADC (AGY_ADC_AUTH=true) |
- agy shows every ADC token failure as
authentication required. Run 'agy' to log in.The harness runs a startup probe and fails setup withagy failed to authenticate with ADC. The error includes theadcAuth:lines from the agy log, which state the real cause. - Outside GCP, ensure fresh ADC credentials exist by running
gcloud auth application-default login. - On GCP without a key file, ensure the metadata-server service account can
call Vertex AI (for example,
roles/aiplatform.user).
- Ensure the configured
modelinmodel_config.yamlis supported under ADC (useGemini 3.1 Pro (Low)or Flash models, notGemini 3.1 Pro (High)). - If every model is rejected, including agy's own default, the credential is
missing
quota_project_idand the model registry came back empty. Ensure a project is resolvable fromenv.GOOGLE_CLOUD_PROJECTor from the key itself.
- Confirm
httpUrlis valid and points to an active MCP endpoint. - For Google APIs, ensure
authProviderType: google_credentialsis set and ADC is active.
- Inspect setup logs for
agy plugin install '<target>' failedfor the exit code and stderr. Ensure the target directory contains a valid manifest (plugin.jsonorgemini-extension.json). - Confirm the plugin registered in
<fake_home>/.gemini/config/import_manifest.json.
- A missing quota project makes agy reject every model. See Model Authorization.