Status:current — Provenance platform overview (English); the entry point of this repository
AI value formation is observable, governable and auditable.
Built on mavisframework v1.3.4 (a self-developed generative multi-agent simulation engine, versioned independently). The application scenario is investment advisory (secondary market): agents make context-based judgments, move and converse within a spatial environment, with every step configurable, explainable and visualizable in real time.
English | 简体中文
Abstract
This is a multi-agent simulation platform. The application scenario is investment advisory (secondary market): agents make context-based judgments, move and converse within a spatial environment, with every step configurable, explainable and visualizable in real time. It serves the Global Trust Challenge process-alignment story — AI value formation can be observed, governed and audited.
Platform and engine are separated: mavisframework is maintained and released independently (v1.3.4); this repo depends on it via
mavisframework>=1.2.0,<2.0.0inrequirements.txt. Constraints never enter the prompt — an expert edit only weights the consequence feedback, so the tendency converges only through later experience (lagged convergence = observable evidence of internalization).→ 0. Current state | 7. IVD Governance Platform | Expert tour
Scope
| Covered | Not covered |
|---|---|
Real-time simulation & visualization; 5010 as the single integration entry (live face + data face) IVD governance: constraint edits, tendency curves, intervention audit, three-layer explanation Pluggable expert interventions (InterventionStrategy registry) Read-only discovery surfaces: expert tour, platform contract, field reading guide Embed surfaces /embed/* (iframe, no CORS) |
Turnkey support for any other business scenario (new scenario = a cases/*/scenario.yaml + an engine) Hardened security for public exposure (no authentication; loopback-only by default) A real market model (consequence feedback is a lightweight embedding-similarity stand-in) Final judgment on research conclusions (the AI only guarantees mechanical correctness) |
Handing over / integrating? Start with provenance/docs/文档索引.md — it shows at a glance which documents are current and which are superseded. The single contract handed to a governance platform is 平台对接契约_5010唯一入口.md.
- 0. Current state
- 1. Architecture
- 2. Environment & Engine Setup
- 3. Configure the LLM
- 4. Run the Live Simulation
- 5. Role Configuration
- 6. Run Options
- 7. IVD Governance Platform
- 8. Deploy & Embed
- 9. Notes
- 10. Custom Maps
- 11. References
- 12. Security & exposure
(checked 2026-10-01)
Layers: scenario declaration (data) -> engine (case_engine/, registrable & replaceable)
-> case (case01/, case00/) -> kernel (mavisframework, independently versioned, v1.3.4)
-> presentation (packages/mavis-vizkit, a mavis plugin).
Interventions are pluggable: the three expert write endpoints run through an
InterventionStrategy registry (live/interventions.py; built-in goals / undo
/ mark / corrective_feedback — expert corrections injected into the agent's
memory stream). New intervention = one subclass + one register() call, exposed
automatically via POST /api/intervention/{strategy_id} and listed at
GET /api/interventions. Legacy paths (/api/goals etc.) remain as thin shells.
Branch judging is LLM-first: on 25 recorded T0 answers the keyword rules
table matched the recorded branch 1/25 (long answers always trip a conditional
keyword), while an LLM judge (GLM-4.7-flash, thinking off) matched 15/16
judge-mode records (93.8%). Rules stay as the offline fallback; the evaluation
harness is case01/tools/branch_judge_eval.py (report:
results/analysis/branch_judge_eval/).
Services (localhost):
5010the single integration entry (live face + data face: aggregate/api/runs, expert-safe/api/run-detail/..., six/embed/*surfaces)5002case01 read-only contract ·5003case00 archive (read-only) ·5020staged pixel scene ·8060config tool (5004multi-scenario panel retired — review panel debug mode only)
Tests: python -m pytest tests (outer) · python -m pytest case01/tests case_engine/tests
· the packages/mavis-vizkit suite · the ../mavis kernel suite.
Docs: start at provenance/docs/文档索引.md (status index); the platform-facing contract is
provenance/docs/平台对接契约_5010唯一入口.md.
Provenance (platform, this repo)
├── provenance/ # platform core
│ ├── live_fastapi.py # real-time simulation + visualization (FastAPI + WebSocket, single entry)
│ ├── case_engine/ # engine layer: scenario declaration -> registrable/replaceable runtime
│ ├── cases/ # scenario declarations (data): case00_village / case01_stock / case02_minimal
│ ├── case01/ # case: investment advisory (secondary market) baseline
│ ├── case00/ # case: village (frozen; kept for comparison & display)
│ ├── live/ # live face: intervention strategies, network guard, ...
│ ├── frontend/ # Phaser frontend + texture pool (agents_pool/)
│ ├── scenarios/ # legacy business scenario configs (investment: roles/relations/story)
│ ├── data/ # configs & prompts
│ └── results/ # checkpoints & decision traces (decisions.json)
├── packages/ # local packages: mavis-vizkit (presentation) / mavis-case01-injector
└── depends on mavisframework # engine (separate repo hellobs/mavis, installed as wheel)
The platform and the engine are separated: mavisframework lives in
hellobs/mavis; this platform depends on it
via mavisframework>=1.2.0,<2.0.0 in requirements.txt. The role configuration tool
(config_tool) also belongs to the engine repo.
The platform depends on mavisframework>=1.2.0,<2.0.0 (not on PyPI; built from
source). Do not go below 1.2.0: 1.0.0 predates the three case01 injection hooks
(external_state / interaction_request / role_directive) and constructing
Simulator with them raises TypeError, while 1.1.0 predates the generic plugin
surface (mavisframework.plugin, Simulator(plugins=),
agent_core.subscribe_chat_line). CI asserts that surface is present, via
tools/check_engine_baseline.py, before running any test.
Execute in order:
# 2.1 Clone the engine repo and build the wheel
# HTTPS (recommended for read-only, no SSH key needed):
git clone https://github.com/hellobs/mavis.git ../mavis
# or SSH (requires a configured SSH key added to your GitHub account):
# git clone git@github.com:hellobs/mavis.git ../mavis
cd ../mavis
uv build # produces dist/mavisframework-1.3.4-py3-none-any.whl
# pip users (no uv): pip install build && python -m build --wheel
cd ../provenance
# 2.2 Create the environment (uv or conda; Python 3.12)
uv venv .venv --python 3.12
# conda users: conda create -n provenance python=3.12 && conda activate provenance
# 2.3 Install dependencies **in this exact order** (all three steps matter)
# a) the framework (wheel built in 2.1; `pip install -e ../mavis` also works)
uv pip install ../mavis/dist/mavisframework-1.3.4-py3-none-any.whl
# b) the two **local packages** in this repo (not on PyPI; skipping this makes
# the next step fail with "No matching distribution found")
uv pip install -e packages/mavis-vizkit -e packages/mavis-case01-injector
# c) the rest of the runtime deps + the test runner (requirements.txt has NO pytest)
uv pip install -r requirements.txt pytestRequires uv or conda;
pipworks the same.Self-check (
pytest.inisetspythonpath=provenance, so running from the repo root or fromprovenance/both work):pytest tests # outer: API / pages / guards pytest provenance/case_engine/tests provenance/case01/testsA fresh clone has no run data.
provenance/case01/runs/andprovenance/results/checkpoints/are deliberately gitignored (they are large, and were once submitted by accident), so the live UI at 5010 starts empty and a handful of evidence-dependent tests skip rather than run — that is expected, not a broken setup. To produce your own:python -m case01.run --run-id <id>(~1-2 min/run), or a batch viapython -m case01.tools.batch_run --duration-min 180 --lanes 3 --model qwen3:8b. Tests that need the missing data say which path they want (e.g.tests/test_metric_semantics.pypoints atresults/checkpoints) — feed it from your own runs, or drop the project-supplied demo package intoprovenance/case01/runs/(seedocs/专家导览_怎么看provenance.md).
-
Local Ollama (free, recommended for development): install Ollama and pull models
ollama pull qwen3:4b-instruct-2507-q4_K_M ollama pull qwen3-embedding:0.6b-q8_0
No configuration change needed (Ollama is the default).
-
OpenRouter (or any OpenAI-compatible API), needs a key: one command writes the key and verifies it online (the key is never printed):
python provenance/tools/setup_api.py --key sk-xxxx python provenance/tools/setup_api.py --show # just show where the key comes fromOr copy
.env.example(repo root) to.envand fill inOPENROUTER_API_KEY=...—.envis loaded automatically (it never overrides real env vars), so you do not need to know how to set environment variables. Both.envand.secrets.jsonare gitignored. Resolution order:env var → .env → .secrets.json(repo root / package root).
cd provenance/provenance
python live_switch.py --start case01 # 5010 is the single live entry (case00/case01 are mutually exclusive)
# UI only, no simulation: python live_switch.py --start case00 --no-simOpen http://127.0.0.1:5010/ . Read-only faces: python -m case00.serve --port 5003 /
python -m case01.serve --port 5002; the config tool lives in this repository:
cd provenance/config_tool && python app.py (8060). It discovers the co-located
platform directory; no environment variables are needed in the standard layout.
Experts/reviewers should read provenance/docs/专家导览_怎么看provenance.md first:
conclusion boundaries, which faces exist only on the case01/case00 side, the two read-only
data endpoints, how to read the fields, how to use the two manual-annotation tables, and a
fifteen-minute walkthrough.
Roles, relations and story are configured through web forms (no hand-written JSON). The tool lives in this repository:
cd provenance/config_tool
python app.py/— role configuration form (generates validated JSON)/relationships— relation input (appended to relationships.json)/story— story input (appended to story.json)/agents— list of configured roles
See provenance/config_tool/角色字段清单.md for the field list. config_tool
writes into this platform's provenance/frontend/static/assets/village/agents/
and provenance/scenarios/ by default (override with MAVIS_ASSETS_ROOT /
MAVIS_SCENARIOS_DIR). Restart the simulation server (5010) after adding roles.
| Option | Description |
|---|---|
--name |
simulation name (unique; checkpoints stored per name) |
--start |
starting time |
--stride |
game minutes per step (2 for finer detail) |
--step |
step count, 0 = run forever |
--resume |
resume from a checkpoint |
--port |
server port |
This platform is the reference implementation of IVD's process alignment story: AI value formation can be observed, governed and audited.
Expert-set constraints/expectations live in
provenance/governance.json (NOT in agent bodies). Each role maps to a
{goal: weight} vector summing to 1, where goal names are behavior-bound
(designed so embedding feedback can distinguish them — e.g. "Risk Control"
for stress-testing, "Data Rigor" for cross-verification):
{ "roles": { "AI投顾助手": { "Serve Users": 0.35, "Compliance Rigor": 0.3, "Risk Control": 0.2, "Data Rigor": 0.15 } } }Each role's agent.json also carries initial_tendency (persona baseline,
slightly offset from the constraints). On --resume, value_tendency and the
experience count are restored from the checkpoint so the tendency curve stays
continuous across restarts.
Constraints never enter the prompt; they only weight the consequence feedback, so an expert adjustment is felt by the agent through later experience (lagged convergence = internalization evidence).
The browser panel (right side) lets an expert:
- Read each role's value tendency (internalized result, read-only) as a live curve — one line per constrained goal, plus a stepped dashed line for the constraint expectation (steps at each expert intervention) and a vertical marker at each intervention time;
- Adjust constraint weights with sliders (sum enforced to 1; submitted on slider release, not per drag tick — avoids flooding the audit log);
- Export the tendency chart as PNG via the backend
(
GET /api/export-chart?agent=..., matplotlib-rendered, stepped constraint lines, compact bottom legend).
interventions.json— every expert edit:{time, sim_time, agent, old_constraints, new_constraints, operator};decisions.json— per-step decision stream withgoal_alignment(instant) andvalue_tendency(accumulated) for each role;- the tendency curve itself: lag between an intervention and the tendency's convergence is the observable evidence of internalization.
action → embedding similarity vs behavior-bound goals → relative share × weight → sliding window → tendency (blend with persona baseline) → prompt → action. Goal names are designed to be semantically distinguishable so the
embedding feedback can tell actions apart (see §7.1); scenario events and
role daily plans rotate behaviors to keep the curves lively instead of flat.
See the engine's README §7 for the formalization.
GET /api/explain?agent=<name> returns three explanation layers for why a
role's value tendency is what it is:
- Decomposition —
tendency = α×persona baseline + (1−α)×experience window mean, per goal, with α and cumulative experience count; - Window details — recent experiences (action description, per-goal alignment, feedback) that drove the internalization;
- Intervention chain — each expert intervention with constraint jump, tendency before/after 2h, and the quantified shift (lagged internalization evidence).
The browser panel shows these via the "解释倾向成因" button per role.
| Component | Notes |
|---|---|
| Python 3.12 + venv | pip install -r requirements.txt + build/install mavis wheel |
| LLM | Local Ollama (qwen3-instruct + qwen3-embedding) or OpenAI-compatible API (set in data/config.json, see §3) |
| Frontend assets | Vendored locally (static/vendor/: phaser/jquery/bootstrap) — no CDN dependency |
# from provenance/provenance
python live_fastapi.py --name stock-en6 --resume --step 0 --port 5010
# fresh sim (no --resume) starts at the configured date; --step 0 = run foreverBehind a reverse proxy (nginx/caddy) for HTTPS when embedding into an external platform. The service is self-contained (FastAPI + WS + static); no build step needed.
The service exposes dedicated embed routes — slim pages that reuse the same
WebSocket/data but hide unrelated UI. Embed via <iframe> from any web
platform (e.g. a governance dashboard); iframe pages connect their own WS, so
no CORS setup is required.
| Route | Content |
|---|---|
/embed/scene |
Phaser canvas only (no floating panels) — for a "simulation" slot |
/embed/goals |
Governance panel only (sliders + tendency curve + explain button) |
/embed/explain |
Governance panel with the explanation panel auto-expanded |
/embed/timeline |
Intervention timeline panel (all agents, full page) — for an "audit trail" slot |
Example (React/Next.js):
<iframe src="https://sim.example.com/embed/scene" style={{width:'100%',height:'480px',border:0}} />
<iframe src="https://sim.example.com/embed/goals" style={{width:'380px',height:'70vh',border:0}} />Deployment topology: run provenance on its own domain; the host platform embeds it. This keeps the two codebases independent while sharing the same live simulation.
- Real-time visualization via WebSocket (
/ws) pushing engine contract messages (agent/time/chat_line/snapshot); the client watchdog reloads on dead connections. Server sends an independent heartbeat every 5s (asyncio task, not queue-driven — reliable even during long LLM-only gaps such as schedule building); client 20s staleness timeout + focus-return check - The live service is driven by mavisframework (Game + Simulator + LiveCompressor)
- API endpoints:
GET /api/goals— constraints/tendency/interventions (scoped to the current simulation viasimulationfield)/role_types/embedding_healthPOST /api/goals— expert constraint edit → writes governance.json + interventions.json audit (withsimulationtag + optionalnotereason); rejects numeric/zero garbage goals; sum must equal 1POST /api/undo-intervention— roll back a past intervention to itsold_constraints(matched by agent + sim_time + record time); appends anoperator=undoaudit record and marks the original recordrevoked— history is never deleted, only amendedGET /api/timeline— intervention timeline across all agents (sorted by sim_time): each event carries old→new constraints, note, operator, and tendency shift (same windowing as/api/explain); revoked/undo events flagged for the frontendGET /api/export-chart?agent=<name>— matplotlib PNG of tendency curveGET /api/explain?agent=<name>— explainability panel: tendency decomposition (α blend), experience window details (action/alignment/ feedback), intervention causal chain (tendency shift after each intervention). Checkpoint series loading is cached by directory mtime.
- Decision export:
decisions.json(time/role/action/others/importance) for governance platforms and expert UI - Tests:
tests/test_live_api.py(pytest, fake server injection, no real simulation/LLM needed) covers goals read/write, intervention scoping, explain three layers, export error handling. Engine tests live in the mavis repo (tests/). - Phaser script: the server prefers the local
frontend/static/vendor/phaser.min.js(works offline) and falls back to CDN. For offline use, downloadhttps://cdn.jsdelivr.net/npm/phaser@3.55.2/dist/phaser.min.js(~1.3MB) into that folder before first run - Localization: modify the engine's
mavisframework/prompt/scratch.pyand frontend copy; no logic changes required - Role/scenario config tool: agents, relationships and story events are
generated by the in-repository
provenance/config_tool/service (port 8060); it writes directly into this platform'sagents/,scenarios/, andcases/directories.
- Follow the maze.py logic in the original generative_agents project to support tiled-exported json/csv files
- Follow the existing maze.json format to merge tiled exports (maze_meta_info.json, collision_maze.csv, sector_maze.csv) into a new maze.json
- Recommended: use the bundled converter
tools/tilemap_to_maze.py(CLI, no external deps) — converts a Tiled.tmx/.jsonmap intomaze.jsondirectly (seetools/tilemap_to_maze_README.md). The legacy GUI tool is at https://github.com/jiejieje/tiled_to_maze.json
- Paper: Generative Agents: Interactive Simulacra of Human Behavior
- Code: mavisframework (self-developed engine) / Generative Agents (original) / wounderland
- Map tool:
tools/tilemap_to_maze.py(bundled) / tiled_to_maze (legacy GUI)
(2026-09-23)
None of these services has authentication. Anyone who can reach a port can read every run record; 5010 additionally lets them rewrite governance weights, undo interventions, mark reflections and restart a run (4 write endpoints in total); 8060 (config tool) can edit scenarios, delete roles and launch a run.
- Loopback-only by default. Binding a non-loopback address requires an explicit
LIVE_ALLOW_REMOTE=1; otherwise the process refuses to start and prints exactly what would be exposed (a one-line warning is not enough — the port would already be open). - CORS defaults to loopback origins only (when
EMBED_ALLOW_ORIGINSis unset). For cross-origin data access from the platform, setEMBED_ALLOW_ORIGINS=https://<platform-host>. Plain<iframe>embedding does not use CORS and is unaffected. - Prefer not exposing ports at all for cross-machine integration: same-host deployment,
a read-only reverse proxy (only
/embed/*andGET /api/*), or an SSH tunnel to 5010. If you must expose it, use the origin allowlist plus a firewall rule per source IP — and never expose 8060. - Secrets: the OpenRouter key lives in
case01/.secrets.json(gitignored, not in git); a scan of 5000+ artifacts and logs found no key material.
Code locations: live/netguard.py (binding & allowlist policy), tests/test_netguard.py
(behaviour assertions). Integration details: provenance/docs/给平台侧_嵌入与数据接入.md;
health-check conclusions: provenance/docs/0923_体检报告.md.
Apache License 2.0, see LICENSE.