A machine-readable rulebook for AI coding agents, and the build checks that enforce it.
AI proposes a change. The file's own toolchain proves it. The build refuses what skipped a check.
Agents start at llms.txt; chats at CHAT.md.
Thea tells an AI coding agent which commands prove a change to a file, and fails the build when a
change skipped them. One declaration, atlas.yaml, answers what proves this change
is correct?
---
config:
theme: base
themeVariables:
primaryColor: "#e6f2e7"
primaryBorderColor: "#6f9f73"
primaryTextColor: "#14301a"
lineColor: "#7f9483"
fontSize: "15px"
flowchart:
padding: 6
nodeSpacing: 14
rankSpacing: 18
---
flowchart TB
accTitle: How Thea is built
accDescr: the contract says what must be checked and what counts as proof, one core resolves the file, selects its gates, runs the checks and records the verdict, and the CLI, MCP, hooks with CI and the TheaOS all reach that one core
R[contract · atlas.yaml<br>what counts as proof] --> C[one core: resolve, gate,<br>check, record the verdict]
C --> doors
subgraph doors [ ]
direction TB
L[CLI<br>commands] ~~~ H[hooks · CI<br>commit, merge]
P[MCP<br>agent tools] ~~~ B[TheaOS<br>results, history]
end
classDef core fill:#cfe9d2,stroke:#4f9a58,color:#103d17
class C core
classDef band fill:none,stroke:#6f9f73,stroke-dasharray:4 3
class doors band
$ thea gate scripts/doctor.py
1. formatter: ruff format --check scripts/doctor.py
2. compiler_or_typechecker: python3 -c 'import ast,sys; [ast.parse(open(f, encoding="utf-8").read(), f) for f in sys.argv[1:]]' scripts/doctor.py
3. unit_tests: pytestNot an app framework or a runtime optimizer.
# no checkout
uv tool install thea-software && thea doctor
# or, editable
git clone --depth 1 --branch v3.53.0 https://github.com/HLIntel/thea-software ~/thea && uv tool install --editable ~/thea && thea doctorcd ~/thea
thea port scripts/doctor.py # route, gates, lessons
thea gate scripts/doctor.py # what proves a change
thea verify # exit 0 only if all PASSRead-only MCP: thea-mcp. Other repos: CONSUMING.
$ thea port scripts/doctor.py --line
◉ scripts/doctor.py │ ⠟backend │ python │ ⌂scripts │ ✓3 │ ⚠1 │ → thea gate scripts/doctor.py
$ thea port scripts --line
◎ scripts │ ⠟103 │ ⌂scripts │ → thea brainstorm
$ thea port . --line
○ . │ ⠟130 ⠿4 ⠁3 │ → thea checkRecorded runs, each naming its instrument: evidence for routing and refusals, not for end-to-end task success (limits).
With Thea: the model sees what thea gate prints. Blind: only the language names. Token savings
compare against pasting every language's tool list.
On Claude (76 questions per model, abtest.py v3.49.0)
- Opus: 100% right with Thea, 41% blind; reads 90% fewer tokens.
- Sonnet: 100% right with Thea, 39% blind; reads 90% fewer tokens.
- Haiku: 100% right with Thea, 39% blind; reads 91% fewer tokens.
- Claude Code start-up: loads
CLAUDE.mdand its imports, 1,053 tokens.
Beyond routing (blind → with Thea, taskbench.py v2.29.0)
- Name a failure from its symptom: Opus 93% → 100%; Sonnet 57% → 100%; Haiku 64% → 96%.
- List the checks a change needs: Opus 12% → 100%; Sonnet 12% → 100%; Haiku 0% → 100%.
- Spot a line the build refuses (yes/no, so a coin flip scores 50%): Opus 60% → 100%; Sonnet 60% → 100%; Haiku 40% → 100%.
- Not measured: visual design, open-ended strategy, arithmetic — nothing declares a right answer.
Across all 11 models tested (5 providers, 2,409 questions, abtest.py v2.27.0 / v2.28.0 / v3.49.0)
- Right answers: 99% (95% interval 98–99%) with Thea, 59% (95% interval 55–63%) blind; every model 97–100% with Thea. A random guess scores 2.8%.
- Tokens: 89% fewer than pasting every tool list, 51% fewer than blind.
The repository itself (recomputed on every build)
- 1,734 tokens read before routing; the other 205 documents (627 KiB) load only when a route names one.
- 378 language × check pairs (42 languages × 9 checks), all answered: 143 with a command, 235 with a declared no tool, 0 silently.
- 501 mistake kinds planted in the tests, each refused.
- 17/17 planted breaks refused in 12/42 languages; 11 unchecked (cloudflare, dockerfile, elixir, fsharp, haskell, json, markdown, sql, thea, toml, yaml), 19 no example (
enforce.py, v3.53.0). - 18/18 handoffs carry the right checks (schema only → with Thea): Opus 0/6 → 6/6; Sonnet 0/6 → 6/6; Haiku 0/6 → 6/6 (
workflowbench.py). - 24/24 solo commits clean with or without the hook on these tasks; a planted broken commit is refused.
- 113 failure shapes in the ledger: 197 sightings, 45 recurred; 94 guarded.
- In use (
agents.py --field, v3.53.0): 49 refusals (12 shapes), 32 re-fired; verify 13 pass / 7 fail; 12/12 lands armed; lessons unmeasured. - 6 agent controls that block, not warn: narrow_tools, sandbox, budget, approval, effects, audit.
- 10 KiB install: 1 module, 1 dependency.
- Plug in with
thea port <file>, or parse .agent/bootstrap.json; load only what it names. - Read machine output, not this page. Add
--jsonto any command: one record per gate, frozen in tools/atlas-output.schema.json. - Before a shell command,
thea shell --json "<cmd>": exit 3 means its verdict would be misread. - Land with
python scripts/branchstate.py --land; askthea landed <branch>before deleting one. - Judgments answer on rules, teacher or a keyless student.
- When something breaks, file it the same turn with the
theaskill.
Every number here is generated on each build; check fails when one drifts. Sites and dashboards
read the same figures from .agent/facts.json.
- 3.53.0 contract version —
VERSION, asserted at a declared line in 7 other files - 42 tool manifests —
languages/<route>/tools.yaml, validated againsttools/tools.schema.json - 401 declared tool entries — distinct entries per manifest, summed;
packprobe.pyclassifies every one - 5 entry kinds —
tools/tools.schema.json$defs.entry.x-kinds - 8 change classes (verification profiles) —
atlas.yaml/verification_policy/profiles - 14 task profiles —
atlas.yaml/task_profiles - 102 python files in the harness —
scripts/*.py, every one held by thelint·format·typecheckgates
- Every document: docs index · wiki
- Languages: atlas · routing · verification
- Agents: harness ·
.thea· runtimes - Why a rule exists:
thea why·thea failures· instruments - Security: policy · OpenSSF
The version tracks the contract, not the content: docs/VERSIONING.md.
Built by HLIntel LLC, public on purpose: no secret, private path or internal host enters
it (ABOUT.md). Contributions pass thea verify. MIT License.
