Skip to content

Repository files navigation

Terfyn

CI Release Go 1.22+ License: Apache-2.0 Go Reference

A statically analyzable, capability-oriented execution platform for nondeterministic programs. Review the authority granted to agents as a plan diff — before they run.

Terfyn bounds and diffs that grant. It does not verify what remote systems do with it.

Architecture

Source graph → validate / plan → SQLite desired state → apply → engine → tools + models → trace / logs / audit.

flowchart LR
  Source[Source graph] --> VP["validate / plan"]
  VP --> SQLite[(SQLite desired state)]
  SQLite --> Apply[apply]
  Apply --> Engine[engine]
  Engine --> Tools[tools + models]
  Engine --> Trace[trace / logs / audit]
Loading

Expanded diagram, plan-time bounds, and closed-world caveats: docs/architecture.md. Product spec: docs/DESIGN_DOC.md.

The differentiator: plan-time bounds on authority

This is not another orchestrator (Temporal, Dagger, LangGraph). Those schedule work. Terfyn's direction is a plan-time effect bound — a sound static upper bound on what an autonomous agent can do, reviewable as a diff (#189 / #191). terfyn plan prints that bound and the authority delta vs stored deployment state.

Today terfyn plan diffs permissions, approvals, models, budgets, C1 risk items, and effect/capability/authority against SQLite desired state:

Plan: 0 to add, 3 to change, 0 to delete
~ update Agent/reviewer
    spec.model: "mock/gpt-4" -> "mock/gpt-4o"
    spec.tools.1:  -> "github"
~ update Policy/default
    spec.approvals.requiredFor.0: "tool.helper.echo" -> "tool.github.issues.write"
    spec.execution.maxTotalCostUsd: 3 -> 10
~ update Tool/github
    spec.permissions.allow.1:  -> "issues.write"

Risk delta:
high:
- [high] approval_removal: Approval requirements removed for "tool.helper.echo" (Policy/default).
- [high] budget_relaxation: Cost ceiling increased (Policy/default).
- [high] permission_widening: New write-like tool permission "issues.write" added (Tool/github).
- [high] tool_surface_change: Agent tools list gained write-like tool "github" (Agent/reviewer).
medium:
- [medium] model_change: Agent model changed (Agent/reviewer).

When the graph declares tool operations, plan also prints the effect bound and an authority delta (bound(desired) vs bound(deployed)). Capability changes and effect changes are separate lines; AUTONOMOUS WIDENED means a nondeterministic agent's action space grew — including when a new grant does not add named effects.

Effect bound (Workflow/pr-review):
high:
- [high] effect_bound: github.write       autonomous  Agent/reviewer may select tool.github.post_comment
medium:
- [medium] effect_bound: github.read        static      step fetch_pr

Capability delta:
Agent/reviewer
+ tool.github.post_comment

Authority:
  static      -> unchanged
  autonomous  -> WIDENED

Agents and workflows are authored in .agent, the surface syntax fixed by ADR 002; the loader compiles .agent (type/effect checking + argument rebind) into the resource graph. Workflows run end-to-end, including conditionals, loops, and dynamic fan-out (#199/#259): a control-flow workflow lowers to the execution IR, is pinned into the deployment snapshot, and runs on the execir interpreter (see examples/agent-control-flow). YAML is the compilation output and interchange format (ADR 003): the loader still accepts it (so machine-generated resources and the existing fixtures work), and terfyn export --format yaml materializes the compiled graph on demand. Lead on capability, not format.

Flagship: incident triage

Start here: examples/incident-triage — an offline agent that can page, read logs, and file a ticket, but cannot restart a service unless policy approvals.requiredFor includes tool.restart.restart and that uses string is pre-approved with --approve tool.restart.restart. Inner agent-loop tools do not HITL; unapproved terfyn run fail-closes with exit 5 (approval_required). With --approve it completes and audit verify passes.

Other walkthroughs: policy-blocked PR review in examples/pr-review-demo (no API keys); live GitHub read/write with a mock reviewer in examples/pr-review-github; the same flow with OpenAI gpt-4o-mini plus GitHub Actions in examples/pr-review-github-actions (PR workflow; optional manual owner/repo/number).


Why this exists

Most agent stacks bury prompts, tool wiring, and permissions in application code. That makes it hard to answer: Is this config valid? What authority changed? What are we about to grant? What actually ran? Did policy allow it?

Terfyn is a small Go CLI (terfyn) plus a resource graph so teams can:

  • Review capability diffs (plan) before changes land
  • Track deployment state separately from runtime traces
  • Enforce policies (budgets, approvals, tool rules) at execution time
  • Stay local-first while the architecture leaves room for a future remote control plane

Mental model

Idea Analogy
Desired resources in Git GitOps
plan / apply / drift Terraform
Typed resources (Project, Workflow, Policy, …) Kubernetes-style API
Tool and IO contracts OpenAPI-style explicitness

Compared with adjacent tools

Terfyn is the declarative governance/config layer for agent systems — not a replacement for durable-execution engines or code-first agent runtimes.

Capability This project OpenAI Agents SDK LangGraph Temporal Terraform
Role Governance/config: versioned resources, plan/apply, policy Code-first agent runtime Code-first graph orchestration Durable workflow execution Infrastructure as code
Durable execution / distributed scheduling No No Optional checkpointers; not a durable-execution engine Yes N/A
Code-first agent runtime No (resource graph; .agent authoring, YAML compilation output) Yes Yes Workflow SDK, not an agent runtime No
Desired-state plan / apply Yes (terfyn plan / apply vs SQLite) No No No Yes
Plan-time effect bound Shipped (#189 / #190 / #191): bound over the callable operation set, including autonomous tool selection; terfyn plan prints the bound and authority delta. No listed comparable. No No No No

The bound is not over what those operations do at the far end; the trust anchor is human review of the tool manifest. Manifest pin (#204) is not enforced today (MCP tools/list can still expand the world). terfyn plan diffs permissions, approvals, models, budgets, C1 risk items, and effect/capability/authority.


Features (MVP today)

  • terfyn init — scaffold a .agent-led project (main.agent workflow plus project.yaml, policies, tools)
  • terfyn export --format yaml — materialize the compiled resource graph as YAML on demand (nothing written to disk by default; --output DIR writes a loadable project)
  • terfyn fmt — format .agent sources to canonical form (and normalize project YAML)
  • terfyn validate — load project, apply project defaults (spec.defaults), then environment overlays (-e / --env, Environment resources §7.6), then validate graph, schemas, and references; runs policy lint (ungated sensitive tools, invalid HITL config, etc.) as advisory output — use --strict to exit 2 on high-severity lint findings (fail-closed safety metadata still gates at run even when lint passes)
  • terfyn plan — diff desired graph vs SQLite deployment state; risk hints including policy lint, effect bound, and authority delta; JSON/YAML output includes policyLint, deploymentBaseline, effectBound, and authority
  • terfyn apply — persist plan (TTY confirm or --auto-approve / TERFYN_AUTO_APPROVE); optimistic concurrency — if the deployment store changed after the plan snapshot (e.g. another process applied the same --state file while this run waited at the prompt), apply fails with exit code 3; re-run plan then apply
  • terfyn run — execute a workflow locally; JSON Schema for inputs where configured; policy gates pause for human-in-the-loop (HITL) approval when a tool call requires it
  • terfyn logs — read trace events from SQLite (--run, --workflow, or recent runs)
  • terfyn audit verify — re-walk hash-linked trace chains and detect tampering (see docs/AUDIT_CHAIN.md)
  • Toolsnative, http, mock, and mcp — MCP supports stdio (subprocess) or streamable HTTP (spec.mcp.transport: http, url, optional headers with env: tokens)
  • Project defaults — besides model and policy, optional runtime flows to spec.runtime on agents/workflows when omitted (MVP: local or unset; see spec validation)
  • Output — table, JSON, or YAML (-o / --output)
  • State — single SQLite file (default .agentic/state.db under the project root; override with --state)
  • Tests — unit/integration coverage, golden CLI output tests, end-to-end init → … → logs in test/integration

See section 18 (MVP) and section 19 (End Goal) in docs/DESIGN_DOC.md for the full included/excluded list.


Quick start

Prerequisites

  • From source: Go 1.22+
  • From a release binary: no Go toolchain; put terfyn (or terfyn.exe) on your PATH after extracting the archive

Build

git clone https://github.com/LAA-Software-Engineering/terfyn.git
cd terfyn
make build   # writes bin/terfyn

Or: make install / go install ./cmd/terfyn (honours GOBIN / GOPATH/bin; ensure that directory is on PATH).

Prebuilt binaries

GitHub Releases ship terfyn for common platforms (.tar.gz on Linux/macOS, .zip on Windows) plus SHA256SUMS.txt. Pick the archive that matches your machine, for example:

Platform Asset suffix
Linux x86_64 linux-amd64.tar.gz
Linux arm64 linux-arm64.tar.gz
macOS Intel darwin-amd64.tar.gz
macOS Apple Silicon darwin-arm64.tar.gz
Windows x86_64 windows-amd64.zip (contains terfyn.exe)

terfyn version reports the release tag (e.g. v0.1.4).

Releases are created automatically when changes land on main, using a patch semver bump over the latest vMAJOR.MINOR.PATCH tag (merges that only touch Markdown or the root Makefile do not trigger a release). To cut minor or major bumps on demand, run the Release workflow manually (Actions → Release → Run workflow) and choose the bump type.

Create a project and run the loop

From the repo root (or anywhere):

terfyn init my-agent-system
terfyn validate --project my-agent-system
terfyn plan   --project my-agent-system
terfyn apply  --project my-agent-system --auto-approve
terfyn run    workflow/hello --project my-agent-system
terfyn logs   --project my-agent-system --workflow hello
terfyn audit verify --project my-agent-system --run <run-id>
terfyn inspect --web --project my-agent-system   # read-only local UI on http://127.0.0.1:8787

inspect --web binds to localhost only and opens the state DB read-only. Avoid running it while terfyn run is writing the same SQLite file (you may see database is locked without WAL); use it when runs are idle or on a copy of the DB.

Authoring: .agent plus project.yaml

Agents and workflows are authored in .agent (grammar reference); .agent files anywhere under the project root are discovered and compiled automatically. Tools, policies, and project configuration stay YAML (the interchange format). After terfyn init my-agent-system, my-agent-system/main.agent is:

workflow hello() {
    helper.echo(message: "hello")
}

and my-agent-system/project.yaml holds the config — a Project resource with spec.imports listing the YAML resources (policies, tools); .agent sources are not imported, they are discovered:

apiVersion: agentic.dev/v0
kind: Project
metadata:
  name: my-agent-system
spec:
  imports:
    - ./policies/default.yaml
    - ./tools/helper.yaml
  defaults:
    policy: default
    model: openai/gpt-4o-mini
    runtime: local
  providers:
    models:
      openai:
        type: openai
        apiKeyFrom: env:OPENAI_API_KEY
      # Optional: Claude via Messages API (set ANTHROPIC_API_KEY and use e.g. defaults.model: anthropic/claude-sonnet-4-20250514)
      # anthropic:
      #   type: anthropic
      #   apiKeyFrom: env:ANTHROPIC_API_KEY

To see the compiled resource graph as YAML — for inspection or handoff to another tool — run terfyn export --format yaml (it prints to stdout; nothing is written to disk unless you pass --output DIR). YAML remains valid ingress, so you can still author or generate resources directly in it.

Field-by-field rules, extra kinds, env overlays, MCP HTTP tools, and defaults.runtime are in docs/DESIGN_DOC.md. See docs/EXAMPLES.md for Anthropic fragments, MCP over HTTP, and structured-output notes.

Notes:

  • init creates my-agent-system/ with apiVersion: agentic.dev/v0 resources and a hello workflow (native echo tool only — no network).
  • apply in non-interactive environments needs --auto-approve or TERFYN_AUTO_APPROVE=1.
  • run HITL: gated tool calls exit with Status: interrupted (exit 0). Resume with --resume <run-id> --decision approve|reject|edit|switch (use --decision-edit-json / --decision-switch-target when needed), or skip prompts with --auto-approve / TERFYN_AUTO_APPROVE=1. Pre-approve a specific call with repeated --approve <uses>. Set TERFYN_HITL_ACTOR to attribute decisions in trace logs.
  • Policy.spec.hitl.interruptOn keys are Tool metadata.name values; they configure review options (edit rules, switch targets) for calls already gated by approvals.requiredFor or safety metadata — they do not gate tools on their own.
  • run stores traces in the same SQLite file used for plan/apply (default .agentic/state.db under --project). Optional OTLP export (spec.telemetry, off by default) is additive only — see docs/OTEL.md. When enabled you need serviceName plus either consoleExport: true or an endpoint (https://… or env:VAR, e.g. env:OTEL_EXPORTER_OTLP_ENDPOINT). Export that variable before run if you use env:; if it is missing or the collector is unreachable, terfyn logs a warning, skips OTLP, and the workflow still completes (SQLite traces unchanged).
  • If spec.traces.retentionDays is a positive integer, runs older than that many UTC calendar days (by runs.started_at) are deleted lazily on run and logs (child trace rows cascade). Unset or non-positive means no pruning.
  • Trace payload redaction (issue #110): before SQLite storage, event JSON is sanitized, key-redacted, and size-capped. Defaults mask common secret key names (substring match on map keys). Optional project knobs:
    • spec.traces.redactKeys / maxPayloadBytes — merged with defaults; also available under spec.traces.redaction together with maxDepth, maxStringChars, and maxBytes (max bytes for binary previews in sanitized values, not the overall JSON cap).
    • Stored events may show [REDACTED], payload_truncated / preview, or depth/binary placeholders in logs / inspect --web.
  • Use logs --run <id> after a run if you want a single run’s trace (IDs are printed by run).

Global flags (common)

Flag Purpose
--project <path> Project root (default .)
--state <path> SQLite file override
-e / --env Environment overlay name
-o / --output table, json, or yaml
--no-color ASCII-friendly validate output

Exit codes are summarized in section 11.2 of docs/DESIGN_DOC.md (0 success, 2 validation, 3 plan/apply conflict when deployment state changed after plan or resolved config drifted before run, 4 execution, 5 policy denial, …).

User-local config (per-developer overrides)

Config is resolved in this order (highest wins): CLI flagsenvironment overlay (-e) → project YAMLuser-localbuilt-in defaults.

Optional user-local files (git-ignored, strict YAML — typos fail validate):

Path Scope
$XDG_CONFIG_HOME/terfyn/config.yaml or ~/.config/terfyn/config.yaml Global per-user defaults (defaults, state, providers, traces, telemetry)
.agentic/local.yaml under --project Project-scoped overrides (same fields; wins over the global file)

validate, plan, and apply write .agentic/resolved-config.json (digest of the resolved graph + env + state path). run rejects drift from that snapshot with exit 3 — re-run validate or plan after changing config.


Repository layout (high level)

Path Role
cmd/terfyn CLI entrypoint
internal/cli Cobra commands, flags, golden tests
internal/spec YAML types, normalize, validate
internal/config Layered config resolution, immutable snapshot
internal/project Load project + imports
internal/plan Planner and risk summary
internal/apply Apply plan to deployment store
internal/engine Workflow execution
internal/policy Policy evaluation
internal/state/sqlite SQLite deployment + runtime/trace tables
internal/audit Tamper-evident hash chain for trace events (issue #116)
test/integration End-to-end CLI flow tests
docs/DESIGN_DOC.md Spec, CLI UX, architecture, roadmap
docs/GITHUB_ACTIONS.md Running terfyn from GitHub Actions (tokens, exit code 5, template path)
examples/pr-review-github-actions/ Full gpt-4o-mini project; PR workflow .github/workflows/terfyn-pr-review.yml; optional publish .github/workflows/terfyn-pr-review-publish.yml

Development

make defaults to help, which lists targets; the table below mirrors the Makefile (## comments and recipes).

Target What it does
help Show usage and target list (default goal)
all fmtvettestbuild (handy before a push)
build go buildbin/terfyn
install go install ./cmd/terfyn (-trimpath; uses GOBIN / GOPATH/bin)
clean Remove bin/ and coverage.out
fmt go fmt ./...
verify-fmt Fail if gofmt -l would list files (matches CI-style formatting check)
vet go vet ./...
test go test ./... -race
test-coverage Tests with -coverprofile=coverage.out and a one-line go tool cover -func summary
check vet + test only (no formatting writes)
ci verify-fmt + vet + test (no build)

CI (.github/workflows/ci.yml) runs Linux, macOS, and Windows on Go 1.22.x, plus Go 1.23.x on Linux, with race and shuffle enabled (workflow steps are defined in YAML, not via make ci).

Updating CLI golden files

When table output is intentionally changed:

GO_UPDATE_GOLDEN=1 go test ./internal/cli/... -run TestGolden_

Roadmap

Near term (MVP hardening)

Recent landings already cover much of “hardening”: plan/apply optimistic concurrency (exit 3 when deployment state drifts), MCP over streamable HTTP as well as stdio, trace retention (spec.traces.retentionDays), defaults.runtime / spec.runtime (MVP local), and clearer defaults vs environment overlay documentation. What is still open for near-term polish:

  • More diff / drift UX where the design doc calls for it (beyond today’s resource-level diff)
  • Richer logs filtering (see sections 10.2 and 17.3 in docs/DESIGN_DOC.md); inspect --web covers read-only run/state browsing (#109)
  • terfyn test-style workflow fixtures (stretch per design doc)

Post-MVP (from design doc section 19)

  • Modules/registry, remote shared state, reconciliation controllers
  • Parallel steps, subworkflows, schedules/events
  • Stronger drift semantics and multi-runtime targets
  • Deeper approval workflows and multi-tenant controls

The recommended implementation phases are outlined in section 20 of docs/DESIGN_DOC.md.


Documentation

  • docs/DESIGN_DOC.md — design document v0 (problem statement, spec, CLI, engine, state model, testing strategy, MVP vs end state, section 23 recommendation).
  • docs/AGENT_LOOP.md — bounded agent tool-calling loop: grants, advertised uses, maxIterations, policy on every inner call, traces, HITL vs exit 5 (issues #160 / #161 / #175).
  • docs/AUDIT_CHAIN.md — hash-linked trace audit chain and terfyn audit verify (issue #116).
  • docs/ATTRIBUTION.md — tenant, thread, and actor fields on runs and traces (issue #111).
  • docs/OTEL.md — optional OTLP trace export alongside SQLite (issue #108).
  • docs/TESTING.mdterfyn test fixture format; CI gate walkthrough in examples/regression-test.
  • examples/pr-review-demo/README.md — end-to-end demo: structured review output, traceable run, approval-gated write (validateplanapplyrunlogs).
  • examples/regression-test/README.mdterfyn test is green on a gated publish and red after dropping requiredFor (issue #176).
  • docs/EXAMPLES.md — copy-paste YAML and CLI examples (init, mock vs OpenAI, workflows, environment overlays).
  • CODE_OF_CONDUCT.md — Contributor Covenant 2.1; participation expectations and reporting.
  • License: MIT

Contributing

Issues and pull requests are welcome. See CONTRIBUTING.md for local setup, tests, golden updates, and pull request expectations.

About

A statically analyzable execution platform for nondeterministic programs, with reviewable authority and reproducible deployment state.

Topics

Resources

Code of conduct

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Used by

Contributors

Languages