Skip to content

[Typed evaluation roadmap] Ship reusable, durable decision programs across Tangle #768

Description

@drewstone

Outcome

One caller-owned typed evaluator works in products, campaigns, trace analysts, workflows and agent checkpoints. Code retains authorization and action execution; generative models retain open-ended work. Prove downstream improvement, not only API success.

Ten milestones

  • M1: Released packages and real scoped calls work across product, campaign, workflow and graph with attributable receipts.
  • M2: Multi-day execution survives crashes and credential rotation without silently repeating paid work or external effects; unknown outcomes stay explicit.
  • M3: Independent labels establish judge errors, calibration and useful automation coverage with application-owned escalation policies.
  • M4: Evidence-linked trace triage feeds deep investigation and repairs verified by executing changed behavior.
  • M5: Context, memory and skill selection improves total task economics without losing mandatory constraints.
  • M6: Outcome-based compute allocation beats a fixed policy on independently checked results and total resources.
  • M7: Existing workflows mix code, native evaluations and ordinary agents with structured results and no duplicate runners.
  • M8: Existing optimization methods improve executable decision programs under independent final assessment and complete cost accounting.
  • M9: Verified production feedback advances exact candidates through shadow/limited rollout with exact rollback.
  • M10: Existing UI explains consequential decisions and turns failures into reproducible evaluation cases.

Canonical execution backlog

Behavioral validation extension: #774 covers independent/adversarial evidence for reward gaming, evaluator tampering, policy/authorization/data-boundary violations, injection-following, contradicted completion, oversight bypass, matched-baseline anomalies and cross-run hypotheses. These remain optional caller-owned recipes, not a default safety gate.

Current implementation status

Lane Current work Actual evidence
Eval #776 on fix/evaluation-definition-identity, head 17453b832080b45db8a36e3e84a709906da6c0c9 Full CI passed. Fixes static evaluator identity collisions from question/Choice reordering and post-construction options mutation. Five regressions failed on baseline; eight new tests pass after fixes.
Platform / Intelligence tangle-network/agent-dev-container#7698 on feat/native-workflow-results, head 6267f1d7fb61daef7399dcfd8f383df5b83b0a23 New output.evaluation works through shared action mappings without text parsing. Fixes native metadata loss and dropped child tool summaries (#7699 there). Nineteen focused local tests pass; final-head full CI/browser checks remain queued.
Router tangle-network/tangle-router#539 is merged; part of tangle-network/tangle-router#538 Seven real PostgreSQL tests were added to the existing CI job. Latest database/process/billing proof was not established in this pass; broader issue remains open.
Runtime tangle-network/agent-runtime#1294 and tangle-network/agent-runtime#1295 remain first Existing budget/hook/receipt owners retained. No new Runtime patch claimed in the latest pass.

Eval #776 full CI: https://github.com/tangle-network/agent-eval/actions/runs/35416744633
Platform #7698 queued CI: https://github.com/tangle-network/agent-dev-container/actions/runs/35417116721

Merged foundation and previous evidence

Current Platform tests execute real modified source in a partial checkout with injected runAgent and isolated unrelated boundaries. No local full-monorepo signoff or deployed trace-ingest proof is claimed. Actual committed/tested blob hashes and baseline failures are recorded in both PRs.

Parallel work and integration boundaries

A / Router: tangle-network/tangle-router#538. Reuse durable request, dispatch and billing stores; no parallel settlement implementation.
B / Eval: #769, #770 and #771 can run independently in non-overlapping modules. #772 consumes observations and measured quality.
C / Runtime: tangle-network/agent-runtime#1294 and #1295 in that repo first; then selectors/candidates; optimization after a measured baseline.
D / Platform: #7698 owns current shared action/output changes. Cross-key recovery coordinates with identity/budget owners; inspector follows the observation contract.

Use isolated worktrees/file ownership and one open integration PR per repository. Reload current main/develop and inspect concurrent work before writing. No new scheduler, registry, ledger, memory manager or question language. No auto-deploy or frozen-release bypass.

The earlier GitHub attempt to assign the official coding agent received HTTP 403. No autonomous coding worker was started; the listed PRs were authored directly. Workstream names are ownership boundaries, not claims of background execution.

Existing ownership retained

Reconciliation

Closure evidence

Record exact revision, commands/results and untested boundaries. Distinguish source/types, local tests, final-head CI, real database/process tests, deployed API receipts, wallet reconciliation and independent quality measurements. Merges, queued/skipped CI and model-generated judgments do not satisfy stronger proof. No infrastructure purchase, increased spend, live provider run, release or production promotion was performed in these follow-through patches.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions