Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# Changelog

## 0.243.0

Training, optimization, and bound harnesses now share candidate-validator admission and invocation. **Migration:** validators must return `undefined` synchronously or throw. Promises and other return values are rejected instead of silently accepting an unchecked candidate; use a block body for side effects. Composed optimizer leaves obey the same rule, while original callbacks remain bound into execution identity.

`improve` and `createProfileImprovementHarness` accept their own frozen profile results directly. Retraining updates small-model, subagent, and mode references to the prior receipted checkpoint, preserving unrelated model choices. Serving evidence is snapshotted before asynchronous checkpoint revalidation.

Training's `timeoutMs` is optional and uses the existing chunked deadline timer when supplied; there is no implicit deadline or seven-day ceiling. Cancellation still stops admission and waiting, without claiming remote cleanup. The command trainer snapshots direct-call requests before asynchronous work, honors caller-selected output budgets, accepts empty pinned configuration files, and streams declared input hashes without a one-gigabyte ceiling. Checkpoints remain bounded and nonempty; dataset bounds, partition checks, receipt identity, and publication ordering are unchanged.

Command training and coding harnesses share one confirmed process-group teardown implementation, allowing graceful final writes before escalation. Existing coding-harness timing policy and import paths are preserved. No new dependency, executor, training algorithm, dataset exporter, scheduler, or deployment client is added.

## 0.242.0

`improve(profile, { mode: 'training', ... })` and the bound harness's `train` method now produce a checkpoint-backed candidate with an Interface training receipt. The command trainer pins its executable and inputs, supplies only an explicit public environment, and cancels its POSIX process group. Managed trainers and verified serving adapters use the same typed boundary; they own remote job cleanup.
Expand Down
6 changes: 3 additions & 3 deletions api-surface.json
Original file line number Diff line number Diff line change
Expand Up @@ -48,7 +48,7 @@
"ConversationStreamEvent": "type a39182b3dbfc",
"ConversationTurn": "type d8280ca3c636",
"CreateKnowledgeImprovementActivationExecutorOptions": "type 4b3e5fe02df7",
"CreateProfileImprovementHarnessOptions": "type 36de01aba28e",
"CreateProfileImprovementHarnessOptions": "type ac085504184e",
"D1DatabaseLike": "type ffce9ec5de30",
"D1StmtLike": "type cd8b3c46cbcd",
"DEFAULT_ROUTER_BASE_URL": "value 4db78e51a917",
Expand Down Expand Up @@ -94,7 +94,7 @@
"ImproveScenarioPartitions": "type 37a3508406b1",
"ImproveSkillsOptions": "type c1f5a69faefc",
"ImproveSurface": "type b711b683b151",
"ImproveTrainingOptions": "type 8bfc42270e58",
"ImproveTrainingOptions": "type 0d216f385a91",
"ImproveTrainingResult": "type 81c28cf9d5ab",
"ImprovementCandidate": "type 0c22a91c6396",
"ImprovementCodeCandidate": "type 588fa6d3b2f5",
Expand Down Expand Up @@ -253,7 +253,7 @@
"formatSupervisedKnowledgeTask": "value bdcf6b28157d",
"generateSpanId": "value 2f8329045cac",
"getModels": "value 95cb7c012c48",
"improve": "value 45f11e9feb52",
"improve": "value 4c0da1dd4fbf",
"isDelegatedLoopMode": "value d0f2042750ec",
"knowledgeReadinessDeliverable": "value f6f33b24a926",
"loopEventToOtelSpan": "value 63bec8b09ae0",
Expand Down
4 changes: 4 additions & 0 deletions bench/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,9 @@
# Changelog

## 0.13.5

Consume Runtime 0.243.0 through the existing workspace dependency. Benchmark APIs and grading behavior are unchanged.

## 0.13.4

Require Interface `^2.10.0` and consume Runtime 0.242.0 through the published dependency ranges, keeping benchmark consumers on the checkpoint-training receipt contract.
Expand Down
2 changes: 1 addition & 1 deletion bench/package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@tangle-network/agent-bench",
"version": "0.13.4",
"version": "0.13.5",
"type": "module",
"description": "Benchmark adapters and execution for agent-runtime across coding, tool-use, RAG, memory, browser, and terminal tasks.",
"repository": {
Expand Down
16 changes: 9 additions & 7 deletions docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -3528,7 +3528,7 @@ Exact materialized profile presented for validation before any candidate run.

##### profile

> **profile**: `AgentProfile`
> **profile**: `object`

Exact baseline profile. It is parsed, detached, and frozen at construction.

Expand Down Expand Up @@ -3917,9 +3917,11 @@ Pins trainer, serving adapter and their private dependencies, just like the boun

> **outputDirectory**: `string`

##### timeoutMs
##### timeoutMs?

> `optional` **timeoutMs?**: `number`

> **timeoutMs**: `number`
Optional overall deadline. Omit to rely on caller cancellation; long durations are supported.

##### maxCheckpointBytes

Expand Down Expand Up @@ -3983,6 +3985,8 @@ Explicit public environment only. Ambient credentials are never inherited.

> **maxOutputBytes**: `number`

Total stdout + stderr byte budget. Output is drained, not retained in memory.

***

### CreateKnowledgeImprovementActivationExecutorOptions
Expand Down Expand Up @@ -7503,6 +7507,8 @@ Runs one exact materialized profile on one scenario.

> **ImproveCandidateValidator** = (`input`) => `void`

Accept by returning void synchronously; reject by throwing. Async callbacks are refused.

#### Parameters

##### input
Expand Down Expand Up @@ -9002,8 +9008,6 @@ Train and serve a checkpoint without implying that it improved held-out quality.

###### profile

`AgentProfile`

###### opts

[`ImproveTrainingOptions`](#improvetrainingoptions)
Expand Down Expand Up @@ -9032,8 +9036,6 @@ Optimize one exact profile surface with a complete method.

###### profile

`AgentProfile`

###### opts

[`ImproveMethodOptions`](#improvemethodoptions)\<`TScenario`, `TArtifact`\>
Expand Down
5 changes: 3 additions & 2 deletions docs/api/primitive-catalog.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@

# Primitive catalog — the never-stale anti-reinvention inventory

> **GENERATED** from `@tangle-network/agent-runtime@0.242.0` and `@tangle-network/agent-eval@0.182.0` by `scripts/gen-primitive-catalog.mjs`. Do NOT hand-edit — run `pnpm run docs:api`. This is the mechanical companion to the JUDGMENT in `canonical-api.md` (§2 decision table + §1.5 AgentProfile law): that doc says WHICH primitive to reach for and what NOT to build; this catalog proves WHAT exists. Per-symbol signatures + `file:line` live in the per-module pages under `docs/api/`.
> **GENERATED** from `@tangle-network/agent-runtime@0.243.0` and `@tangle-network/agent-eval@0.182.0` by `scripts/gen-primitive-catalog.mjs`. Do NOT hand-edit — run `pnpm run docs:api`. This is the mechanical companion to the JUDGMENT in `canonical-api.md` (§2 decision table + §1.5 AgentProfile law): that doc says WHICH primitive to reach for and what NOT to build; this catalog proves WHAT exists. Per-symbol signatures + `file:line` live in the per-module pages under `docs/api/`.

## 1. agent-runtime — own public surface

Expand Down Expand Up @@ -156,6 +156,7 @@ Import from `@tangle-network/agent-runtime` — 298 exports.
| `AgentEvalErrorCode` | type | Error taxonomy for `@tangle-network/agent-eval`. |
| `AgenticGeneratorShotDisposition` | type | Worktree decision emitted before a completed shot is retried, accepted, or |
| `AgenticGeneratorShotExecution` | type | Runtime's exact terminal turn plus its complete normalized event stream. |
| `ImproveCandidateValidator` | type | Accept by returning void synchronously; reject by throwing. Async callbacks are refused. |
| `ImproveCodeRunOptions` | type | Runtime-owned code search in isolated git worktrees. |
| `ImproveMethodFactory` | type | Build a complete method after trace findings are available. |
| `ImproveMethodOptions` | type | Complete-method configuration for every non-code profile surface. |
Expand All @@ -175,7 +176,7 @@ Import from `@tangle-network/agent-runtime` — 298 exports.
| `Verifier` | type | Verifies the edited worktree. Sync or async; throws only on a setup fault |
| `WorktreeCheckRunner` | type | The single shell-command-in-worktree runner seam (replaces the per-executor copies). |

**Undocumented supporting types** (add a TSDoc line at the declaration to earn a table row): `AgentAdapter`, `AgentBackendContext`, `AgentBackendInput`, `AgentExecutionBackend`, `AgenticGeneratorOptions`, `AgenticGeneratorShotReceipt`, `AgentKnowledgeProvider`, `AgentKnowledgeReadinessCheckOptions`, `AgentTaskContext`, `AgentTaskRunResult`, `AgentTaskSpec`, `BackendCallPolicy`, `ChatModelCandidate`, `CheckpointServingPort`, `ControlBudget`, `ControlEvalResult`, `ControlledTrainingCommand`, `ControlRunResult`, `ControlStep`, `Conversation`, `ConversationDriveState`, `ConversationJournal`, `ConversationJournalEntry`, `ConversationParticipant`, `ConversationPolicy`, `ConversationResult`, `ConversationTurn`, `CreateKnowledgeImprovementActivationExecutorOptions`, `CreateProfileImprovementHarnessOptions`, `D1StmtLike`, `DataAcquisitionPlan`, `DelegatedLoopResult`, `EvalRunEvent`, `EvalRunGeneration`, `EvalRunsExportConfig`, `EvalRunsExportResult`, `HaltContext`, `HaltSignal`, `ImproveCodeBaseOptions`, `ImproveCodeResult`, `ImproveCustomCodeGeneratorOptions`, `ImprovementCodeCandidate`, `ImprovementProfileCandidate`, `ImproveMethodContext`, `ImproveMethodResult`, `ImproveRuntimeCodeGeneratorOptions`, `ImproveSkillsOptions`, `ImproveTrainingOptions`, `KnowledgeImprovementActivationExecutor`, `KnowledgeImprovementCandidatePair`, `KnowledgeImprovementExperimentBundles`, `KnowledgeImprovementJobMeasurement`, `KnowledgeImprovementJobResult`, `KnowledgeReadinessCheckInput`, `KnowledgeReadinessDecision`, `KnowledgeReadinessReport`, `KnowledgeRequirement`, `LoopRunnerCliArgs`, `LoopRunnerCliResult`, `McpServeSpec`, `OfficialSensitiveCandidateInput`, `OtelAttribute`, `OtelExportConfig`, `OtelExporter`, `OtelSpan`, `PersonaConversationResult`, `ProfileTrainerRequest`, `RawTraceDistillerOptions`, `ReflectiveGeneratorOptions`, `ResearchLoopResult`, `ResearchLoopRunnerOptions`, `ResolvedChatModel`, `RunAgentTaskOptions`, `RunAgentTaskStreamOptions`, `RunConversationOptions`, `RunDelegatedLoopOptions`, `RunKnowledgeImprovementJobOptions`, `RunPersonaConfig`, `RunPersonaConversationOptions`, `RuntimeDecisionEvidenceRef`, `RuntimeDecisionPoint`, `RuntimeEventCollector`, `RuntimeEventOtelOptions`, `RuntimeHookContext`, `RuntimeHookErrorContext`, `RuntimeHookEvent`, `RuntimeRunCompleteInput`, `RuntimeRunCost`, `RuntimeRunHandle`, `RuntimeRunOptions`, `RuntimeRunPersistenceAdapter`, `RuntimeRunRow`, `RuntimeSession`, `RuntimeSessionStore`, `RuntimeStreamEventCollector`, `RuntimeStreamEventSummary`, `RuntimeTelemetryOptions`, `SanitizedKnowledgeReadinessReport`, `SanitizedKnowledgeRequirement`, `ServerSentEventOptions`, `SupervisedKnowledgeUpdateInput`, `SupervisedKnowledgeUpdateOptions`, `SupervisedKnowledgeUpdateResult`, `TrainingDatasetDocument`, `VetoedFact`, `WorktreeLoopRunnerOptions`, `AgenticGeneratorExecutorForWorktree`, `AgentRuntimeEvent`, `AgentRuntimeEventSink`, `AgentTaskStatus`, `AuthSource`, `ChatModelValidation`, `ControlDecision`, `ConversationStreamEvent`, `DeepReadonly`, `DelegatedLoopMode`, `DelegatedLoopRegistry`, `DelegatedLoopRunner`, `HaltPredicate`, `HaltReason`, `ImproveCandidateValidator`, `ImproveCodeOptions`, `ImprovementCandidate`, `ImprovementProfileCandidatePopulation`, `ImprovementProfilePopulationCandidate`, `ImprovementProfilePopulationLineage`, `ImproveMethodSource`, `ImproveOptimizationRunOptions`, `ImproveProfileSurface`, `ImproveResult`, `ImproveTrainingResult`, `KnowledgeReadinessCheck`, `KnowledgeReadinessCheckResult`, `ProfileImprovementHarnessRunOptions`, `ProfileImprovementHarnessTrainOptions`, `RuntimeDecisionKind`, `RuntimeHookTarget`, `RuntimeRunStatus`, `RuntimeStreamEvent`, `RuntimeStreamEventSink`, `SupervisedKnowledgeUpdater`, `TrainingBoundaryResult`, `TurnOrder`.
**Undocumented supporting types** (add a TSDoc line at the declaration to earn a table row): `AgentAdapter`, `AgentBackendContext`, `AgentBackendInput`, `AgentExecutionBackend`, `AgenticGeneratorOptions`, `AgenticGeneratorShotReceipt`, `AgentKnowledgeProvider`, `AgentKnowledgeReadinessCheckOptions`, `AgentTaskContext`, `AgentTaskRunResult`, `AgentTaskSpec`, `BackendCallPolicy`, `ChatModelCandidate`, `CheckpointServingPort`, `ControlBudget`, `ControlEvalResult`, `ControlledTrainingCommand`, `ControlRunResult`, `ControlStep`, `Conversation`, `ConversationDriveState`, `ConversationJournal`, `ConversationJournalEntry`, `ConversationParticipant`, `ConversationPolicy`, `ConversationResult`, `ConversationTurn`, `CreateKnowledgeImprovementActivationExecutorOptions`, `CreateProfileImprovementHarnessOptions`, `D1StmtLike`, `DataAcquisitionPlan`, `DelegatedLoopResult`, `EvalRunEvent`, `EvalRunGeneration`, `EvalRunsExportConfig`, `EvalRunsExportResult`, `HaltContext`, `HaltSignal`, `ImproveCodeBaseOptions`, `ImproveCodeResult`, `ImproveCustomCodeGeneratorOptions`, `ImprovementCodeCandidate`, `ImprovementProfileCandidate`, `ImproveMethodContext`, `ImproveMethodResult`, `ImproveRuntimeCodeGeneratorOptions`, `ImproveSkillsOptions`, `ImproveTrainingOptions`, `KnowledgeImprovementActivationExecutor`, `KnowledgeImprovementCandidatePair`, `KnowledgeImprovementExperimentBundles`, `KnowledgeImprovementJobMeasurement`, `KnowledgeImprovementJobResult`, `KnowledgeReadinessCheckInput`, `KnowledgeReadinessDecision`, `KnowledgeReadinessReport`, `KnowledgeRequirement`, `LoopRunnerCliArgs`, `LoopRunnerCliResult`, `McpServeSpec`, `OfficialSensitiveCandidateInput`, `OtelAttribute`, `OtelExportConfig`, `OtelExporter`, `OtelSpan`, `PersonaConversationResult`, `ProfileTrainerRequest`, `RawTraceDistillerOptions`, `ReflectiveGeneratorOptions`, `ResearchLoopResult`, `ResearchLoopRunnerOptions`, `ResolvedChatModel`, `RunAgentTaskOptions`, `RunAgentTaskStreamOptions`, `RunConversationOptions`, `RunDelegatedLoopOptions`, `RunKnowledgeImprovementJobOptions`, `RunPersonaConfig`, `RunPersonaConversationOptions`, `RuntimeDecisionEvidenceRef`, `RuntimeDecisionPoint`, `RuntimeEventCollector`, `RuntimeEventOtelOptions`, `RuntimeHookContext`, `RuntimeHookErrorContext`, `RuntimeHookEvent`, `RuntimeRunCompleteInput`, `RuntimeRunCost`, `RuntimeRunHandle`, `RuntimeRunOptions`, `RuntimeRunPersistenceAdapter`, `RuntimeRunRow`, `RuntimeSession`, `RuntimeSessionStore`, `RuntimeStreamEventCollector`, `RuntimeStreamEventSummary`, `RuntimeTelemetryOptions`, `SanitizedKnowledgeReadinessReport`, `SanitizedKnowledgeRequirement`, `ServerSentEventOptions`, `SupervisedKnowledgeUpdateInput`, `SupervisedKnowledgeUpdateOptions`, `SupervisedKnowledgeUpdateResult`, `TrainingDatasetDocument`, `VetoedFact`, `WorktreeLoopRunnerOptions`, `AgenticGeneratorExecutorForWorktree`, `AgentRuntimeEvent`, `AgentRuntimeEventSink`, `AgentTaskStatus`, `AuthSource`, `ChatModelValidation`, `ControlDecision`, `ConversationStreamEvent`, `DeepReadonly`, `DelegatedLoopMode`, `DelegatedLoopRegistry`, `DelegatedLoopRunner`, `HaltPredicate`, `HaltReason`, `ImproveCodeOptions`, `ImprovementCandidate`, `ImprovementProfileCandidatePopulation`, `ImprovementProfilePopulationCandidate`, `ImprovementProfilePopulationLineage`, `ImproveMethodSource`, `ImproveOptimizationRunOptions`, `ImproveProfileSurface`, `ImproveResult`, `ImproveTrainingResult`, `KnowledgeReadinessCheck`, `KnowledgeReadinessCheckResult`, `ProfileImprovementHarnessRunOptions`, `ProfileImprovementHarnessTrainOptions`, `RuntimeDecisionKind`, `RuntimeHookTarget`, `RuntimeRunStatus`, `RuntimeStreamEvent`, `RuntimeStreamEventSink`, `SupervisedKnowledgeUpdater`, `TrainingBoundaryResult`, `TurnOrder`.

### Vertical agent — manifest + surface proposal source

Expand Down
32 changes: 29 additions & 3 deletions docs/canonical-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
Generated signatures and the complete export list live in docs/api/.
Run pnpm docs:freshness after editing this file. -->

> **Version 0.242.0.**
> **Version 0.243.0.**
> [`docs/api/primitive-catalog.md`](./api/primitive-catalog.md) lists every export and import path.
> `agent-eval` must satisfy `>=0.182.0 <0.183.0`.
> `sandbox` must satisfy `>=0.36.4 <0.42.0`.
Expand Down Expand Up @@ -292,9 +292,35 @@ artifact-addressed Router model identity. Runtime retains complete receipt ances
checkpoint bytes after serving and candidate validation, and durably publishes the receipt
before the runnable profile. Use the existing benchmark and held-out gates to assess that profile.

The controlled command trainer runs without a shell or inherited credentials and cancels its
POSIX process group. Managed training and serving adapters own their remote jobs and cleanup;
Pass a returned frozen profile directly into `improve` or a new bound harness; no mutable cast
or reconstruction is needed. When retraining a checkpoint-backed profile, references to that
same checkpoint in the small model, subagents, and modes follow the new receipt. Unrelated
model choices remain unchanged.

Training has no implicit deadline. Omit `timeoutMs` to rely on caller cancellation, or specify a
positive safe duration; the shared deadline timer supports long jobs without native timer overflow.
A managed adapter that ignores cancellation may continue working after Runtime stops awaiting it.

The controlled command trainer runs without a shell or inherited credentials. It snapshots the
request before asynchronous work, streams pinned input hashes (empty configuration files are
valid), and drains stdout/stderr within the caller's positive `maxOutputBytes` budget. Checkpoints
must still be nonempty and fit `maxCheckpointBytes`; dataset snapshot bounds remain in force.
It uses the same confirmed POSIX process-group teardown as the coding harnesses: permit graceful
shutdown, then escalate if necessary. This is trusted host execution, not an OS sandbox; separately
detached sessions are outside the owned process group.
Managed training and serving adapters own their remote jobs and cleanup;
inspect the failure stage and the training/serving uncertainty flags rather than assuming a
timeout removed external resources. A cancellation before adapter dispatch starts no job.
Once profile publication commits, later cancellation does not retract the committed result.
This is a local execution primitive, not a durable remote-job scheduler or a Router deployment API.


### Candidate validation has one acceptance rule

Across training, optimization, composed method leaves, and bound harnesses, `validateCandidate`
accepts by returning `undefined` synchronously and rejects by throwing. A promise or another
return value is an error, not an accepted candidate. Use a block body for side effects, rather
than returning the result of an assertion or array operation. Do asynchronous preparation before
calling the improvement API; keep measurement and held-out decisions in the existing evaluators.
The original callback remains part of execution identity, so centralizing its invocation does
not collapse distinct validator implementations into one cache identity.
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@tangle-network/agent-runtime",
"version": "0.242.0",
"version": "0.243.0",
"description": "Shared task-lifecycle skeleton for agents: a recursive loop kernel for chat turns, one-shot tasks, and multi-attempt loops, with trace capture and eval-gated self-improvement. Domain behavior lives in adapters; scoring and ship-gates in @tangle-network/agent-eval.",
"homepage": "https://github.com/tangle-network/agent-runtime#readme",
"repository": {
Expand Down
9 changes: 9 additions & 0 deletions scripts/verify-package-exports.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -307,6 +307,15 @@ try {
const harnessTraining: Promise<ImproveTrainingResult> = profileHarness.train(trainingOptions)
if (trainingResult.succeeded) {
const trainingReceipt: AgentTrainingReceipt = trainingResult.receipt
const { timeoutMs: _timeout, ...withoutDeadline } = trainingOptions
const retrained: Promise<ImproveTrainingResult> = improve(trainingResult.profile, withoutDeadline)
const rebound = createProfileImprovementHarness({
profile: trainingResult.profile,
executionRef: trainingOptions.executionRef,
agent: async () => 'fixture',
})
void retrained
void rebound
void trainingReceipt
}
void training
Expand Down
Loading
Loading