Skip to content

Ship adaptive agents end to end: durable execution, evidence acquisition, measured improvement, product adoption #1306

Description

@drewstone

Deliverable

One shared system, two real consumers: a Discovery pursuit and a Platform-owned product agent can do meaningful work, accept correction, acquire useful evidence, retain capabilities, and improve subsequent work. Both use the same maintained execution, evaluation and activation mechanisms.

The ambition includes better tools, skills, code, information acquisition, collaboration and learning procedures—not only better prompts. Agents choose methods and organization within granted authority. GEPA/Omni, DSPy/RLM and native harnesses supply their existing algorithms. Weight training is optional.

Finish line: reproduce a real artifact → independently checked candidate → exact next-worker adoption → restore walkthrough in both consumers. Demonstrate a meaningful non-prompt change. Report learning value separately against strong unchanged and resource-matched alternatives; a negative or inconclusive experiment is not a fabricated win.

Priority / dependencies

Start A, B and D at their existing owners. C connects them to users. E is one bounded acquisition capability, not a training detour. F measures methods; G qualifies the actual released path. Do not wait for every research hypothesis to settle before shipping useful execution.

One active delivery PR per affected repository. Extend compatible open work; do not overwrite another contributor or create placeholder PRs. Core mechanisms stay in core packages; labs supply tasks, policy and evidence.

A. Control and recover real running agents

Owner: Runtime + Sandbox/provider. Reuse #1165 and merged #1291; no new control bus.

  • Extend existing durable steer/answer/cancel records and acknowledgers to the root and native-harness managers. Preserve request identity across reconnect; distinguish recorded, delivered, consumed where observable, and settled. State whether intervention is in-turn or between turns.
  • Recover original execution, completed child results, pending requests and partial artifacts after interruption. Never redispatch an uncertain external effect. Expose queued settlements and capacity contention so agents can change allocation; no research-specific scheduler.
  • Exercise client disconnect, coordinator restart, late child result, duplicate control delivery and cancellation through actual HTTP/provider boundaries.
    Done: a fresh authenticated client controls the live root and retrieves original work after interruption; no duplicated effects or invented cleanup confirmation.

B. Preserve the exact executable agent and retained state

Owner: Runtime candidate/workspace APIs, profile materializer, Sandbox/provider and Knowledge. Reuse harvestSurfaceDiffs, existing bundles, file/session APIs and stores.

  • Capture supported tool definitions and handlers, skill/helper bytes, knowledge references and native state before disposal; a digest-only observation is not retained content. Carry them through evaluation and next execution.
  • Isolate tenants and independent attempts. Reset removes learned state; carry imports only declared state. Distinguish same-session continuation, fresh-session import, resume and fork. Refuse unsupported capabilities before spend.
  • Verify effective materialization, not merely profile JSON. Reject mismatched bytes, missing handlers and authority expansion; candidate code cannot rewrite independent acceptance checks.
    Done: the next real worker executes an assessed tool/helper change. Reset/carry, corrupted-state and cross-tenant controls pass.

C. Connect improvement to the actual product

Owner: Intelligence/Platform in agent-dev-container, plus agent-app, using Runtime Intelligence.

  • Dispatch the existing runAndStorePlatformAgentProfileImprovement from an authorized durable product job using the actual executor, task source and evaluator. Recheck live callers first; the audit found test callers, not proof of a deployed producer. Reuse the existing run/proposal/review records.
  • Use the same source revision/materializer for product requests and evaluation. Activate the exact assessed revision, handle concurrent source edits, and restore through Platform's existing idempotent activation store.
  • Remove Agent App's duplicate certified-delivery cache in favor of Runtime's source, preserving forced refresh/coalescing/current inspection. Prompt projection must not masquerade as tool/file materialization. Prerequisite: feat(intelligence): share forced certified refresh with product consumers #1307; downstream adoption needs a compatible published release and verified lockfile.
  • Expose request/status/correction, revision diff, evidence, cost completeness, activate and restore in existing product surfaces. Keep authorization server-owned; no second chat shell or activation database.
    Done: a customer initiates, reviews and activates an improvement; the next task uses it; restore returns to the prior behavior.

D. Preserve resource truth from admission to result

Owner: Runtime #1175/#1252, existing provider grants, Intelligence auto-improve adapters.

  • Remove missing-cost-to-zero conversions, including runFixEngine's costUsd ?? 0. Preserve observed, estimated, unknown and explicitly non-token-priced usage. Include losing/failed attempts, retrieval, analysts, optimizer and final assessment without double-counting.
  • Enforce limits through existing pre-dispatch reservations/grants. A post-iteration counter is not a spending ceiling. Preserve settlement through cancellation and response loss.
  • Give compared arms explicit serving-enforced allowances; report actual spend and unused allowance. Fixed N is not equal dollars. Name infrastructure and subscription costs separately.
    Done: missing receipts cannot produce a complete total; joined receipts reconcile for failures, cancellation and repeated cumulative events.

E. Acquire and select better evidence before or during work

Owner: Knowledge hooks + Runtime dispatch; Eval measures utility. Reuse runRagKnowledgeImprovementLoop.acquireKnowledge, research/readiness hooks, delegate/spawn, connectors and native RLMs.

  • Connect an identified information gap to permitted retrieval, web/source inspection, a delegated agent, RLM exploration or a targeted experiment. Selection is agent/policy-owned—not mandatory preflight or an always-on classifier.
  • Record source/version/time/content references, access/usage rights, selection rationale, cost and what the worker consumed in existing evidence records. Treat acquired content as untrusted; preserve tenant boundaries and keep final-assessment answers out of learning.
  • Search acquisition/selection programs through existing GEPA/Omni/Knowledge surfaces. Score downstream accepted work, not document count or self-rated data quality; compare to existing retrieval and appropriate resource-matched alternatives.
    Done: a real missing/stale-evidence task triggers acquisition and the next action consumes it; an acquisition-disabled control isolates the effect. Acquisition is optional for unrelated work.

Prior art: WeCo AutoData, 17 Sep 2026, code: agents search executable document selectors using proxy-model training feedback. This is pre-training corpus selection, not a ready-made live search agent. Reuse the idea of optimizing selection procedures, not a new training stack. Code is MIT; the feature bank is CC BY-NC 4.0; corpus terms remain separate.

F. Improve methods and working checks, not just wording

Owner: Runtime/Eval #1297/#1299; Supervisor Lab #112/#113; Intelligence analyst quality.

  • Exercise existing full-profile/program, scoped/sequential/joint methods through the same real executor. Use one pinned GEPA/Omni route and one appropriate real RLM integration. Native harnesses own inner loops; no local fallback optimizer or rewrite into DSPy solely to access an algorithm.
  • Supply actionable execution feedback. Assess proposed working checks against independent known-outcome controls with Eval's existing evaluator audit; self-authored checks do not certify their own claims.
  • Compare reset, resource-matched retry/search, carry and revise on predeclared fresh source units. Isolate effective state, preserve errors/missingness and hold outer final evidence apart from adaptation. Measure specialist gains separately from transfer of a learning procedure.
    Done: executable candidates and a reproducible positive/negative/inconclusive result, with scope and costs. Extend the existing Lab instrument; do not copy it into Discovery.

G. Release and qualify the same path in both consumers

Owner: release owners, Discovery Lab, and the Platform/Agent App product.

  • Adopt an exact compatible published package set, rebundle frozen Discovery methods and verify served revisions. Peer-range edits alone are not compatibility evidence; historical registrations stay immutable.
  • Run the opening deliverable during real Discovery and product work. Retain effective state, control acknowledgements, artifacts, usage, independent assessment and restore evidence. Document one clean-checkout entrypoint per consumer.
  • Keep research goals/methods agent-authored. Discovery owns doctrine, Lab owns runs/evidence, lower packages own reusable mechanisms. No Discovery-owned execution framework, fixed swarm topology or additional state schema.
    Done: another engineer reproduces both journeys without machine-local patches. Fix observed defects upstream with regressions; synthetic fixtures prove contracts, not native-model capability or learning gains.

Code discipline

Small cohesive owner modules, no omnibus controller/God object. A new abstraction must remove duplication or solve an observed consumer gap. Preserve lower-level composition and remove replaced paths in the same delivery. Failure-first tests, cancellation/concurrency/security controls and packed cross-package checks belong in each PR. No new parallel ledger/cache/bus, algorithm implementation, protection bypass or unrelated version sweep.

Delivery record

Unchecked tasks remain unfinished. PR merge, package publication, deployed capability and measured improvement are separate states.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions