You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
One shared system, two real consumers: a Discovery pursuit and a Platform-owned product agent can do meaningful work, accept correction, acquire useful evidence, retain capabilities, and improve subsequent work. Both use the same maintained execution, evaluation and activation mechanisms.
The ambition includes better tools, skills, code, information acquisition, collaboration and learning procedures—not only better prompts. Agents choose methods and organization within granted authority. GEPA/Omni, DSPy/RLM and native harnesses supply their existing algorithms. Weight training is optional.
Finish line: reproduce a real artifact → independently checked candidate → exact next-worker adoption → restore walkthrough in both consumers. Demonstrate a meaningful non-prompt change. Report learning value separately against strong unchanged and resource-matched alternatives; a negative or inconclusive experiment is not a fabricated win.
Priority / dependencies
Start A, B and D at their existing owners. C connects them to users. E is one bounded acquisition capability, not a training detour. F measures methods; G qualifies the actual released path. Do not wait for every research hypothesis to settle before shipping useful execution.
One active delivery PR per affected repository. Extend compatible open work; do not overwrite another contributor or create placeholder PRs. Core mechanisms stay in core packages; labs supply tasks, policy and evidence.
A. Control and recover real running agents
Owner: Runtime + Sandbox/provider. Reuse #1165 and merged #1291; no new control bus.
Extend existing durable steer/answer/cancel records and acknowledgers to the root and native-harness managers. Preserve request identity across reconnect; distinguish recorded, delivered, consumed where observable, and settled. State whether intervention is in-turn or between turns.
Recover original execution, completed child results, pending requests and partial artifacts after interruption. Never redispatch an uncertain external effect. Expose queued settlements and capacity contention so agents can change allocation; no research-specific scheduler.
Exercise client disconnect, coordinator restart, late child result, duplicate control delivery and cancellation through actual HTTP/provider boundaries. Done: a fresh authenticated client controls the live root and retrieves original work after interruption; no duplicated effects or invented cleanup confirmation.
B. Preserve the exact executable agent and retained state
Owner: Runtime candidate/workspace APIs, profile materializer, Sandbox/provider and Knowledge. Reuse harvestSurfaceDiffs, existing bundles, file/session APIs and stores.
Capture supported tool definitions and handlers, skill/helper bytes, knowledge references and native state before disposal; a digest-only observation is not retained content. Carry them through evaluation and next execution.
Isolate tenants and independent attempts. Reset removes learned state; carry imports only declared state. Distinguish same-session continuation, fresh-session import, resume and fork. Refuse unsupported capabilities before spend.
Verify effective materialization, not merely profile JSON. Reject mismatched bytes, missing handlers and authority expansion; candidate code cannot rewrite independent acceptance checks. Done: the next real worker executes an assessed tool/helper change. Reset/carry, corrupted-state and cross-tenant controls pass.
C. Connect improvement to the actual product
Owner: Intelligence/Platform in agent-dev-container, plus agent-app, using Runtime Intelligence.
Dispatch the existing runAndStorePlatformAgentProfileImprovement from an authorized durable product job using the actual executor, task source and evaluator. Recheck live callers first; the audit found test callers, not proof of a deployed producer. Reuse the existing run/proposal/review records.
Use the same source revision/materializer for product requests and evaluation. Activate the exact assessed revision, handle concurrent source edits, and restore through Platform's existing idempotent activation store.
Remove Agent App's duplicate certified-delivery cache in favor of Runtime's source, preserving forced refresh/coalescing/current inspection. Prompt projection must not masquerade as tool/file materialization. Prerequisite: feat(intelligence): share forced certified refresh with product consumers #1307; downstream adoption needs a compatible published release and verified lockfile.
Expose request/status/correction, revision diff, evidence, cost completeness, activate and restore in existing product surfaces. Keep authorization server-owned; no second chat shell or activation database. Done: a customer initiates, reviews and activates an improvement; the next task uses it; restore returns to the prior behavior.
D. Preserve resource truth from admission to result
Remove missing-cost-to-zero conversions, including runFixEngine's costUsd ?? 0. Preserve observed, estimated, unknown and explicitly non-token-priced usage. Include losing/failed attempts, retrieval, analysts, optimizer and final assessment without double-counting.
Enforce limits through existing pre-dispatch reservations/grants. A post-iteration counter is not a spending ceiling. Preserve settlement through cancellation and response loss.
Give compared arms explicit serving-enforced allowances; report actual spend and unused allowance. Fixed N is not equal dollars. Name infrastructure and subscription costs separately. Done: missing receipts cannot produce a complete total; joined receipts reconcile for failures, cancellation and repeated cumulative events.
E. Acquire and select better evidence before or during work
Connect an identified information gap to permitted retrieval, web/source inspection, a delegated agent, RLM exploration or a targeted experiment. Selection is agent/policy-owned—not mandatory preflight or an always-on classifier.
Record source/version/time/content references, access/usage rights, selection rationale, cost and what the worker consumed in existing evidence records. Treat acquired content as untrusted; preserve tenant boundaries and keep final-assessment answers out of learning.
Search acquisition/selection programs through existing GEPA/Omni/Knowledge surfaces. Score downstream accepted work, not document count or self-rated data quality; compare to existing retrieval and appropriate resource-matched alternatives. Done: a real missing/stale-evidence task triggers acquisition and the next action consumes it; an acquisition-disabled control isolates the effect. Acquisition is optional for unrelated work.
Prior art:WeCo AutoData, 17 Sep 2026, code: agents search executable document selectors using proxy-model training feedback. This is pre-training corpus selection, not a ready-made live search agent. Reuse the idea of optimizing selection procedures, not a new training stack. Code is MIT; the feature bank is CC BY-NC 4.0; corpus terms remain separate.
F. Improve methods and working checks, not just wording
Exercise existing full-profile/program, scoped/sequential/joint methods through the same real executor. Use one pinned GEPA/Omni route and one appropriate real RLM integration. Native harnesses own inner loops; no local fallback optimizer or rewrite into DSPy solely to access an algorithm.
Supply actionable execution feedback. Assess proposed working checks against independent known-outcome controls with Eval's existing evaluator audit; self-authored checks do not certify their own claims.
Compare reset, resource-matched retry/search, carry and revise on predeclared fresh source units. Isolate effective state, preserve errors/missingness and hold outer final evidence apart from adaptation. Measure specialist gains separately from transfer of a learning procedure. Done: executable candidates and a reproducible positive/negative/inconclusive result, with scope and costs. Extend the existing Lab instrument; do not copy it into Discovery.
G. Release and qualify the same path in both consumers
Owner: release owners, Discovery Lab, and the Platform/Agent App product.
Adopt an exact compatible published package set, rebundle frozen Discovery methods and verify served revisions. Peer-range edits alone are not compatibility evidence; historical registrations stay immutable.
Run the opening deliverable during real Discovery and product work. Retain effective state, control acknowledgements, artifacts, usage, independent assessment and restore evidence. Document one clean-checkout entrypoint per consumer.
Keep research goals/methods agent-authored. Discovery owns doctrine, Lab owns runs/evidence, lower packages own reusable mechanisms. No Discovery-owned execution framework, fixed swarm topology or additional state schema. Done: another engineer reproduces both journeys without machine-local patches. Fix observed defects upstream with regressions; synthetic fixtures prove contracts, not native-model capability or learning gains.
Code discipline
Small cohesive owner modules, no omnibus controller/God object. A new abstraction must remove duplication or solve an observed consumer gap. Preserve lower-level composition and remove replaced paths in the same delivery. Failure-first tests, cancellation/concurrency/security controls and packed cross-package checks belong in each PR. No new parallel ledger/cache/bus, algorithm implementation, protection bypass or unrelated version sweep.
feat(intelligence): share forced certified refresh with product consumers #1307 — open, non-draft: small Runtime prerequisite for C; forced shared refresh, concurrent-read correctness and construction-bound options. Clean exact-base verification: 4,262 tests passed, six existing skips; source/example types, lint, packed consumers and docs pass. Normal PR CI/review tracked on that PR. This is not closure of C.
Agent App source migration is not yet delivered: its Runtime 0.231.x peer window excludes the new API. Adopt the real release and remove the duplicate cache together; no hidden shim, broken install or empty PR.
Supervisor Lab #112/#113 retain experiment implementation and live acceptance. This issue coordinates shared delivery, not a replacement history.
Unchecked tasks remain unfinished. PR merge, package publication, deployed capability and measured improvement are separate states.
Deliverable
One shared system, two real consumers: a Discovery pursuit and a Platform-owned product agent can do meaningful work, accept correction, acquire useful evidence, retain capabilities, and improve subsequent work. Both use the same maintained execution, evaluation and activation mechanisms.
The ambition includes better tools, skills, code, information acquisition, collaboration and learning procedures—not only better prompts. Agents choose methods and organization within granted authority. GEPA/Omni, DSPy/RLM and native harnesses supply their existing algorithms. Weight training is optional.
Finish line: reproduce a real artifact → independently checked candidate → exact next-worker adoption → restore walkthrough in both consumers. Demonstrate a meaningful non-prompt change. Report learning value separately against strong unchanged and resource-matched alternatives; a negative or inconclusive experiment is not a fabricated win.
Priority / dependencies
Start A, B and D at their existing owners. C connects them to users. E is one bounded acquisition capability, not a training detour. F measures methods; G qualifies the actual released path. Do not wait for every research hypothesis to settle before shipping useful execution.
One active delivery PR per affected repository. Extend compatible open work; do not overwrite another contributor or create placeholder PRs. Core mechanisms stay in core packages; labs supply tasks, policy and evidence.
A. Control and recover real running agents
Owner: Runtime + Sandbox/provider. Reuse #1165 and merged #1291; no new control bus.
Done: a fresh authenticated client controls the live root and retrieves original work after interruption; no duplicated effects or invented cleanup confirmation.
B. Preserve the exact executable agent and retained state
Owner: Runtime candidate/workspace APIs, profile materializer, Sandbox/provider and Knowledge. Reuse
harvestSurfaceDiffs, existing bundles, file/session APIs and stores.Done: the next real worker executes an assessed tool/helper change. Reset/carry, corrupted-state and cross-tenant controls pass.
C. Connect improvement to the actual product
Owner: Intelligence/Platform in
agent-dev-container, plusagent-app, using Runtime Intelligence.runAndStorePlatformAgentProfileImprovementfrom an authorized durable product job using the actual executor, task source and evaluator. Recheck live callers first; the audit found test callers, not proof of a deployed producer. Reuse the existing run/proposal/review records.Done: a customer initiates, reviews and activates an improvement; the next task uses it; restore returns to the prior behavior.
D. Preserve resource truth from admission to result
Owner: Runtime #1175/#1252, existing provider grants, Intelligence
auto-improveadapters.runFixEngine'scostUsd ?? 0. Preserve observed, estimated, unknown and explicitly non-token-priced usage. Include losing/failed attempts, retrieval, analysts, optimizer and final assessment without double-counting.Done: missing receipts cannot produce a complete total; joined receipts reconcile for failures, cancellation and repeated cumulative events.
E. Acquire and select better evidence before or during work
Owner: Knowledge hooks + Runtime dispatch; Eval measures utility. Reuse
runRagKnowledgeImprovementLoop.acquireKnowledge, research/readiness hooks, delegate/spawn, connectors and native RLMs.Done: a real missing/stale-evidence task triggers acquisition and the next action consumes it; an acquisition-disabled control isolates the effect. Acquisition is optional for unrelated work.
Prior art: WeCo AutoData, 17 Sep 2026, code: agents search executable document selectors using proxy-model training feedback. This is pre-training corpus selection, not a ready-made live search agent. Reuse the idea of optimizing selection procedures, not a new training stack. Code is MIT; the feature bank is CC BY-NC 4.0; corpus terms remain separate.
F. Improve methods and working checks, not just wording
Owner: Runtime/Eval #1297/#1299; Supervisor Lab #112/#113; Intelligence analyst quality.
Done: executable candidates and a reproducible positive/negative/inconclusive result, with scope and costs. Extend the existing Lab instrument; do not copy it into Discovery.
G. Release and qualify the same path in both consumers
Owner: release owners, Discovery Lab, and the Platform/Agent App product.
Done: another engineer reproduces both journeys without machine-local patches. Fix observed defects upstream with regressions; synthetic fixtures prove contracts, not native-model capability or learning gains.
Code discipline
Small cohesive owner modules, no omnibus controller/God object. A new abstraction must remove duplication or solve an observed consumer gap. Preserve lower-level composition and remove replaced paths in the same delivery. Failure-first tests, cancellation/concurrency/security controls and packed cross-package checks belong in each PR. No new parallel ledger/cache/bus, algorithm implementation, protection bypass or unrelated version sweep.
Delivery record
Unchecked tasks remain unfinished. PR merge, package publication, deployed capability and measured improvement are separate states.