feat(planner): generate plans as validated task files instead of a reply - #2588
Conversation
The planning agent now writes one JSON task per file under tasks/ in a scratch workspace and runs `node validate-plan.mjs --tasks tasks` until it passes; ProPR re-assembles and re-validates the files with its own validator before accepting the plan. The plan no longer depends on the agent's final chat message, which is cut at the per-message output limit for very large plans, and no single tool call has to hold the whole plan. - validate-plan.mjs gains an additive --tasks mode that assembles task files in name order and names the file behind each error. - runPlanFileAgent gains a task-file contract, reads back only regular files (never a symlink out of the workspace), bounds what it reads, raises the task's turn limit (Claude ships CLAUDE_MAX_TURNS=10) and reports setup/agent failures as PlanFileAgentUnavailableError; usage limits keep their type. - PROPR_PLAN_GENERATION_MODE=file (default) | response. Only an unavailable agent falls back to the reply path; an invalid plan fails. - Heavy modules (agent registry, model aliases, log helpers) load on use, so planning modules and their tests stay free of connections. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pair tests Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
02ad806 to
6fd0b5d
Compare
|
/merge |
|
🔀 Resolved merge conflicts from Conflicting Files
Resolution SummaryResolved the conflict in
Refreshed the generated JavaScript containing the same conflict. The repository-wide scan found no remaining unresolved conflict markers; remaining matches are intentional strings or Markdown separators. Validation passed: TypeScript checking, JavaScript syntax checks—including the embedded validator—and all three focused test files using demo mode. Git could not mark the file resolved in its index because creating 🤖 Resolution Details
System-triggered merge conflict resolution |
Automatically resolved merge conflicts after merging target branch into PR branch. Resolved conflicts in: - packages/core/src/services/taskPlanning/planValidation.ts PR: #2588 Model: gpt-6-astra
|
/review astra |
|
✅ AI Code Review Complete requested by @integry Posted 1 review: |
🔍 AI Code Review — astraOverall EvaluationThe file-based approach addresses reply truncation, but the PR needs changes before merge: it can still accept partial plans, read outside the workspace, and incorrectly restart generation after an agent has produced output. ✅ Independent validation — ProPR validates copied task contents using its own validator, so editing the workspace validator cannot bypass schema checks. ✅ Granularity preserved — The new generation path applies This review is based on static tracing of the supplied code; no commands were run. Current checks show 18 passed, 5 pending, and none failed. Merge blockersEvery finding below was introduced by this PR and must be resolved before merging. F1: 🔴 File bounds silently discard tasks
F2: 🔴 Failed execution can save an unfinished plan
F3: 🔴 Symlinked task directory escapes the workspace
F4: 🔴 Exceptions after output incorrectly trigger fallback
SuggestionsThese are optional follow-ups and are not sent to S1: 🟢 Exercise real agent integrationsThe fake-agent tests verify orchestration, but do not establish that each supported CLI follows the file contract in its actual container. A small real generation per agent would improve rollout confidence, particularly because file mode becomes the default. This is optional integration evidence, separate from the demonstrated code defects above. ScoreThe approach and independent validation are sound foundations, but four correctness and boundary defects remain. Passing checks do not cover these failure sequences, and five checks are still pending. Score: 5/10 🤖 Review Details
🤖 Review by ProPR |
|
/fix |
|
✅ Applied the requested follow-up changes in commit 8c44459
Summary of ChangesImplemented F1–F4 only:
Added focused regressions and checked shared paths and callers. All 57 relevant tests passed, plus core typecheck, ESLint, and diff checks. No visual preview needed. 🤖 Implementation Details
Undo Changes • View Task Execution 💡 ProPR Slash Commands
🤖 Generated by ProPR |
Implemented F1–F4 only: - Reject excessive file counts and unacceptable task files. - Reject unsuccessful generation even when valid tasks exist. - Reject symlinked task directories and files. - Prevent fallback after output or an inconclusive output check; preserve usage-limit errors. Added focused regressions and checked shared paths and callers. All 57 relevant tests passed, plus core typecheck, ESLint, and diff checks. No visual preview needed. PR: #2588 Comment by: @integry (ID: 5879713126) Model: gpt-6-astra
|
/review astra |
|
✅ AI Code Review Complete requested by @integry Posted 1 review: |
🔍 AI Code Review — astraOverall EvaluationThe PR addresses reply truncation with incremental task files and independent validation. It needs changes before merge because synthetic routing retries can mix tasks from different generation attempts. ✅ Independent validation — ProPR validates copied task contents with its own validator, rejecting malformed tasks even when the workspace validator is modified. ✅ Controlled fallback — Execution failures after output exists do not trigger reply regeneration; thrown usage-limit errors retain their identity. This review is based on the supplied code; no commands were run. Current checks show 19 passed, four pending, and no failures. Merge blockersEvery finding below was introduced by this PR and must be resolved before merging. F5: 🔴 Isolate task files across routing retries
SuggestionsThese are optional follow-ups and are not sent to S2: 🟢 Exercise real agent adaptersA monitored smoke run across the supported agents would verify workspace access, tool permissions, validator execution, and turn-limit behavior. The fake-agent coverage verifies orchestration well, but cannot establish those adapter behaviors. This is optional integration confidence work, separate from the demonstrated retry defect. ScoreThe implementation has useful validation and failure-path coverage, but accepting stale tasks across routing retries is a plan-integrity blocker. The supplied checks show no failures, with four test shards still pending. Score: 6/10 🤖 Review Details
🤖 Review by ProPR |
|
/fix |
|
✅ Applied the requested follow-up changes in commit 9e07c8a
Summary of ChangesImplemented F5: each physical routing attempt gets a fresh workspace, and only the final attempt’s output is validated. Earlier output still prevents inappropriate response fallback. Added regressions for retry contamination, empty replacements, late writes, and preparation failures. All 75 relevant tests passed, along with core typecheck and ESLint. No visual preview needed. 🤖 Implementation Details
Undo Changes • View Task Execution 💡 ProPR Slash Commands
🤖 Generated by ProPR |
…workspace, and only the final attempt’s output is validated. Earlier output still prevents inappropriate response fallback. Implemented F5: each physical routing attempt gets a fresh workspace, and only the final attempt’s output is validated. Earlier output still prevents inappropriate response fallback. Added regressions for retry contamination, empty replacements, late writes, and preparation failures. All 75 relevant tests passed, along with core typecheck and ESLint. No visual preview needed. PR: #2588 Comment by: @integry (ID: 5879884452) Model: gpt-6-astra
|
/ultrafix |
|
🔄 Ultrafix loop started (goal: 8/10, max cycles: 10) First action:
|
|
✅ AI Code Review Complete requested by @integry Posted 1 review: |
🔍 AI Code Review — codex:gpt-6-astraOverall EvaluationThe PR addresses reply truncation with per-task files and independent validation. It needs a localized correction to failure propagation before merge. ✅ Independent validation — ProPR assembles copied task files using its own validator, so editing the workspace validator cannot authorize invalid output. ✅ Retry isolation — Each routed attempt receives a fresh workspace, with tests covering abandoned files, late writes, and cleanup. The supplied current-head checks report 24 passed and no failures or pending checks. This review used the supplied code only; no commands were run. Merge blockersEvery finding below was introduced by this PR and must be resolved before merging. F6: 🔴 Preserve explicit cancellation before fallback
SuggestionsThese are optional follow-ups and are not sent to S3: 🟢 Exercise real agent adaptersA monitored smoke generation with each supported adapter would supplement the fake-agent coverage by checking workspace access, validator execution, and completion reporting. This is an optional integration follow-up: the supplied review evidence does not establish an adapter-specific defect. S4: 🟢 Document practical plan limitsDocument the 200-file ceiling and allow turn-budget headroom for validation when describing maximum plan sizes. The current boundary test writes 200 files directly through a fake agent, so it does not demonstrate that an agent can write that many tasks and finish validation within 200 turns. This is optional operational clarification. ScoreThe implementation has substantial focused coverage and passing current-head checks. The cancellation-to-fallback conversion requires a small but important correction. Score: 6/10 🤖 Review Details
🤖 Review by ProPR |
|
✅ Applied the requested follow-up changes in commit 56f739c
Summary of ChangesImplemented F6 only. Explicit cancellation and other terminal execution errors retain their identity before output inspection, preventing response fallback. The shared repair path is covered too. Added 12 regressions. All 69 targeted tests passed, along with core typecheck, ESLint, and diff checks. No visual preview needed for this backend change. 🤖 Implementation Details
💡 ProPR Slash Commands
🤖 Generated by ProPR |
…al execution errors retain their identity before output inspection, preventing response fallback. The shared repair path is covered too. Implemented F6 only. Explicit cancellation and other terminal execution errors retain their identity before output inspection, preventing response fallback. The shared repair path is covered too. Added 12 regressions. All 69 targeted tests passed, along with core typecheck, ESLint, and diff checks. No visual preview needed for this backend change. PR: #2588 Comment by: @propr-ultrafix (ID: 0) Model: gpt-6-astra
|
/review astra |
|
✅ AI Code Review Complete requested by @integry Posted 1 review: |
🔍 AI Code Review — astraOverall EvaluationThe PR addresses reply truncation with incremental task files and independent validation. It needs a localized correction before merge: response fallback reuses an exhausted synthetic routing session. ✅ Independent validation — ProPR validates copied task contents with its own validator, rejecting incomplete objects and workspace-validator tampering. ✅ Attempt isolation — Each routed execution receives a fresh workspace, preventing abandoned task files from contaminating a successful retry. ✅ Failure coverage — Added tests cover invalid output, file bounds, symlinks, cancellation, and retry isolation. This review is based on static inspection; no commands were run. The supplied current checks show zero failures and eight pending checks. Merge blockersEvery finding below was introduced by this PR and must be resolved before merging. F7: 🔴 Reset routing for response fallback
SuggestionsThese are optional follow-ups and are not sent to S5: 🟢 Exercise real agent adaptersThe fake-agent tests establish orchestration behavior but do not verify that real adapters follow the file contract in the scratch workspace. A monitored smoke generation for each supported adapter would strengthen rollout confidence, particularly around workspace access and validator execution. This is optional integration validation, not a demonstrated adapter defect. ScoreThe implementation has strong validation and retry-isolation coverage, but the synthetic-routing fallback defect requires correction. Current checks have no reported failures, with eight still pending. Score: 6/10 🤖 Review Details
🤖 Review by ProPR |
|
/fix |
|
✅ Applied the requested follow-up changes in commit 5b61103
Summary of ChangesImplemented F7 only: response fallback now forks the routing session, restoring member eligibility after file execution exhausts the pool. Regressions cover thrown and returned failures, response retries, and explicit response mode. Related callers revealed no additional occurrences. Validation: 76 tests passed; core typecheck passed; lint reported only a test-file length warning. No visual preview needed. 🤖 Implementation Details
Undo Changes • View Task Execution 💡 ProPR Slash Commands
🤖 Generated by ProPR |
…g session, restoring member eligibility after file execution exhausts the pool. Implemented F7 only: response fallback now forks the routing session, restoring member eligibility after file execution exhausts the pool. Regressions cover thrown and returned failures, response retries, and explicit response mode. Related callers revealed no additional occurrences. Validation: 76 tests passed; core typecheck passed; lint reported only a test-file length warning. No visual preview needed. PR: #2588 Comment by: @integry (ID: 5880340697) Model: gpt-6-astra
|
/review astra |
|
✅ AI Code Review Complete requested by @integry Posted 1 review: |
🔍 AI Code Review — astraOverall EvaluationThe PR replaces reply-based plan delivery with individually written task files and host-side validation. The implementation is coherent and merge-ready within the reviewed scope, conditional on the remaining checks passing. ✅ Independent validation — ProPR assembles copied task files using its own validator, so modifying the workspace validator cannot authorize malformed tasks. ✅ Retry isolation — Each physical routing attempt receives a separate workspace, preventing abandoned tasks from contaminating a successful replacement plan. ✅ Controlled fallback — Output from failed attempts prevents response fallback, while explicit thrown cancellation errors retain their identity. This was a static review of the supplied code; no commands or tests were run. The authoritative status reports 20 passed checks, three pending, and no failures. Merge blockersNo merge blockers. SuggestionsThese are optional follow-ups and are not sent to S6: 🟢 Exercise real agent deliveryRun a representative file-mode generation through the supported physical adapters, including a plan requiring more than ten turns. The fake-agent tests cover orchestration extensively, but cannot establish actual CLI adherence to the file contract or scratch-workspace compatibility. This is useful integration follow-up rather than a demonstrated code blocker. S7: 🟢 Clarify repair documentationIn ScoreStrong validation and regression coverage support merging within scope. The remaining checks are pending, and physical-agent behavior has not been demonstrated by the supplied tests. Score: 8/10 🤖 Review Details
🤖 Review by ProPR |
Stacked on the
fix/plan-generation-integrityPR (PR A): merge that first. GitHub retargets this PR tomainonce A merges.Why
Plan 027d8f35's regeneration (Opus 5.5, 129,562 output tokens) was cut at the per-message output limit. Only the last 3,542 characters reached the parser, and a malformed "task" was saved. As long as the plan is taken from the final chat message, any plan larger than one message is at risk.
What
tasks/001.json,tasks/002.json, and so on, in plan order. It then runsnode validate-plan.mjs --tasks tasksand fixes the files it names until the validator exits 0. Each task is a separate tool call, so neither the reply nor any singleWritehas to hold the whole plan.planValidation.ts, still one source of truth): an additive--tasks <dir>mode assembles the files in name order intoplan.json. Errors name the file (task 2 (tasks/002.json) has no non-empty "body"). Without the flag, output is unchanged.planFileAgent.ts):taskFilesoption. ProPR reads the task files back and assembles and validates them with its own copy of the validator..env.exampleshipCLAUDE_MAX_TURNS=10, and a granular plan needs a turn per task.AgentTaskOptions.maxTurnsfield, which Claude uses; agents without a turn limit ignore it.planFileGeneration.ts): the existing planner prompt (fullContext) plus the file contract, which replaces the reply format. It hooks intocallLLMForPlanafter the token check, estimation and trace update. Granularity enforcement still applies.Mode, default and fallback
PROPR_PLAN_GENERATION_MODE:file(default) orresponse(the previous reply parsing).fileis the default:executeTask.node:22, so the validator runs everywhere.responsestays available as a switch.PlanFileAgentUnavailableErrorfalls back toresponsefor that run: the workspace couldn't be created, no agent could be resolved, orexecuteTaskthrew before producing anything.UsageLimitErrorpropagates unchanged, so requeueing still works.Tests
test/planFileGeneration.test.ts(12 tests) covers:callLLMForPlan, including granularity enforcementtest/planGenerationJsonRepair.test.tsis pinned toresponsemode, since it covers reply parsing.Not verified
docs/docs/features/planning.mdanddocs/docs/operations/configuration-reference.md(PROPR_PLAN_GENERATION_MODE,PROPR_PLAN_WORKSPACE_ROOT).🤖 Generated with Claude Code