Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
86 changes: 86 additions & 0 deletions docs/scheduler-migration-handoff-protocol.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
# Scheduler migration handoff protocol requirements

## Decision

Workflow-to-CHASM migration must remain disabled until History can durably fence mutation
admission for the source workflow incarnation. Scheduler-only staging, retries, or channel drains
cannot establish a lossless ownership boundary.

This is a safety decision, not an implementation deferral. The deterministic evidence in
`scheduler-migration-defects.md` shows that History can accept a signal and deliver it to an SDK
channel while the migration local activity is pending. The workflow can then complete from the
older snapshot without consuming the signal. Any additional stage or finalize activity yields
again and recreates the same window.

## Required invariants

For one namespace, schedule ID, and source incarnation:

1. At most one implementation may execute schedule actions.
2. Every API operation acknowledged before cutover is represented exactly once at the active
destination or remains durably pending there.
Comment on lines +20 to +21

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve operation request IDs across cutover

If a V1 PatchSchedule signal is accepted just before sealing but its response is lost, the signal is included in the final snapshot while a client retry with the same request ID can reach CHASM after activation. History's signal-request-ID deduplication is not part of the listed migrated state, and the CHASM patch handler currently performs an ordinary component update without attaching the frontend request ID, so a trigger or backfill can be applied twice despite this exactly-once invariant. Transfer the accepted request-ID deduplication state or retain a routing-side dedupe record across the handoff.

Useful? React with 👍 / 👎.

3. Destination ownership is proven by an immutable migration ID, never inferred from a shared
schedule ID or an `AlreadyExists` response.
4. A snapshot revision can activate only after History has stopped accepting mutations at the
source and the snapshot covers History's final accepted-operation watermark.
5. Delete and reverse migration fence older migration IDs so an ambiguous retry cannot resurrect
an earlier schedule incarnation.
Comment on lines +26 to +27

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Persist a deletion fence even when no stage exists

A delete can reach CHASM before the concurrent stage write and get NotFound, then terminate V1 successfully, after which the in-flight stage commits with no workflow left to abort it. The current frontend deliberately invokes CHASM deletion before V1 termination, and the required durable state contains no tombstone or ordered incarnation generation that lets a later stage distinguish this delete from an earlier one; an opaque migration ID alone cannot be ordered. Require deletion to durably fence the source migration ID even when the destination is absent, otherwise an acknowledged delete can leave an orphan stage that blocks recreation of the schedule ID.

Useful? React with 👍 / 👎.


## Required durable state

History must replicate a source fence containing the workflow first-execution identity, migration
ID, admission phase, and final accepted-event watermark. CHASM must persist the same migration ID,
snapshot revision, covered watermark, and staged/active phase. Workflow state must preserve the
migration ID and last acknowledged revision across continue-as-new.
Comment on lines +31 to +34

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Prevent reset from resurrecting the V1 scheduler

A user can call ResetWorkflowExecution with an explicit scheduler run after migration completes, rebuilding a new run from a workflow task before the migration ID and frozen state were recorded. Preserving state only across continue-as-new does not cover this branch, and the described History fence blocks mutation admission rather than timer processing or the V1 StartWorkflow activity, so the reset V1 run can execute actions concurrently with the active CHASM scheduler. The protocol must reject resets for a fenced chain or force every reset branch to inherit the terminal migration fence.

Useful? React with 👍 / 👎.


The stage operation is idempotent for an equal ID and revision, replaces only an older revision
for the same inactive ID, and rejects a different ID. Staging starts no generator, invoker,
backfiller, or callback task. Finalize atomically activates only the matching staged revision.
Abort removes only a matching inactive stage and can never deactivate an active scheduler.
Comment on lines +36 to +39

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reject API mutations against an inactive stage

While a staged CHASM execution exists under the schedule's normal business ID, CHASM-first frontend routing can send PatchSchedule or UpdateSchedule to it before activation. The requirements only prohibit scheduler tasks, so an implementation may acknowledge and persist that mutation; step 4 can then replace the stage with a higher V1 snapshot revision and silently discard it. Require inactive stages to reject or redirect every external mutation, with only migration reconciliation, finalize, abort, and deletion allowed to access them.

Useful? React with 👍 / 👎.


## Cutover sequence

1. V1 records a stable migration ID and freezes action execution.
2. V1 stages snapshot revision 1. Ambiguous responses are reconciled by migration ID and revision.
3. History seals mutation admission for that source incarnation and returns the last accepted
event watermark. Later callers receive a retryable redirect and are never acknowledged by V1.
4. V1 drains through that watermark. If state changed, it stages a higher revision.
Comment on lines +45 to +47

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Make the accepted-event watermark observable to V1

The workflow cannot currently prove that it has drained through the event ID returned by History: SDK signal channels expose only the decoded signal payload, not the source history event ID, and the patch/update payloads contain no admission sequence. Consequently V1 can empty the channels it sees and label a revision with the sealed watermark without knowing whether the watermark event has actually been delivered, recreating the loss window this protocol is intended to close. Require an SDK-visible sequence on each mutation or a deterministic barrier event that is delivered only after every event through the watermark.

Useful? React with 👍 / 👎.

5. CHASM finalizes the matching revision and starts scheduler tasks in the same durable transition.
6. V1 records the finalize receipt and completes. A lost response reconciles the active migration
ID; it never resumes local action execution.

If rollout is disabled before activation, V1 confirms abort of its exact inactive stage before
resuming. If activation has committed, recovery must finish the handoff or perform an explicit
Comment on lines +52 to +53

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Treat sealing as the point of no return

If the rollout flag is disabled after step 3 seals History but before step 5 activates CHASM, this instruction aborts the only staged destination and resumes V1 even though the protocol explicitly makes source admission permanently sealed. The schedule then either rejects all subsequent mutations or requires reopening the fence, contradicting the failover invariant and reintroducing the stale-request race. Abort-and-resume must be limited to pre-seal failures; once sealing commits, recovery must finish activation regardless of the rollout flag.

Useful? React with 👍 / 👎.

reverse migration.

## Compatibility and deployment

Protocol support must reach every possible active and failover History host before a new scheduler
workflow version can emit fence commands. Older create requests without migration fields retain
legacy behavior for existing histories, but the rollout flag must not select them for safe
cutover. An older host that cannot enforce the ingress fence is a capability-gating failure.
Comment on lines +58 to +61

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Quiesce legacy migration RPCs before deploying staged creation

If migration was previously enabled, an old V1 local activity can already have passed its rollout guard and still be issuing a legacy create while CHASM phase support is deployed first. Because this compatibility rule keeps requests without migration fields on legacy behavior, that in-flight request can immediately activate CHASM before the History fence exists, preserving the acknowledged-signal loss race despite disabling new selections. The deployment protocol must disable migration, wait out or reconcile all pending legacy attempts, and only then install handling that can accept legacy migration creates.

Useful? React with 👍 / 👎.


The workflow change requires a new recorded scheduler version and replay fixture. CHASM protocol
support deploys first; History fence support and frontend redirect handling deploy next; only then
may the workflow version and migration rollout advance.

## Failure and load assessment

- A crash after staging replays the same ID and revision.
- A crash after sealing cannot reopen source admission; failover must replicate the fence.
- A lost finalize response reconciles the active ID and keeps V1 frozen.
- A namespace failover with an unconfirmed fence retries instead of activating or resuming.
- At 10x migration load, staging adds migration-only writes proportional to snapshot revisions;
ordinary scheduler execution gains no reads or writes.
- Dirty-snapshot rounds are bounded because sealing stops new V1 acceptance. Without sealing,
an input rate at or above forwarding capacity creates an unbounded outbox.

## Rejected partial fixes

- Treating every `AlreadyExists` as success does not prove ownership or snapshot freshness.
- Reusing a request ID prevents some duplicate creates but does not preserve accepted operations.
- Draining SDK channels before or after an activity moves the race to the next yield.
- A timeout or fixed grace period cannot prove that no stale frontend request remains in flight.
- Activating CHASM before source retirement permits duplicate action execution after response loss.
- Retaining a forwarder without an ingress fence has no finite safe retirement point and is
unbounded under sustained traffic.
Loading