Skip to content

speculative: --spec-draft-window and --spec-draft-n-max-tail (MTP drafting at every depth) - #320

Open
professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/spec-draft-window
Open

professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/spec-draft-window

Conversation

@professorpalmer

Copy link
Copy Markdown

One of the five pieces of #285, split per bri-prism's request on #317. Single commit on current prism (eaecb50c7), applies without conflicts.

What it does

--spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows before each batch, and the draft context is sized for the window (common_speculative_init), so its cells are reused and a draft pass costs the same at any depth. This retires the --spec-draft-depth-max 24576 cutoff from #221 (the flag stays, default 0). --spec-draft-n-max-tail K sets the draft size from the tiered-KV line (--kv-vram-cells) on; the MTP draft loop now honours the per-draft cap. Without the tiered-KV PR the tail flag is inert; the window works on its own.

Receipts (RTX 4070, Bonsai 2 27B PTQ1_0 with the MTP head, q8_0 K/V)

Serving recipe and receipts: https://github.com/professorpalmer/bonsai-ada-surgery/blob/main/docs/Q8_FULL_CONTEXT.md

…fting at every depth)

--spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows
before feeding each batch and the draft context is sized for the window, so its cells are reused, its
cache stays small and on the device, and a draft pass costs the same at any depth. The MTP head predicts
the next few tokens from recent context: a 16k window accepts as many drafts as the full history at 131k.
Also sizes the draft context for --spec-draft-depth-max when no window is set.

--spec-draft-n-max-tail N: draft size once the sequence reaches --kv-vram-cells. Past the tiered-KV line a
step is bound by reading the host tail over PCIe, and a wider verify reads it once for all columns. The
drafter and the output limits are built for max(n_max, n_max_tail); each slot caps a draft by depth, and the
MTP draft loop now honours that per-draft cap.

Bonsai 2 27B on an RTX 4070 with the MMA decode route: drafting pays at every depth, so the
--spec-draft-depth-max cutoff is no longer needed (32k 54.5 -> 103.6 tok/s, 64k 48.0 -> 90.1, same
acceptance; window alone +17-18% at 131k; tail draft 4 vs 2: +26% code / +15% prose at 180k).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit d2e2964)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant