Repository navigation
speculative: --spec-draft-window and --spec-draft-n-max-tail (MTP drafting at every depth) - #320
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
…fting at every depth) --spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows before feeding each batch and the draft context is sized for the window, so its cells are reused, its cache stays small and on the device, and a draft pass costs the same at any depth. The MTP head predicts the next few tokens from recent context: a 16k window accepts as many drafts as the full history at 131k. Also sizes the draft context for --spec-draft-depth-max when no window is set. --spec-draft-n-max-tail N: draft size once the sequence reaches --kv-vram-cells. Past the tiered-KV line a step is bound by reading the host tail over PCIe, and a wider verify reads it once for all columns. The drafter and the output limits are built for max(n_max, n_max_tail); each slot caps a draft by depth, and the MTP draft loop now honours that per-draft cap. Bonsai 2 27B on an RTX 4070 with the MMA decode route: drafting pays at every depth, so the --spec-draft-depth-max cutoff is no longer needed (32k 54.5 -> 103.6 tok/s, 64k 48.0 -> 90.1, same acceptance; window alone +17-18% at 131k; tail draft 4 vs 2: +26% code / +15% prose at 180k). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit d2e2964)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One of the five pieces of #285, split per bri-prism's request on #317. Single commit on current
prism(eaecb50c7), applies without conflicts.What it does
--spec-draft-window N: the draft (MTP) context keeps only the last N rows. The server drops older rows before each batch, and the draft context is sized for the window (common_speculative_init), so its cells are reused and a draft pass costs the same at any depth. This retires the--spec-draft-depth-max 24576cutoff from #221 (the flag stays, default 0).--spec-draft-n-max-tail Ksets the draft size from the tiered-KV line (--kv-vram-cells) on; the MTP draft loop now honours the per-draft cap. Without the tiered-KV PR the tail flag is inert; the window works on its own.Receipts (RTX 4070, Bonsai 2 27B PTQ1_0 with the MTP head, q8_0 K/V)
ggml_cuda_fattn_mma_kv_native_supportedguard); the window alone removes the depth cutoff, the speed at depth depends on the attention kernel the decode step uses.Serving recipe and receipts: https://github.com/professorpalmer/bonsai-ada-surgery/blob/main/docs/Q8_FULL_CONTEXT.md