Skip to content

Add exact GLM-5.3 DFlash7 Python-overlay runtime - #133

Closed
FujitsuPolycom wants to merge 1 commit into
codex/cuda-restore-terminologyfrom
codex/glm53-dflash7-python-overlay
Closed

Add exact GLM-5.3 DFlash7 Python-overlay runtime#133
FujitsuPolycom wants to merge 1 commit into
codex/cuda-restore-terminologyfrom
codex/glm53-dflash7-python-overlay

Conversation

@FujitsuPolycom

Copy link
Copy Markdown
Owner

Result

Adds an exact-contract GLM-5.3 DFlash7 image path and TP4/DCP1 runtime profiles on top of the SparkCache CUDA restore terminology branch.

The image retains vLLM da4d7be native extensions and wheel metadata, overlays the 31 Python files from 0b67266, installs B12X b1d541f, and composes SparkCache reconstructed-page placement at 5d571018. Prepared image labels and receipts identify external DFlash7; they do not claim adaptive MTP.

Profiles

  • Safetensors: implemented but unqualified on the composed 0b image.
  • Fastsafetensors queue one: research-only until external-draft loading and peak GPU memory pass live TP4 gates.

Both profiles use the external DFlash2 weights digest b33c0347, seven speculative tokens, draft TP4, FP8 target KV, 32 sequences, 256-token blocks, page-tail copy-on-write publication, and canonical SparkCache CUDA restore keys.

Compatibility

The DFlash checkpoint digest and page-tail publication schema select a cache namespace distinct from embedded MTP and snapshot-v1 entries. Loader choice does not alter model identity; the two unqualified profiles use separate cache roots and one-shot clear tokens. No wire values, digest salts, or logical geometry change.

Validation

  • 24 focused GPU-free tests passed.
  • Exact prepared context verified 31 vLLM overlay files and DFlash7-labelled image metadata.
  • Ruff, Python compilation, JSON parsing, Bash syntax, git diff, and changed-prose checks passed.
  • A broader pre-existing GLM README assertion remains failing on the stacked base and is unchanged by this branch.

No image was built or published, and no serving host was modified.

Construct a DFlash7-labelled image from the shared 31-file vLLM Python overlay while retaining da4d7be native extensions, B12X b1d541f, and SparkCache reconstructed-page placement at 5d571018. Prepared receipts name the external DFlash7 workload and preserve exact source, native, CUDA placement, patch, and lease-contract verification.

Add TP4/DCP1 profiles for global safetensors and fastsafetensors. Both use seven DFlash tokens, draft TP4, FP8 target KV, 32 sequences, 256-token blocks, page-tail copy-on-write publication, and canonical SparkCache CUDA restore keys. Safetensors is implemented but unqualified on the composed image; fastsafetensors is research-only pending live draft-loading and peak-memory gates.

Cache compatibility: the external DFlash weights digest and page-tail publication schema select a namespace distinct from embedded MTP and snapshot-v1 entries. The two loader profiles share model identity but use separate test roots and clear tokens. No wire value, digest salt, or logical geometry changes.

Validation: 24 focused GPU-free tests passed; an exact prepared context verified 31 vLLM files and DFlash7 image labels; Ruff, Python compilation, JSON parsing, Bash syntax, git diff, and prose checks passed. A broader unchanged GLM README assertion remains failing at the stacked base.
@FujitsuPolycom

Copy link
Copy Markdown
Owner Author

The resulting runtime, operator contract, and evidence are consolidated in retained draft stack #146#147#150. Independent page-base research remains in #149. Closing this superseded draft and deleting only its remote head branch.

@FujitsuPolycom
FujitsuPolycom deleted the codex/glm53-dflash7-python-overlay branch August 31, 2026 01:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant