Skip to content

Record adaptive-MTP and DFlash7 decode diagnostic - #132

Closed
FujitsuPolycom wants to merge 2 commits into
codex/cuda-restore-terminologyfrom
codex/glm53-adaptive-vs-dflash-diagnostic
Closed

Record adaptive-MTP and DFlash7 decode diagnostic#132
FujitsuPolycom wants to merge 2 commits into
codex/cuda-restore-terminologyfrom
codex/glm53-adaptive-vs-dflash-diagnostic

Conversation

@FujitsuPolycom

Copy link
Copy Markdown
Owner

Result

Status: research-only.

Add a bounded GLM-5.3 decode-log diagnostic for the adaptive embedded-MTP runtime and the retained DFlash7 runtime. The record preserves exact target, runtime, image, draft, topology, and UTC interval identities.

The adaptive interval observed depths 2–4, draft acceptance from 48.9% to 81.2%, and server generation gauges from 24.5 to 38.7 tok/s. The retained DFlash7 interval observed generation gauges from 40.6 to 136.0 tok/s. A separate operator-reported matched coding peak of approximately 70 versus 33.5 tok/s is labeled unaudited.

The prompt, generated output, request shape, and client receipt were not preserved. The record therefore makes no benchmark, quality, or isolated speculator claim. The two runtime compositions differ in multiple components.

Neither interval contains a SparkCache store, restore, placement, publication, or maintenance event. The evidence does not measure or implicate SparkCache work.

Validation

  • python -m pytest performance/test_glm53_decode_log_diagnostic.py scripts/test_sparkcache_terminology.py -q: 9 passed
  • ruff check --select E,F,W --ignore E501 performance scripts: passed
  • The fixture test verifies required evidence sections, exact interval endpoints, recomputed ranges, missing-receipt markers, and the local receipt link.

This draft is stacked on #131 because the adaptive runtime identity uses the CUDA-terminology profile from that branch.

Add a research-only GLM-5.3 server-log record comparing the observed adaptive-MTP and retained DFlash7 decode intervals. Preserve exact runtime, model, image, timestamp, depth, acceptance, and generation-gauge evidence while marking the absent prompt/output receipt and user-reported coding peak as unaudited. State that SparkCache emitted no work in either decode interval, so the observation does not measure cache behavior. Validation: 9 focused fixture, link, and terminology tests passed; Ruff E/F/W checks passed.
@FujitsuPolycom

Copy link
Copy Markdown
Owner Author

The resulting runtime, operator contract, and evidence are consolidated in retained draft stack #146#147#150. Independent page-base research remains in #149. Closing this superseded draft and deleting only its remote head branch.

@FujitsuPolycom
FujitsuPolycom deleted the codex/glm53-adaptive-vs-dflash-diagnostic branch August 31, 2026 01:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant