Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 14 additions & 7 deletions docs/GLM53_B12X_KDA_ADAPTIVE_MTP_SPARKCACHE_TP4_QUICKSTART.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,9 +2,10 @@

Status: **implemented, not qualified**. This guide builds vLLM commit
`0b67266a0f37d6146a8403fb8482403c62f412d5` and the SparkCache overlay from
commit `5d571018de5b63a9a90e5c11e6d6e86bbff4a957`, Git tree
`e864ed9ad64f771188fdb59aa9738e348134d636`, for four DGX Spark systems at
TP4/DCP1.
commit `5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3`, Git tree
`94c236b9dfbf5f70075eb47877fd9caaa5d8c249`, for four DGX Spark systems at
TP4/DCP1. The adaptive-MTP composition has GPU-free contract coverage but no
four-rank persistent-restore or performance qualification.

The serving profile uses embedded MTP with maximum depth five, initial depth
three, and a 32-step acceptance window. Fastsafetensors uses queue size one.
Expand All @@ -26,7 +27,7 @@ storage. Clone both repositories beside each other:
git clone https://github.com/FujitsuPolycom/sparkring.git sparkring
git -C sparkring checkout --detach <revision-containing-this-guide>
git clone https://github.com/FujitsuPolycom/sparkcache.git sparkcache
git -C sparkcache checkout --detach 5d571018de5b63a9a90e5c11e6d6e86bbff4a957
git -C sparkcache checkout --detach 5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3

IMAGE='sparkring-glm53-runtime:b12x-kda-adaptive-mtp-0b67266a-arm64' \
BUILD_RECEIPT="$PWD/glm53-b12x-kda-adaptive-mtp-runtime-receipt.json" \
Expand All @@ -39,8 +40,8 @@ python sparkcache/deploy/glm53_flash/build_image.py \
--containerfile deploy/glm53_flash/Containerfile.b12x-kda-adaptive-mtp \
--base-image "${runtime_image}" \
--base-image-id "${runtime_image_id}" \
--source-sha256 f7c0565521fddeff7085e4cc08043cb8d1e2bde33abc67f83b8608a162d05b88 \
--sparkcache-revision 5d571018de5b63a9a90e5c11e6d6e86bbff4a957 \
--source-sha256 bc238f96e550c7ec27d4081dd1f2e741d404aaf5c8572d89ccc5e76812be4d63 \
--sparkcache-revision 5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3 \
--output-image sparkring-glm53-sparkcache:b12x-kda-adaptive-mtp-0b67266a-arm64
```

Expand All @@ -58,7 +59,9 @@ test "${#cuda_placement_sha256}" -eq 64
The runtime builder verifies the complete first-parent vLLM history from
`da4d7be` through adaptive MTP and the three live-tensor B12X KDA commits. The
SparkCache build verifies LF Linux preimages, four exact patches, and eleven
postimage source files.
postimage source files. The pinned SparkCache source accepts the canonical
`spark_cache_cuda_*` keys in the profile directly; no legacy-key translation
is part of this composition.

## Resolve the TP4 profile

Expand All @@ -85,6 +88,10 @@ loader queue, source identities, or attestation command.
The one-shot clear token is recorded only after a successful SparkCache-owned
cache removal. Restarting this unchanged profile does not clear again.

`--prefill-schedule-interval` is not part of this implemented profile. Test
interval `8` as a separate research-only profile so its mixed prefill/decode
tradeoff cannot be confused with adaptive-MTP or SparkCache results.

## Verify and launch

```bash
Expand Down
97 changes: 78 additions & 19 deletions docs/GLM53_DFLASH7_PYTHON_OVERLAY_SPARKCACHE_TP4_QUICKSTART.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,14 @@
# Serve GLM-5.3 with external DFlash7 and the exact Python-overlay runtime

Status: **implemented**, not qualified. The image builder, profile resolver,
and four-rank dry-run contract pass without GPUs. No image digest from this
path has completed TP4/DCP1 model loading, semantic generation, SparkCache
store/restart/restore, or concurrency qualification.
Status: **implemented**. The image builder, profile resolver, and four-rank
dry-run contract pass without GPUs. Local image ID
`sha256:eef863d8bc578815a80b0e2d9f0d745102b6363415225101fd92171a2e5a55cb`
is **qualified** only for the TP4/DCP1 startup, health, semantic generation,
arbitrary page-boundary replay, and 131,072- and 262,144-token restore cases
in the
[bounded validation record](../performance/records/glm53-flash/dflash7-python-overlay-pr30-live-validation.md).
The exact image has no retained C2/C8/C16 or DFlash response-quality evidence.
A rebuilt image has a different identity and requires its own live checks.

## Runtime contract

Expand Down Expand Up @@ -57,18 +62,21 @@ Two profiles share the same image and DFlash7 cache identity:
| Profile | Status | Loader behavior |
|---|---|---|
| `glm53-flash-dflash7-python-overlay-safetensors-sparkcache-tp4-dcp1.example.json` | **implemented**, not qualified | Uses global safetensors for target and draft. This follows the qualified-compatible loader shape but still requires live qualification on the composed 0b image. |
| `glm53-flash-dflash7-python-overlay-fastsafetensors-sparkcache-tp4-dcp1.example.json` | **implemented**, not qualified | Uses global fastsafetensors with queue size one for the target and `draft_load_config={"load_format":"safetensors"}` for DFlash. |
| `glm53-flash-dflash7-python-overlay-fastsafetensors-sparkcache-tp4-dcp1.example.json` | **qualified** only for the recorded image and bounded cases | Uses global fastsafetensors with queue size one for the target and `draft_load_config={"load_format":"safetensors"}` for DFlash. |

The image applies an exact-input vLLM patch that passes
`SpeculativeConfig.draft_load_config` to the DFlash model loader. The image
receipt verifies patch SHA-256
`39b567013ee7aed79f63200ed460129587933dc77fb430decdf19f78178de279` and
postimage SHA-256
`98acbae2b3bb4482d83f9637c163ce7c92707ccdf6561b7e431f23337f151cf4`.
Both profiles remain unqualified until live four-rank gates pass.
The all-safetensors profile remains unqualified. The fastsafetensors result
belongs only to the image ID and cases named above; it does not transfer to a
rebuild.

SparkCache PR #26 accepts the canonical CUDA keys used by both profiles. No
local PR25 compatibility profile or legacy-key rewrite is part of this path.
SparkCache commit `5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3`
accepts the canonical CUDA keys used by both profiles. No legacy-key rewrite
is part of this path.

## Resolve the profile and inspect the plan

Expand All @@ -79,7 +87,7 @@ device, host path, and image identity. Select one profile template:
```bash
receipt="$PWD/glm53-dflash7-python-overlay-image-receipt.json"
image='sparkring-glm53-sparkcache:dflash7-vllm-python-0b67266-native-da4d7be-b12x-b1d541f-arm64'
profile_template='scripts/config/glm53-flash-dflash7-python-overlay-safetensors-sparkcache-tp4-dcp1.example.json'
profile_template='scripts/config/glm53-flash-dflash7-python-overlay-fastsafetensors-sparkcache-tp4-dcp1.example.json'

python scripts/prepare_glm53_dflash7_python_overlay_profile.py \
--profile-template "$profile_template" \
Expand All @@ -101,22 +109,73 @@ python scripts/sparkring_generic_launcher.py \

`plan` is offline. Inspect every rank action before a lifecycle command.

## Start and observe the four-rank service

Run the strict placeholder check before starting containers:

```bash
python scripts/preflight.py \
--site /path/to/glm53-dflash7-site.yaml \
--strict-placeholders \
--json /path/to/glm53-dflash7-preflight.json
```

Starting the profile replaces the named four-rank deployment:

```bash
python scripts/sparkring_generic_launcher.py \
--site /path/to/glm53-dflash7-site.yaml \
--profile /path/to/glm53-dflash7-profile.json \
--execute \
--confirmation START_GLM53_FLASH_DFLASH7_PYTHON_OVERLAY_FASTSAFETENSORS_TP4 \
start
```

Tail rank zero from another terminal:

```bash
ssh operator@rank0.example.net \
'docker logs --follow --tail 120 glm53-flash-dflash7-python-overlay-fastsafetensors-sparkcache-tp4-r0 2>&1'
```

Wait for health and send a bounded generation smoke request with the model name from
the selected profile:

```bash
api_endpoint='http://rank0.example.net:8015'
served_model='glm-5.3-flash-nvfp4-dflash7-python-overlay-0b67266-on-da4d7be-b12x-b1d541f-tp4'
until curl --fail --silent "${api_endpoint}/health" >/dev/null; do sleep 5; done
curl --fail --silent --show-error "${api_endpoint}/v1/completions" \
-H 'Content-Type: application/json' \
-d "{\"model\":\"${served_model}\",\"prompt\":\"The capital of France is\",\"max_tokens\":16,\"temperature\":0}"
```

The configured `spark_cache_clear_once` value is a durable operation token,
not a permanent clear-on-start switch. SparkCache removes its owned cache
content, writes the token's completion marker only after successful removal,
and treats later starts with the same token as no-ops. Change the token only
when another intentional cache reset is required.

`--prefill-schedule-interval` is not part of the qualified DFlash7 profile.
Test interval `8` in a separate research-only profile so its mixed
prefill/decode tradeoff is measured independently.

## Cache namespace impact

The external DFlash weights SHA-256 is stored as
`spark_cache_draft_checkpoint_sha256`. It cannot share entries with embedded
MTP profiles. `tail-cow-v1` also separates these entries from snapshot-v1
manifests. The two target-loader profiles share a namespace because loader
choice does not change target or draft model state; each profile uses a
different cache root and one-shot clear token while qualification is pending.

[SparkCache pull request #30](https://github.com/FujitsuPolycom/sparkcache/pull/30)
combines canonical CUDA configuration names, replacement of a partial terminal
HMA page when an authenticated cache boundary falls inside that page, and an
eight-worker reader for authenticated page-delta chunks. The reader preserves
manifest descriptor order after concurrent reads.
Moving from SparkCache commit
`5d571018de5b63a9a90e5c11e6d6e86bbff4a957` to the pinned commit
choice does not change target or draft model state. Each template uses a
different cache root and one-shot clear token so loader observations remain
isolated.

The pinned SparkCache source combines canonical CUDA configuration names,
replacement of a partial terminal HMA page when an authenticated cache
boundary falls inside that page, and an eight-worker reader for authenticated
page-delta chunks. The reader preserves manifest descriptor order after
concurrent reads. Moving from SparkCache commit
`5d571018de5b63a9a90e5c11e6d6e86bbff4a957` to
`5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3` does not change the namespace.
Checkpoint identities, page-delta wire schemas, record vocabulary, digest
salts, parallel geometry, vLLM patches, the lease contract, and the CUDA
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
{
"schema": "sparkring-glm53-dflash7-pr30-live-validation/v1",
"status": {
"artifact_construction": "implemented",
"startup_health": "qualified",
"semantic_generation": "qualified",
"arbitrary_page_boundary_restore": "qualified",
"persistent_128k_restore": "qualified",
"persistent_256k_restore_correctness": "qualified",
"persistent_256k_restore_performance": "research-only",
"shared_prefix_concurrency": "unsupported",
"dflash_response_quality": "unsupported"
},
"artifact": {
"image": "sparkring-glm53-sparkcache:dflash7-vllm-python-0b67266-native-da4d7be-b12x-b1d541f-pr30-clean-arm64",
"image_id": "sha256:eef863d8bc578815a80b0e2d9f0d745102b6363415225101fd92171a2e5a55cb",
"published_digest": null,
"oci_revision_label": "b7a7265c62c1df05b17d1221d8fd0c97c54240d1",
"vllm_native_revision": "da4d7be6c97434f6942292ed8abbf4b32dc44355",
"vllm_python_revision": "0b67266a0f37d6146a8403fb8482403c62f412d5",
"vllm_python_tree": "ba9484ccb33aa56e90ff2f447f15ca9b9da97639",
"b12x_revision": "b1d541f9e71a35f030d45fae437630fff7507c2a",
"b12x_tree": "c69cdec1c59a08e8e0e549f930fa8abcfb5134ae",
"sparkcache_revision": "5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3",
"sparkcache_tree": "94c236b9dfbf5f70075eb47877fd9caaa5d8c249",
"sparkcache_source_sha256": "bc238f96e550c7ec27d4081dd1f2e741d404aaf5c8572d89ccc5e76812be4d63",
"cuda_placement_library_sha256": "d57509052b73853bcc8e3c3f47bb81748d87b9cbd8d908fc20d4c79a09aa400c",
"source_receipt_sha256": "497b068d4fb789499c540ea5ccc1638ff6d31a7a871b706f61f2a8b64606d88d",
"dflash_loader_patch_sha256": "39b567013ee7aed79f63200ed460129587933dc77fb430decdf19f78178de279",
"dflash_loader_postimage_sha256": "98acbae2b3bb4482d83f9637c163ce7c92707ccdf6561b7e431f23337f151cf4"
},
"conditions": {
"observed_at_utc": "2026-08-30T05:32:01Z/2026-08-30T05:40:04Z",
"hardware": "four NVIDIA DGX Spark systems",
"topology": {
"tensor_parallel_size": 4,
"decode_context_parallel_size": 1,
"pipeline_parallel_size": 1
},
"served_model": "glm-5.3-flash-nvfp4-dflash7-python-overlay-0b67266-on-da4d7be-b12x-b1d541f-tp4",
"target_loader": "fastsafetensors",
"target_loader_queue_size": 1,
"draft_loader": "safetensors",
"draft_method": "dflash",
"draft_tokens": 7,
"draft_tensor_parallel_size": 4,
"kv_cache_dtype": "fp8",
"kv_cache_bytes_per_rank": 21474836480,
"vllm_block_size_tokens": 256,
"max_num_seqs": 32,
"sparkcache_publication_schema": "tail-cow-v1",
"sparkcache_effective_publication_schema": "page-tail-cow-v1",
"sparkcache_config_schema": "canonical-v1",
"sparkcache_cuda_restore_io_workers": 8,
"sparkcache_load_threads": 2
},
"measurements": {
"startup": {
"target_fastsafetensors_seconds": 53.93,
"draft_safetensors_seconds": 4.85,
"total_model_load_seconds": 64.375343,
"ready_seconds_approx": 146.0,
"healthy_ranks": 4,
"restart_count_by_rank": [0, 0, 0, 0],
"oom_killed_by_rank": [false, false, false, false]
},
"clear_once": {
"worker_result": "already completed",
"scheduler_result": "already completed"
},
"arbitrary_page_boundary_restore": {
"base_tokens": 7168,
"restored_tokens": 12032,
"rank_0_end_to_end_milliseconds": 487.657,
"rank_0_read_milliseconds": 402.873,
"rank_0_h2d_submit_milliseconds": 52.12,
"rank_0_cuda_sync_milliseconds": 16.698,
"chunk_count": 47,
"connector_outcome": "verified"
},
"persistent_restore_128k": {
"restored_tokens": 131072,
"completion_token": 13,
"rank_restore_milliseconds": {
"0": 151.4,
"1": 123.6,
"2": 130.6,
"3": 138.8
}
},
"persistent_restore_256k": {
"restored_tokens": 262144,
"prime_request_seconds": 67.164,
"prime_completion_token": 916,
"replay_completion_token": 916,
"replay_wall_seconds": 7.835,
"rank_restore_milliseconds": {
"0": 5842.1,
"1": 7217.9,
"2": 6208.7,
"3": 7121.6
},
"rank_0_page_bytes": 1575821491,
"rank_0_chunk_count": 1024,
"rank_0_read_milliseconds": 4437.355,
"rank_0_h2d_submit_milliseconds": 961.72,
"rank_0_cuda_sync_milliseconds": 313.987
}
},
"limitations": [
"The exact image has no retained C2, C8, or C16 shared-prefix wave on this SparkCache revision.",
"The 262144-token result is one prime and one committed replay; it has no variability estimate.",
"The 1024-object page-delta layout makes the 262144-token restore a correctness result, not a performance qualification.",
"The first 262144-token replay began before all-rank manifest visibility and recomputed the unavailable tail; only the replay after committed visibility is reported as a restore.",
"No scored DFlash response-quality benchmark was run.",
"Sanitized raw HTTP response bodies and a complete command transcript were not retained with this record.",
"The image has no published OCI digest and exists only on the observed deployment hosts.",
"The prefill schedule interval was not changed from the profile default."
]
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# GLM-5.3 DFlash7 with SparkCache bounded validation

Status: **qualified** only for the startup, health, semantic generation, and
exact SparkCache restore cases described below. The 262,144-token restore
latency is **research-only**. Shared-prefix concurrency and DFlash response
quality are **unsupported** by this record.

## Conditions

- Four NVIDIA DGX Spark systems at TP4/DCP1/PP1.
- Local image ID
`sha256:eef863d8bc578815a80b0e2d9f0d745102b6363415225101fd92171a2e5a55cb`
on every rank. The image has no published OCI digest.
- SparkRing image revision
`b7a7265c62c1df05b17d1221d8fd0c97c54240d1`; retained vLLM native commit
`da4d7be6c97434f6942292ed8abbf4b32dc44355`; vLLM Python commit
`0b67266a0f37d6146a8403fb8482403c62f412d5`; B12X commit
`b1d541f9e71a35f030d45fae437630fff7507c2a`.
- SparkCache commit `5ec6a9953ad5d39120298bbfc26e95a6fa4b1dc3`, Git tree
`94c236b9dfbf5f70075eb47877fd9caaa5d8c249`, and deployable-source SHA-256
`bc238f96e550c7ec27d4081dd1f2e741d404aaf5c8572d89ccc5e76812be4d63`.
- GLM-5.3 NVFP4 target loaded with fastsafetensors queue size one. The external
BF16 DFlash checkpoint loaded with safetensors. Serving used seven draft
tokens, draft TP4, FP8 KV, 32 sequences, and 256-token vLLM blocks.
- SparkCache used `tail-cow-v1`, resolved to `page-tail-cow-v1` for opaque GLM
pages, with the canonical CUDA configuration keys, eight storage-read
workers, and two request-placement lanes.

The machine-readable artifact identities, observations, and limitations are
in
[`validation.json`](../../receipts/glm53-flash/dflash7-python-overlay-pr30/validation.json).

## Measurement

Startup values came from the target-loader, draft-loader, model-runner, and API
readiness logs. Container inspection checked the same complete image ID on all
four ranks, running state, restart count, and the runtime OOM flag.

Restore values came from each rank's SparkCache completion log. The client
observed the same completion token after the 131,072-token prime and replay,
and the same completion token after the 262,144-token prime and committed
replay. The first 262,144-token replay began before all-rank manifest
visibility and recomputed the unavailable tail; it is not counted as a cache
restore. Each reported shape has one retained observation and no variability
estimate.

## Result

| Gate or measurement | Observed result |
|---|---|
| Target fastsafetensors load | 53.93 seconds |
| Draft safetensors load | 4.85 seconds |
| Total model load | 64.375 seconds |
| API readiness | approximately 146 seconds after container start |
| Four-rank health | four running containers; restart counts `0,0,0,0`; OOM flags `false,false,false,false` |
| One-shot cache reset | scheduler and worker reported that the configured operation token had already completed; no repeated removal occurred |
| Arbitrary page boundary | 7,168-token base followed by verified 12,032-token restore; rank-zero total 487.657 ms; 47 objects |
| Persistent 131,072-token restore | prime and replay completion token 13; rank range 123.6-151.4 ms |
| Persistent 262,144-token restore | prime and replay completion token 916; 7.835-second request; rank range 5,842.1-7,217.9 ms |
| Rank-zero 262,144-token phases | 1,024 objects; 4,437.355 ms read; 961.720 ms CUDA submission; 313.987 ms CUDA synchronization |

## Conclusion

The exact four-rank image loaded the target and external draft, remained
healthy, continued generation, and restored the authenticated arbitrary-page,
131,072-token, and 262,144-token cases without changing their observed
completion tokens. Those bounded cases are qualified for correctness. The
large restore is too slow to support a performance claim.

## Limitations

- This SparkCache revision has no retained C2, C8, or C16 shared-prefix wave on
the exact image. Concurrency observations from another image do not qualify
this artifact.
- The 262,144-token result is one prime and one committed replay.
- The 1,024-object physical layout dominates large page-delta restore. A
macro-object layout requires separate implementation and qualification.
- No scored DFlash response-quality benchmark was run.
- Sanitized raw HTTP bodies and a complete command transcript were not
retained with this record.
- The image exists only on the observed hosts and cannot be pulled by digest.
- `--prefill-schedule-interval 8` was not part of the recorded profile and
requires an independent research-only mixed-traffic test.
Loading
Loading