Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,7 +81,8 @@ measurements, and known limits out of the generic cache design.

| Model family | Guide |
|---|---|
| GLM-5.3 Flash | [Four-node SparkRing quickstart](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_JJ_R8_GB10_SPARKCACHE_TP4_QUICKSTART.md) · [SparkCache integration notes](deploy/glm53_flash/README.md) |
| GLM-5.3 Flash, native MTP3 | [Four-Spark MTP3 cache/checkpoint quickstart](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_MTP3_CACHE_CHECKPOINTS_QUICKSTART.md) · [Implementation and issue status](deploy/glm53_flash/MTP3_STATUS.md) |
| GLM-5.3 Flash, DFlash2 | [Four-node DFlash2 quickstart](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_JJ_R8_GB10_SPARKCACHE_TP4_QUICKSTART.md) · [SparkCache integration notes](deploy/glm53_flash/README.md) |
| GLM-5.2 EXL3 3.5-bpw | [`deploy/glm52_35bpw/README.md`](deploy/glm52_35bpw/README.md) |
| DeepSeek-V4 | [`deploy/deepseek_v4/README.md`](deploy/deepseek_v4/README.md) |

Expand Down
80 changes: 80 additions & 0 deletions deploy/glm53_flash/MTP3_STATUS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# GLM-5.3 MTP3 cache implementation and issue status

Status: **implemented**, with bounded live evidence. The four-Spark native-MTP3
deployment remains **research-only**. Open performance issues are not evidence
that their related implementation PRs are unmerged; merged PRs are not evidence
that every reported workload is fixed.

## Deployable composition

Use SparkRing's
[MTP3 cache/checkpoint quickstart](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_MTP3_CACHE_CHECKPOINTS_QUICKSTART.md).
It selects a compatible image, transport bundle, recurrent checkpoints,
native placement library, and cache namespace together. Installing the Python
package alone does not install the runtime's GPU prefix-retention fixes.

The
[published image contract](https://github.com/FujitsuPolycom/sparkring/blob/main/runtime/glm53-spark-mtp3-mesh/performance/public-image.json)
pins SparkCache `48bbd2be4a7b972e56632a2d7b934bac5460f272` and includes
[SparkCache #62](https://github.com/FujitsuPolycom/sparkcache/pull/62) and
[#63](https://github.com/FujitsuPolycom/sparkcache/pull/63).
[SparkRing #236](https://github.com/FujitsuPolycom/sparkring/pull/236) integrates
the compute, transport, and cache-runtime changes from SparkRing #219, #226,
and #227. Those component PR numbers are not separate installation steps.

[SparkRing #237](https://github.com/FujitsuPolycom/sparkring/pull/237) implements
admission control that rejects public inference until readiness warmup completes.
That source change is merged but is **not included** in the published image
identified above. It requires a rebuilt image and startup validation.

## What remains in the performance issues

| Issue | Implemented | Remaining work |
|---|---|---|
| [#60: sustained publication and eviction slowdown](https://github.com/FujitsuPolycom/sparkcache/issues/60) | Fewer repeated restore reads and hashes, bounded restore arenas, tiled native placement, publication-dependency protection, publication-backlog gauges, optional periodic full captures, and capacity guidance. | Maintenance still scans and sorts the inventory and evicts toward the low watermark; it has no per-pass time or entry budget. Reproduce the reported slowdown with a near-full 40 GiB store and matched before/after probes. |
| [#61: growing conversations lose local prefix reuse](https://github.com/FujitsuPolycom/sparkcache/issues/61) | Runtime GPU-lease accounting, preference for longer local prefixes, recurrent checkpoint retention, and opt-in restore/lease traces. | Exact per-request local-hit, verified-restore, and recompute token counters are not implemented. Validate local retention on the reported long-context, multi-turn workload with occasional images. |

One saver admission per rank bounds concurrent optional work; it does not bound
the duration of an inventory scan or eviction pass. See
[capacity and cleanup](../../sparkcache/README.md#capacity-and-cleanup).
Backlog reports distinguish pending publications, their oldest age, and active
maintenance, but reports may not refresh while the connector is idle.

Reuse traces distinguish restore offers, verified worker completion, and GPU
lease attachment. Their scheduler prefix-token field is block-aligned input,
not an exact local hash-hit measurement. API `cached_tokens` alone is not enough
to attribute reuse to GPU retention rather than persistent restoration.

Both issues should remain open until their remaining implementation questions
and workload-specific results are recorded explicitly.

## Evidence and its limits

The
[cache-pressure validation record](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/glm53-flash/mtp3-cache-history-validation.md)
contains 551 successful responses on a related composition. It used text-only
traffic, a 2 GiB cache per rank, and opt-in periodic full captures. It is useful
regression evidence, but it does not reproduce the original 40 GiB workload.

The published image's
[source-equivalence record](https://github.com/FujitsuPolycom/sparkring/blob/main/performance/records/glm53-flash/mtp3-integrated-image-source-equivalence.md)
records 5,308 identical runtime files and excludes SparkCache source from that
equality claim. File equality does not constitute a serving soak.

The [exact-image serving record](https://github.com/FujitsuPolycom/sparkring/blob/86d2ebdff794da3d89f3ba1c6aca649ff6b052a0/performance/records/glm53-flash/mtp3-cache-checkpoints-serving-smoke-20260906.md)
qualifies eight short correctness checks, including a growing conversation,
streaming, reasoning-mode rejection, and repeated red/blue/red image responses.
All ranks published a 4,096-token context and were healthy afterward. The
40 GiB cache limit was configured, not filled; these checks do not establish
near-capacity performance or multimodal persistent restoration.

The published profile defaults to 40 GiB per rank, a 32 GiB low watermark, and
periodic full captures disabled. Setting a capacity is not a near-capacity test.
Keep the successful small-cache evidence; use short semantic and multimodal
checks after deployment, then collect near-capacity and local-reuse evidence
during representative operation before closing #60 or #61.

For an affected workload, `SPARKCACHE_ACCESS_MODE=restore-only` with
`SPARKCACHE_ASYNC_PAGE_CAPTURE=0` is a diagnostic workaround documented in the
issue reports. It stops persistence of newly encountered contexts. It is not
required merely because the issues remain open.
17 changes: 11 additions & 6 deletions deploy/glm53_flash/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,16 +6,21 @@ repository assembles the model runtime, transport, image, and operator settings.

## Run the model

Use the
[four-node SparkRing quickstart](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_JJ_R8_GB10_SPARKCACHE_TP4_QUICKSTART.md).
It uses one Linux/ARM64 image for TP4 with DCP1, DCP2, or DCP4 and documents
both modes:
Choose the guide matching the speculative runtime:

- [Native MTP3 with caching and recurrent checkpoints](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_MTP3_CACHE_CHECKPOINTS_QUICKSTART.md)
uses four Sparks at TP4/DCP4, with no external draft model. Its
[implementation and issue status](MTP3_STATUS.md) separates merged code,
published images, live evidence, and remaining work for issues #60 and #61.
- [DFlash2 with SparkCache](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/GLM53_JJ_R8_GB10_SPARKCACHE_TP4_QUICKSTART.md)
uses a Linux/ARM64 image for TP4 with DCP1, DCP2, or DCP4 and documents
both modes below.

- `SPARKCACHE_ENABLED=1` enables persistent SparkCache alongside vLLM's GPU
prefix cache.
- `SPARKCACHE_ENABLED=0` runs with vLLM's GPU prefix cache alone.

The quickstart is the source of truth for the image digest, model revisions,
The selected quickstart is the source of truth for the image digest, model revisions,
launch variables, storage paths, and four-host procedure. Keeping those values
in SparkRing prevents a copied deployment recipe here from becoming stale.

Expand All @@ -39,7 +44,7 @@ topology, launch settings, and cache namespace supplied by the deployment.

GLM-5.3 execution and GB10 performance depend primarily on Local Inference
Lab's [Jovian Judgement vLLM work](https://github.com/local-inference-lab/vllm/tree/dev/jovian-judgement)
and [B12X kernels](https://github.com/local-inference-lab/b12x). The SparkRing
and [B12X kernels](https://github.com/local-inference-lab/b12x). The DFlash2 SparkRing
quickstart identifies the exact revisions and model artifacts, including
[GLM-5.3-Flash-NVFP4](https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4)
and the BF16
Expand Down