Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
5c8e722
Distribute verified image archives through direct fabric
FujitsuPolycom Aug 30, 2026
75809e4
Bound GLM-5.3 research summary to valid observations
FujitsuPolycom Aug 30, 2026
d2b9d80
Merge direct-fabric image fanout tooling
FujitsuPolycom Aug 30, 2026
0e00b07
Merge exact GLM runtime evidence
FujitsuPolycom Aug 30, 2026
aeb7162
Consolidate GLM-5.3 operator guidance
FujitsuPolycom Aug 30, 2026
38148eb
Bound GLM concurrency by resident KV capacity
FujitsuPolycom Aug 30, 2026
d78f175
Correct GLM hybrid concurrency guidance
FujitsuPolycom Aug 30, 2026
47d8f5d
Default DFlash7 to safe snapshot restore
FujitsuPolycom Aug 30, 2026
98a2b10
Bind snapshot fallback to exact image source
FujitsuPolycom Aug 30, 2026
72df8b0
Add a concise four-rank GLM quickstart
FujitsuPolycom Aug 30, 2026
65d71d1
Record rejected four-reader restore candidate
FujitsuPolycom Aug 30, 2026
7de9c16
Align quickstart artifact evidence test
FujitsuPolycom Aug 30, 2026
8e4d8d8
Record qualified split-page C8 runtime
FujitsuPolycom Aug 31, 2026
ca91fa7
Expose qualified GLM SparkCache runtime settings
FujitsuPolycom Aug 31, 2026
f695f9a
Credit GLM runtime and quantization sources
FujitsuPolycom Aug 31, 2026
ecfb502
Publish GLM-5.3 JJ r7-compatible GB10 run contracts
FujitsuPolycom Aug 31, 2026
54d9df7
Clarify historical GLM runtime names
FujitsuPolycom Aug 31, 2026
e9b9d8f
Default GLM serving to 512K context
FujitsuPolycom Aug 31, 2026
539b7a0
Separate SparkRing subsystems from model profiles
FujitsuPolycom Aug 31, 2026
b5660e0
Add GLM DCP2 and DCP4 launch contracts
FujitsuPolycom Aug 31, 2026
3b5e662
Build consolidated GLM R8 SparkCache image
FujitsuPolycom Aug 31, 2026
d803c6b
Correct R8 SparkCache image defaults
FujitsuPolycom Aug 31, 2026
89886ee
Enable R8 prefill cadence by default
FujitsuPolycom Aug 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -1,7 +1,10 @@
*.sh text eol=lf
*.patch text eol=lf
runtime/** text eol=lf
performance/receipts/glm53-flash/dflash7-snapshot-v1-safe/*.json text eol=lf
performance/receipts/glm53-flash/split-page-shared-base-c8-20260830/*.json text eol=lf
runtime/exl3/patches/*.patch whitespace=-trailing-space
runtime/glm53-flash-split-page-sparkcache/patches/*.patch whitespace=-trailing-space

# Runtime receipts hash these source files byte-for-byte before copying them,
# and the plugin vendors byte-identical copies whose SHA-256 it records in
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ jobs:
--index-url https://download.pytorch.org/whl/cpu \
"torch==${TORCH_VERSION}"
- name: Test maintained Python trees
run: python -m pytest spark_transport runtime/exl3-r7 runtime/glm53-flash runtime/glm53-flash-adaptive-mtp-python-overlay runtime/glm53-flash-dflash7-python-overlay runtime/deepseek0731-gb10 runtime/qwen38 runtime/test_public_overlay.py performance/harnesses scripts -q -rs
run: python -m pytest spark_transport runtime/exl3-r7 runtime/glm53-flash runtime/glm53-flash-jj-r7-gb10 runtime/glm53-flash-split-page-sparkcache runtime/glm53-flash-adaptive-mtp-python-overlay runtime/glm53-flash-dflash7-python-overlay runtime/deepseek0731-gb10 runtime/qwen38 runtime/test_public_overlay.py performance/harnesses scripts -q -rs

docs-links:
name: docs links
Expand Down
34 changes: 18 additions & 16 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ conversation or development history.
- Keep operator instructions focused on what to do and what result to expect.
- Put lane, maturity, hardware, and evidence metadata in one table or callout
instead of repeating it throughout the prose.
- Prefer `the new settings are still being tested` over phrases such as
- Prefer `the selected settings are still being tested` over phrases such as
`candidate target changes`, `silent promotion`, `historical qualification
scope`, or `revalidate the composition` when the plain statement is accurate.
- Use exact formal vocabulary only where a machine-readable contract or release
Expand All @@ -40,10 +40,11 @@ SparkRing supports four model families across ten deployment profiles:

- GLM-5.2 EXL3 3.5-bpw at four-Spark TP4/DCP4, as a base profile and a
SparkCache composition, using the R7 runtime and site/candidate contracts.
- GLM-5.3 Flash with the public BF16 DFlash2 drafter at four-Spark TP4/DCP1,
as a cache-disabled base profile and a SparkCache composition, using the
immutable source contract in `runtime/glm53-flash/` and the sanitized site
and runtime templates in `scripts/config/`.
- GLM-5.3 Flash with the public BF16 DFlash2 drafter at four-Spark TP4. The
cache-disabled profile supports DCP1, DCP2, and DCP4; the SparkCache
composition remains DCP1, using the
immutable published JJ r7-compatible artifacts and environment contract in
`runtime/glm53-flash-jj-r7-gb10/`.
- DeepSeek-V4-Flash-0731 at two-Spark TP2/DCP1 and four-Spark TP4/DCP1, as
base profiles and SparkCache compositions, using the published serving image
and per-rank environment contracts.
Expand Down Expand Up @@ -71,8 +72,9 @@ choosing one.
| Subject | Canonical source |
|---|---|
| GLM-5.2 EXL3 runtime build | `runtime/exl3-r7/README.md` |
| GLM-5.3 Flash runtime, model, DFlash, NCCL, and SparkCache pins | `runtime/glm53-flash/pins.json` |
| GLM-5.3 Flash four-rank site and runtime profiles | `scripts/config/glm53-flash-tp4-site.example.yaml`, then the selected `scripts/config/glm53-flash-dflash2-bf16-tp4-dcp1*.example.json` |
| Published JJ r7-compatible GLM-5.3 images, source/model identities, and smoke evidence | `runtime/glm53-flash-jj-r7-gb10/artifacts.json` |
| Published JJ r7-compatible GLM-5.3 four-rank operator settings | `runtime/glm53-flash-jj-r7-gb10/runtime.env.example` |
| Historical GLM-5.3 split-page local artifact | `runtime/glm53-flash-split-page-sparkcache/qualified-artifact.json` |
| Qwen3.8-27B runtime build | `runtime/qwen38/README.md` |
| GLM/DeepSeek runtime base and model pins | `runtime/faststart-lock.json` |
| Qwen runtime/source pins and model identities | `runtime/qwen38/pins.json`, then `recipes/qwen38-27b-exl3-k5k6{,-pair}.json` |
Expand Down Expand Up @@ -104,7 +106,7 @@ test imports torch:
python -m pip install -r requirements-dev.txt
python -m pip install --index-url https://download.pytorch.org/whl/cpu "torch==2.11.0"
ruff check --select E,F,W --ignore E501 spark_transport runtime scripts performance
python -m pytest spark_transport runtime/exl3-r7 runtime/glm53-flash runtime/glm53-flash-adaptive-mtp-python-overlay runtime/glm53-flash-dflash7-python-overlay runtime/deepseek0731-gb10 runtime/qwen38 runtime/test_public_overlay.py performance/harnesses scripts -q -rs
python -m pytest spark_transport runtime/exl3-r7 runtime/glm53-flash runtime/glm53-flash-jj-r7-gb10 runtime/glm53-flash-split-page-sparkcache runtime/glm53-flash-adaptive-mtp-python-overlay runtime/glm53-flash-dflash7-python-overlay runtime/deepseek0731-gb10 runtime/qwen38 runtime/test_public_overlay.py performance/harnesses scripts -q -rs
```

The test suite is CPU-only contract coverage. It does not validate CUDA,
Expand All @@ -113,8 +115,10 @@ RDMA, live pair/cycle serving, or a performance result.
## Runtime and configuration work

`runtime/exl3-r7/` builds the GLM-5.2 EXL3 R7 image.
`runtime/glm53-flash/` records the GLM-5.3 Flash runtime, model, public BF16
DFlash2, patched NCCL, SparkCache, and vLLM lease-contract identities.
`runtime/glm53-flash-jj-r7-gb10/` records the published JJ r7-compatible GLM-5.3 base and
SparkCache images, source/model identities, operator environment, and bounded
TP4 smoke evidence. `runtime/glm53-flash/` retains the historical published
artifact contract.
`runtime/deepseek0731-gb10/` builds the hardened DeepSeek-V4-Flash-0731 image,
and `runtime/qwen38/` builds the Qwen3.8-27B ARM64 image.
`runtime/faststart-lock.json` pins the generic GLM/rollback image, the hardened
Expand All @@ -129,12 +133,10 @@ version control. Use `scripts/config/deepseek-v4-flash-0731-pair.env.example`
for a two-rank pair and `scripts/config/deepseek-v4-flash-0731.env.example` for
a four-rank cycle.

Use `scripts/config/glm53-flash-tp4-site.example.yaml` with exactly one of the
two GLM-5.3 runtime-profile templates. Keep resolved site addresses, host
paths, image IDs, and credentials outside version control. Both GLM-5.3
profiles use the same SparkCache-capable image and preserve asynchronous
scheduling, native prefix caching, and chunked prefill; only the
SparkCache-named profile enables the external connector.
Use `runtime/glm53-flash-jj-r7-gb10/runtime.env.example` for the published JJ r7-compatible GLM-5.3
serving. Keep resolved site addresses, host paths, and credentials outside
version control. `IMAGE_VARIANT` selects the base or SparkCache digest; only
the SparkCache variant enables the external connector.

Use the topology-specific Qwen environment in `scripts/config/` with the image
built by `runtime/qwen38/build-image.sh`, as described in
Expand Down
5 changes: 3 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,9 @@ code are welcome. You do not need DGX Spark hardware or maintainer approval to
participate. Partial reports are useful; share what you know and maintainers
can help identify what is missing.

The supported deployments and their maturity are listed in
[`README.md`](README.md). This guide intentionally does not repeat that matrix.
Deployment profiles and their evidence scope are listed in the
[profile registry](docs/profiles/README.md). This guide intentionally does not
repeat that matrix.

## Issues and discussions

Expand Down
159 changes: 70 additions & 89 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,123 +3,104 @@
> Repository branches are mutable. Pin an immutable commit when reproducing a
> deployment.

SparkRing is a low-latency collective transport and vLLM-based
inference-serving stack for switchless clusters of NVIDIA DGX 'Spark' systems
powered by the GB10 Grace Blackwell Superchip.
SparkRing is a collective-communication and inference-serving stack for
switchless clusters of NVIDIA DGX Spark systems. It connects each system
directly to its neighbours and runs distributed workloads without an external
Ethernet or InfiniBand fabric switch.

SparkRing supports GB10 pairs and four-node rings. Six-node ring work is
research-only.
The repository contains:

Models run as tensor-parallel deployments over the direct fabric without an
external Ethernet or InfiniBand switch. [SIRCL](https://github.com/FujitsuPolycom/sparkring/blob/main/docs/SIRCL.md) provides custom RDMA collectives
where applicable, CUDA-graph command rings support repeated decode
work, and [patched NCCL](https://github.com/FujitsuPolycom/sparkring/blob/main/spark_transport/nccl/README.md) handles communication outside SIRCL's supported paths.
- cluster discovery, configuration, and diagnostics;
- native RDMA collectives and vLLM integration;
- patched NCCL support for communication outside the native transport;
- reproducible runtime builders and deployment profiles; and
- benchmark methods, records, and sanitized receipts.

The repository provides launch tooling, model profiles, test evidence, and [performance data](https://github.com/FujitsuPolycom/sparkring/tree/main/performance).
## Get started

## Setup
Start on the system that will serve as rank 0:

Start with ssh to node0 and enough disk space for the intended model weights.
Use the [bootstrap guide](docs/BOOTSTRAP.md). The model-independent `sparkring cluster
init` workflow enrolls nodes, discovers management and ConnectX-7 hardware,
generates the cluster inventory, and launches Ring Doctor before any model
profile is selected.
```bash
curl -fLO \
https://raw.githubusercontent.com/FujitsuPolycom/sparkring/main/bootstrap.sh
less bootstrap.sh
bash bootstrap.sh
sparkring cluster init --size 4
sparkring doctor --verify
```

## Resources
The bootstrap workflow enrolls nodes over the management network, discovers
ConnectX-7 hardware, writes a local cluster inventory, and checks the direct
fabric before any serving profile is selected.

- [Supported models and profiles](#profiles)
- [Benchmark results](#benchmark-results)
- [Deployment prerequisites](docs/PREREQUISITES.md) — then choose a profile quickstart below
Read the [bootstrap guide](docs/BOOTSTRAP.md) before applying network changes.
Then choose a deployment from the [profile registry](docs/profiles/README.md).

## Profiles
## Topologies

### GLM-5.3 Flash profiles
| Topology | Status | Fabric |
|---|---|---|
| Two-system pair | **implemented** | One direct 200 Gb/s link |
| Four-system cycle | **implemented** | Four direct links in a closed `0-1-2-3-0` cycle |
| Six-system cycle | **research-only** for serving | Six direct links in a closed cycle |

| Profile | Deployment | Context | Seqs | Batch | KV / cache | Start here |
|---|---|---:|---:|---:|---|---|
| GLM-5.3 Flash + BF16 DFlash2 | 4 Sparks · TP4/DCP1 | 512K | 32 | 8,192 | 12 GiB/rank FP8 KV | [Quickstart](docs/GLM53_FLASH_DFLASH2_BF16_TP4_QUICKSTART.md) |
| GLM-5.3 Flash + BF16 DFlash2 + SparkCache | 4 Sparks · TP4/DCP1 | 512K | 32 | 8,192 | 12 GiB/rank FP8 KV + 48 GiB nvme/rank SparkCache | [Quickstart](docs/GLM53_FLASH_DFLASH2_BF16_SPARKCACHE_TP4_QUICKSTART.md) |
Each rank also needs a management-network connection for SSH, rendezvous, and
the rank-0 API. The management network is not an inference-fabric edge.

### Other model profiles
## Communication paths

| Profile | Deployment | Context | Seqs | Batch | KV / cache | Start here |
|---|---|---:|---:|---:|---|---|
| GLM-5.2 EXL3 3.5-bpw | 4 Sparks · TP4/DCP4 | 1M | 16 | 4,096 | NVFP4 DS-MLA · 9.25 GB/rank | [Quickstart](docs/GLM52_35BPW_QUICKSTART.md) |
| DeepSeek-V4-Flash-0731 | 2 Sparks · TP2/DCP1 | 1M | 32 | 4,096 | FP8 DS-MLA · 16 GiB/rank | [Quickstart](docs/DEEPSEEK_V4_FLASH_QUICKSTART.md) |
| DeepSeek-V4-Flash-0731 | 4 Sparks · TP4/DCP1 | 1M | 32 | 4,096 | FP8 DS-MLA · 16 GiB/rank | [Quickstart](docs/DEEPSEEK_V4_FLASH_QUICKSTART.md) |
| Qwen3.8-27B EXL3 K5/K6 | 2 Sparks · TP2/DCP1 | 1M | 32 | 8,192 | FP8 | [Quickstart](docs/QWEN38_27B_EXL3_K5K6_PAIR_QUICKSTART.md) |
| Qwen3.8-27B EXL3 K5/K6 | 4 Sparks · TP4/DCP1 | 1M | 64 | 8,192 | FP8 | [Quickstart](docs/QWEN38_27B_EXL3_K5K6_QUICKSTART.md) |
| GLM-5.2 EXL3 3.5-bpw + SparkCache | 4 Sparks · TP4/DCP4 | 1M | 16 | 4,096 | NVFP4 DS-MLA + SparkCache | [SparkCache compositions](recipes/sparkcache/README.md) |
| DeepSeek-V4-Flash-0731 + SparkCache | 2 Sparks · TP2/DCP1 | 1M | 32 | 4,096 | FP8 DS-MLA + SparkCache | [SparkCache compositions](recipes/sparkcache/README.md) |
| DeepSeek-V4-Flash-0731 + SparkCache | 4 Sparks · TP4/DCP1 | 1M | 32 | 4,096 | FP8 DS-MLA + SparkCache | [SparkCache compositions](recipes/sparkcache/README.md) |
[SIRCL](docs/SIRCL.md) provides persistent RDMA sessions and graph-replayable
collectives for supported four-rank shapes. [Patched
NCCL](spark_transport/nccl/README.md) handles other collective shapes and
phases. A deployment profile states which path it uses.

The public BF16 DFlash2 checkpoint is licensed CC BY-NC-ND 4.0 for research
and evaluation; review its model card before use. The qualified GLM-5.3 community
images are published by immutable digest in the two quickstarts. Both guides also
link the complete source-build recipes, SBOM workflow, source commits, applied patches, and license record.
See the [profile registry](docs/profiles/README.md) for recipe identities and evidence scope.
See [architecture](docs/ARCHITECTURE.md), [deployment
prerequisites](docs/PREREQUISITES.md), and [cable
qualification](spark_transport/CABLE_QUALIFICATION.md) for the system
contracts.

## Benchmark results
## Profiles and evidence

### GLM-5.3 Flash research observation
Model, runtime, topology, memory, and scheduler settings belong to deployment
profiles rather than the transport definition:

**Research-only — 16K context, single observation.** The SparkCache-enabled
profile recorded 2,371 tok/s prefill and 36.06 tok/s sustained C1 decode on
random tokens. No A/B baseline has been completed. C4 and C8 were capacity-limited
and are omitted rather than reported as throughput results.
- [profile registry](docs/profiles/README.md);
- [runtime builders](runtime/README.md);
- [serving configuration templates](scripts/config/README.md);
- [benchmark results](docs/RESULTS.md); and
- [methods, records, and receipts](performance/README.md).

| Profile | Prefill | C1 decode | C8 decode | Highest valid decode | Coding peak |
|---|---:|---:|---:|---:|---:|
| [GLM-5.3 Flash + BF16 DFlash2 + SparkCache · 4 Sparks](performance/records/glm53-flash/sparkcache-dflash2-bf16-tp4-16k-run1-20260829.md) | 2,371 | 36.06 | — | C1: 36.06 | — |

### Other model profiles

**Qualified — 16K context.** All values are tokens per second. Prefill uses a
cold prompt with caching disabled. Decode uses unique, cold prompt contexts at
temperature 1.0; decode values are aggregate throughput across active streams.

| Profile | Prefill | C1 decode | C8 decode | Highest tested decode | Coding peak |
|---|---:|---:|---:|---:|---:|
| [GLM-5.3flash NVFP4 · 4 Sparks](IN PROGRESS) | 2300 | 40 | 130 | C8: 130 | 70 |
| [GLM-5.2 EXL3 3.5-bpw · 4 Sparks](performance/records/glm-3.5bpw/normalized-base-20260822.md) | 671 | 20.15 | 64.13 | C8: 64.13 | 25.39 |
| [DeepSeek-V4-Flash DSpark · 2 Sparks](performance/records/deepseek-v4-flash/normalized-tp2-base-temp1-n5-20260823.md) | 1,926 | 58.36 | 162.69 | C32: 307.13 | 59.31 |
| [DeepSeek-V4-Flash-0731 · 4 Sparks](performance/records/deepseek-v4-flash/normalized-tp4-base-temp1-n5-20260823.md) | 2,488 | 68.84 | 265.16 | C32: 508.11 | 95.77 |
| [Qwen3.8-27B EXL3 K5/K6 · 2 Sparks](performance/records/qwen38-27b/normalized-tp2-1m-probmtp-temp1-20260823.md) | 1,367 | 29.50 | 142.20 | C8: 142.20 | 39.95 |
| [Qwen3.8-27B EXL3 K5/K6 · 4 Sparks](performance/records/qwen38-27b/normalized-tp4-1m-probmtp-temp1-20260823.md) | 1,964 | 35.07 | 191.02 | C8: 191.02 | 48.46 |

See [benchmark results and throughput tables](docs/RESULTS.md) for full
matrices, sample counts, exact settings, and limitations.

## Architecture

Two-Spark profiles use one direct 200 Gb/s cable and patched NCCL. Four-Spark
profiles use a switchless `0-1-2-3-0` cable cycle. SIRCL serves its tested
collective paths; patched NCCL handles the remaining paths.

See [architecture](docs/ARCHITECTURE.md), [SIRCL](docs/SIRCL.md), and the
[deployment prerequisites](docs/PREREQUISITES.md).
A successful offline test or process start does not establish serving
correctness, output quality, capacity, or performance. Each profile documents
its own evidence and limitations.

## Repository map

| Path | Purpose |
|---|---|
| `spark_transport/` | Native transport and vLLM adapters |
| `runtime/` | Pinned runtime inputs and builders |
| `scripts/` | Site validation, preflight, launch, and evidence tooling |
| `recipes/` | Machine-readable serving recipes |
| `spark_transport/` | Native transport, patched NCCL, and framework adapters |
| `runtime/` | Pinned runtime inputs and image builders |
| `scripts/` | Cluster validation, launch, and evidence tooling |
| `recipes/` | Machine-readable serving compositions |
| `performance/` | Benchmark methods, records, and sanitized receipts |
| `docs/` | Profile procedures, architecture, prerequisites, and evidence |
| `docs/` | Architecture, prerequisites, profiles, and operator procedures |

## Development

See [CONTRIBUTING.md](CONTRIBUTING.md) for contribution and validation
guidance. Security reports belong in private GitHub security advisories as
described in [SECURITY.md](SECURITY.md).

## Acknowledgements

SparkRing builds on work from vLLM, NVIDIA NCCL, B12X, SparkInfer, LMCache,
ExLlamaV3, and the
[local inference community](https://github.com/local-inference-lab/).
SparkRing builds on vLLM, NVIDIA NCCL, B12X, SparkInfer, LMCache, ExLlamaV3,
and work published by the [Local Inference
Lab](https://github.com/local-inference-lab/). Model-specific profiles identify
their exact upstream source, checkpoint, revision, and license.

Detailed third-party attribution is in
[`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md).
[THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).

## License

Apache-2.0. See [`LICENSE`](LICENSE).
Apache-2.0. See [LICENSE](LICENSE).
4 changes: 3 additions & 1 deletion SECURITY.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
# Security Policy

SparkRing is pre-release research software. There is no supported-versions table: the only supported version is the latest commit on `main`.
SparkRing is pre-release research software and does not publish a
supported-versions table. Security fixes are applied to `main`; deployment
profiles separately identify the immutable revisions covered by their evidence.

## Reporting a vulnerability

Expand Down
Loading
Loading