Skip to content

Source Based Dataset Layout

Peter Pak edited this page Jul 16, 2026 · 6 revisions

Source-Based Dataset Layout

Design agreed 2026-07-15 — not yet implemented. Effective from build 47; no retroactive migration of builds 1–46 except where noted.

Principle

Every HF dataset repo is self-contained and follows one convention:

datasets/<repo>/
  source/    raw feed data, committed to HF (the archive)
  scripts/   processing: source/ → data/
  data/      published, analysis-ready tables (parquet/jsonl)

Database, ASTM, and Conversations already work this way. This design promotes the recorder's output (the one holdout) to the same pattern and retires data/exports/ as a staging layer. Generated data has one home: the owning dataset's source/ tree.

Write path (unchanged during prints)

The print-time rule stays exactly as deployed 2026-07-01: live capture goes to the NVMe SSD, never to a git tree or the SMR HDD's random-write path. Telemetry/position/frame-metadata spool to ~/.agentic-sls/spool/; frames buffer locally. Post-print, a delivery step moves everything to its final home. The recorder never depends on submodule checkout state — if the dataset submodule is absent, delivery fails loudly and the buffer stays put.

Post-print delivery (new pipeline step)

Runs after the spool import completes, per build:

  1. Raw ticks — spool NDJSON (telemetry, position) gzipped into Agentic-SLS-Telemetry/source/spool/<build>/. Replaces the current ".imported files accumulate on the NVMe until cleaned manually" TODO — delivery IS the cleanup.
  2. Frames (settled 2026-07-15 — see "Frames flow" below) — store-mode zips of ≤10k frames per chunk into Agentic-SLS-Telemetry/source/frames/<build>/<kind>/chunk-NNNN.zip (complete native-rate originals) + data/frames_index/ parquet. The ML serving form stays the existing ticks config (which already embeds tick-aligned frames). frames.path rows updated to chunk-NNNN.zip::<member>; Replay serves via ranged reads. The live buffer is transient, not a permanent home. Tool: sls-deliver-frames.
  3. Recorder metadata — builds/events export retargeted from data/exports/ into Agentic-SLS-Telemetry/source/recorder/.
  4. Processing — the dataset repo's own scripts turn source/ into data/ticks/ parquet etc. (already the existing ETL, re-pointed from data/exports/ to its own source/).
  5. Commit — one dataset commit per build is the archival act.

Unchanged: Conversations (already source-based), Postgres (stays the agent query surface — it is derived state, not publication), sidecar CSVs re-homed 2026-07-16: sensors.csv → Telemetry source/recorder/, build_to_inova_session.csv → Database source/ (beside the PrintSessions it annotates; the matcher already requires that checkout).

Corrected facts (measured 2026-07-15)

  • CORRECTED 2026-07-15 (schema-verified): the ticks parquet already embeds frames — each 10 Hz tick row carries frame_chamber / frame_galvo / frame_thermal (HF image features), the raw bedmatrix matrix, and a position_hf_burst list. load_dataset serves images today. What the dataset does NOT have: the complete native-rate originals (chamber captures ~20 fps; ticks are a 10 Hz grid), loose-file access, or a source/ tree. data/frames/ (408 GB) remains the only complete-rate copy — the source zips are ITS first offsite archive.
  • NVMe headroom: ~90 GB free — comfortable for per-build spool + frame buffer (~9 GB/build average), but check before unusually long builds.
  • data/ end state: postgres/ + transient buffers only; exports/ retired after migration.

Frames flow (settled 2026-07-15)

Measured reality that shaped this: builds produce ~1M loose frame files each (chamber streams at native ~20 fps; build 44 = 1,247,262 files / 38 GB; build 40 = 58 GB). Loose files are non-negotiable-impossible on HF (millions of files) and a single zip per build exceeds HF's hard 50 GB per-file limit (build 40). Hence chunking.

PRINT      frames → NVMe buffer (fast IO; ~90 GB headroom is fine —
           builds are capped ≲12 h and the buffer is cleared before the
           next print; free-space check at print start, HDD fallback)
POST-PRINT 1. verify frame count vs spool metadata
           2. store-mode zip per kind, ≤10k frames per chunk, written
              directly into source/frames/<build>/<kind>/ (no middle copy)
           3. ML serving form = the EXISTING ticks config (already
              embeds tick-aligned frames; regenerated by the dataset's
              own ETL as before). Standalone embedded-image shards are
              available via `sls-deliver-frames --image-shards` but OFF
              by default — they'd be a third copy of the bytes.
           4. index parquet (ts, kind, archive, member, size) +
              frames.path rows updated in Postgres
           5. verify archives (member counts + spot hashes) →
              commit + push to HF
           6. push confirmed → delete NVMe originals

Duplication on HF: frame bytes live twice (source zips = complete native-rate originals; ticks parquet = tick-aligned ML form). Keep zips store-mode so Xet's content-defined chunking has a chance to dedup identical byte runs server-side.

Open decisions

  1. Old-era backfill: builds 1–46 frames (408 GB, single-copy!) and the old-era telemetry parquet in data/exports/telemetry/ move into source/ as a scheduled bulk job, separate from the per-build flow. sls-summarize-builds re-points when that happens.

Migration order

  1. Decide frames repo + packing — settled: Telemetry repo, chunked store-mode zips + embedded-image parquet (see "Frames flow").
  2. Build the delivery step; run it for build 47 onward.
  3. Verify one full cycle (print → spool → import → deliver → process → commit → push), then retire data/exports/ and re-point consumers.
  4. Bulk-backfill the historical archives; only after HF push verification, delete the local duplicates.

Clone this wiki locally