-
Notifications
You must be signed in to change notification settings - Fork 0
Source Based Dataset Layout
Design agreed 2026-07-15 — not yet implemented. Effective from build 47; no retroactive migration of builds 1–46 except where noted.
Every HF dataset repo is self-contained and follows one convention:
datasets/<repo>/
source/ raw feed data, committed to HF (the archive)
scripts/ processing: source/ → data/
data/ published, analysis-ready tables (parquet/jsonl)
Database, ASTM, and Conversations already work this way. This design
promotes the recorder's output (the one holdout) to the same pattern and
retires data/exports/ as a staging layer. Generated data has one home:
the owning dataset's source/ tree.
The print-time rule stays exactly as deployed 2026-07-01: live capture
goes to the NVMe SSD, never to a git tree or the SMR HDD's random-write
path. Telemetry/position/frame-metadata spool to ~/.agentic-sls/spool/;
frames buffer locally. Post-print, a delivery step moves everything to its
final home. The recorder never depends on submodule checkout state — if the
dataset submodule is absent, delivery fails loudly and the buffer stays put.
Runs after the spool import completes, per build:
-
Raw ticks — spool NDJSON (telemetry, position) gzipped into
Agentic-SLS-Telemetry/source/spool/<build>/. Replaces the current ".imported files accumulate on the NVMe until cleaned manually" TODO — delivery IS the cleanup. -
Frames (settled 2026-07-15 — see "Frames flow" below) — store-mode
zips of ≤10k frames per chunk into
Agentic-SLS-Telemetry/source/frames/<build>/<kind>/chunk-NNNN.zip(complete native-rate originals) +data/frames_index/parquet. The ML serving form stays the existing ticks config (which already embeds tick-aligned frames).frames.pathrows updated tochunk-NNNN.zip::<member>; Replay serves via ranged reads. The live buffer is transient, not a permanent home. Tool:sls-deliver-frames. -
Recorder metadata — builds/events export retargeted from
data/exports/intoAgentic-SLS-Telemetry/source/recorder/. -
Processing — the dataset repo's own scripts turn
source/intodata/ticks/parquet etc. (already the existing ETL, re-pointed fromdata/exports/to its ownsource/). - Commit — one dataset commit per build is the archival act.
Unchanged: Conversations (already source-based), Postgres (stays the agent
query surface — it is derived state, not publication), sidecar CSVs
re-homed 2026-07-16: sensors.csv → Telemetry source/recorder/,
build_to_inova_session.csv → Database source/ (beside the
PrintSessions it annotates; the matcher already requires that checkout).
- CORRECTED 2026-07-15 (schema-verified): the ticks parquet already
embeds frames — each 10 Hz tick row carries
frame_chamber/frame_galvo/frame_thermal(HFimagefeatures), the rawbedmatrixmatrix, and aposition_hf_burstlist.load_datasetserves images today. What the dataset does NOT have: the complete native-rate originals (chamber captures ~20 fps; ticks are a 10 Hz grid), loose-file access, or asource/tree.data/frames/(408 GB) remains the only complete-rate copy — the source zips are ITS first offsite archive. - NVMe headroom: ~90 GB free — comfortable for per-build spool + frame buffer (~9 GB/build average), but check before unusually long builds.
-
data/end state:postgres/+ transient buffers only;exports/retired after migration.
Measured reality that shaped this: builds produce ~1M loose frame files each (chamber streams at native ~20 fps; build 44 = 1,247,262 files / 38 GB; build 40 = 58 GB). Loose files are non-negotiable-impossible on HF (millions of files) and a single zip per build exceeds HF's hard 50 GB per-file limit (build 40). Hence chunking.
PRINT frames → NVMe buffer (fast IO; ~90 GB headroom is fine —
builds are capped ≲12 h and the buffer is cleared before the
next print; free-space check at print start, HDD fallback)
POST-PRINT 1. verify frame count vs spool metadata
2. store-mode zip per kind, ≤10k frames per chunk, written
directly into source/frames/<build>/<kind>/ (no middle copy)
3. ML serving form = the EXISTING ticks config (already
embeds tick-aligned frames; regenerated by the dataset's
own ETL as before). Standalone embedded-image shards are
available via `sls-deliver-frames --image-shards` but OFF
by default — they'd be a third copy of the bytes.
4. index parquet (ts, kind, archive, member, size) +
frames.path rows updated in Postgres
5. verify archives (member counts + spot hashes) →
commit + push to HF
6. push confirmed → delete NVMe originals
Duplication on HF: frame bytes live twice (source zips = complete native-rate originals; ticks parquet = tick-aligned ML form). Keep zips store-mode so Xet's content-defined chunking has a chance to dedup identical byte runs server-side.
-
Old-era backfill: builds 1–46 frames (408 GB, single-copy!) and the
old-era telemetry parquet in
data/exports/telemetry/move intosource/as a scheduled bulk job, separate from the per-build flow.sls-summarize-buildsre-points when that happens.
-
Decide frames repo + packing— settled: Telemetry repo, chunked store-mode zips + embedded-image parquet (see "Frames flow"). - Build the delivery step; run it for build 47 onward.
- Verify one full cycle (print → spool → import → deliver → process →
commit → push), then retire
data/exports/and re-point consumers. - Bulk-backfill the historical archives; only after HF push verification, delete the local duplicates.