-
Notifications
You must be signed in to change notification settings - Fork 0
Datasets
The publication layer: five HuggingFace dataset repos, all git submodules under datasets/ in the main repo. They are downstream of Postgres (or their own lab/firmware sources) — never the agent's query surface.
✅ Migration executed 2026-07-16 (see Source Based Dataset Layout): exporters and ETLs are re-pointed —
sls-export→ Telemetrysource/, linkage CSV → Databasesource/, conversations written in place.data/exports/survives only as a fallback copy until the frames backfill completes.
| Dataset | Contents |
|---|---|
Agentic-SLS-Telemetry |
Raw per-build telemetry parquet (566 GB) |
Agentic-SLS-Database |
Firmware jobs, print profiles, PrintSessions |
Agentic-SLS-ASTM |
D638 tensile / D790 flex specimen results (108 SLS specimens, all with job FKs) |
Agentic-SLS-Knowledge |
SLS knowledge broadly (Inova + general: TDS sheets, machine data) and the agent-facing reference corpus served by reference_* (source/*.md + data/knowledge.jsonl); unlike the big datasets it must ALWAYS be initialized — the MCP tools read it at runtime. Phase-4 playbook will live here. |
Agentic-SLS-Conversations |
Agent transcripts — raw session files are written directly into its source/sessions/ by the harness store (it is the archive; commit it often) |
- The four data-heavy repos have
update = nonein.gitmodulesso a blanket--recurse-submodulesskips them (Telemetry alone is 566 GB). Never remove that guard; init them explicitly. Agentic-SLS-Knowledge is the deliberate exception — noupdate = none, it must always be initialized because thereference_*tools serve from it at runtime. - The
git@hf.co:remotes need the HuggingFace SSH key. - Conversations' HF remote slug is
Agentic-SLS-Conversations(typo) — intentional and baked in; redirects work; don't fix it.
HF repos consume flat-file exports from the main repo's data/exports/ (JSONL + parquet + sidecar CSVs) via relative paths — never the live Postgres. Why: dataset ETL needs clean, container-free, idempotent inputs; hitting the live DB fails when it's down and competes with the live writer.
- Producer:
uv run sls-export. Idempotency contract: existing parquet preserved unless--force; sidecar CSVs (build_to_inova_session.csv,sensors.csv) preserve hand-filled values across re-runs. - Conversations exporter:
uv run sls-export-conversations. - Sidecar CSVs are the pattern for any mapping that needs human authoring — checked in, authoritative. Don't add recorder-DB columns just to serve a downstream HF consumer; push the mapping into
data/exports/. - Path conventions: Telemetry's and Conversations'
scripts/_lib.pyclimb two parents to the repo root; ASTM finds Database as a sibling insidedatasets/.
Postgres holds synced copies of the Database and ASTM repos' data (inova_jobs, inova_print_profiles, inova_print_sessions, astm_specimens). Corrections happen at the origin repo, then uv run sls-sync-reference re-syncs. Never fix reference data in Postgres directly.
The parameter→property picture the GP surrogate will eventually model (energy-density ladder 14–40 mJ/mm, monotonic in both):
- D638 tensile modulus: 365 MPa (profile A) → 2815 MPa (profile H) — TDS target 2800 MPa
- D790 flex modulus (corrected): 285 MPa (C) → 2272 MPa (H) — TDS target 2400 MPa
- Fuse-printed vertical controls: 2599 / 1950 MPa
See also: Data Plane · Operations for the post-print refresh sequence