Skip to content

Datasets

Peter Pak edited this page Jul 16, 2026 · 8 revisions

Datasets

The publication layer: five HuggingFace dataset repos, all git submodules under datasets/ in the main repo. They are downstream of Postgres (or their own lab/firmware sources) — never the agent's query surface.

✅ Migration executed 2026-07-16 (see Source Based Dataset Layout): exporters and ETLs are re-pointed — sls-export → Telemetry source/, linkage CSV → Database source/, conversations written in place. data/exports/ survives only as a fallback copy until the frames backfill completes.

Dataset Contents
Agentic-SLS-Telemetry Raw per-build telemetry parquet (566 GB)
Agentic-SLS-Database Firmware jobs, print profiles, PrintSessions
Agentic-SLS-ASTM D638 tensile / D790 flex specimen results (108 SLS specimens, all with job FKs)
Agentic-SLS-Knowledge SLS knowledge broadly (Inova + general: TDS sheets, machine data) and the agent-facing reference corpus served by reference_* (source/*.md + data/knowledge.jsonl); unlike the big datasets it must ALWAYS be initialized — the MCP tools read it at runtime. Phase-4 playbook will live here.
Agentic-SLS-Conversations Agent transcripts — raw session files are written directly into its source/sessions/ by the harness store (it is the archive; commit it often)

Submodule guards — important

  • The four data-heavy repos have update = none in .gitmodules so a blanket --recurse-submodules skips them (Telemetry alone is 566 GB). Never remove that guard; init them explicitly. Agentic-SLS-Knowledge is the deliberate exception — no update = none, it must always be initialized because the reference_* tools serve from it at runtime.
  • The git@hf.co: remotes need the HuggingFace SSH key.
  • Conversations' HF remote slug is Agentic-SLS-Conversations (typo) — intentional and baked in; redirects work; don't fix it.

The export contract

HF repos consume flat-file exports from the main repo's data/exports/ (JSONL + parquet + sidecar CSVs) via relative paths — never the live Postgres. Why: dataset ETL needs clean, container-free, idempotent inputs; hitting the live DB fails when it's down and competes with the live writer.

  • Producer: uv run sls-export. Idempotency contract: existing parquet preserved unless --force; sidecar CSVs (build_to_inova_session.csv, sensors.csv) preserve hand-filled values across re-runs.
  • Conversations exporter: uv run sls-export-conversations.
  • Sidecar CSVs are the pattern for any mapping that needs human authoring — checked in, authoritative. Don't add recorder-DB columns just to serve a downstream HF consumer; push the mapping into data/exports/.
  • Path conventions: Telemetry's and Conversations' scripts/_lib.py climb two parents to the repo root; ASTM finds Database as a sibling inside datasets/.

Reference data flows the other way too

Postgres holds synced copies of the Database and ASTM repos' data (inova_jobs, inova_print_profiles, inova_print_sessions, astm_specimens). Corrections happen at the origin repo, then uv run sls-sync-reference re-syncs. Never fix reference data in Postgres directly.

Measured signal so far

The parameter→property picture the GP surrogate will eventually model (energy-density ladder 14–40 mJ/mm, monotonic in both):

  • D638 tensile modulus: 365 MPa (profile A) → 2815 MPa (profile H) — TDS target 2800 MPa
  • D790 flex modulus (corrected): 285 MPa (C) → 2272 MPa (H) — TDS target 2400 MPa
  • Fuse-printed vertical controls: 2599 / 1950 MPa

See also: Data Plane · Operations for the post-print refresh sequence

Clone this wiki locally