-
Notifications
You must be signed in to change notification settings - Fork 0
Operations
Deployment model, runbook sequences, and the gotchas that were learned the hard way. (The repo's CLAUDE.md is the compact agent-facing version of this page.)
Systemd is fully retired (2026-07-16) — all six services run in docker
compose, one folder per service under services/: postgres, pgadmin
(:5050), web (:5173, vite dev, hot-reloads), broker (:3100,
services/harness/ — harness CLIs pinned in-image: claude-code 2.1.187,
codex 0.144.1, opencode 1.15.12, agy 1.1.1 bind-mounted; auth =
bind-mounted host home dirs), recorder (:3000, services/recorder/ —
restart: on-failure preserves stop-after-build semantics; never
recreate while recording), defect (:3200, services/defect/ — notify-only
bridge to the v05 defect model server on the MAIL-10 workstation; safe to
restart anytime, losing MAIL-10 never affects recording). The GUI Services
page carries a Harness Auth card (versions, last-session outcomes, guided
login commands).
Recorder/broker run tsx without watch — services/recorder/src and services/harness/ edits do NOT hot-reload. A deploy is an explicit docker compose restart <recorder|broker> at a safe moment.
Service folders are included by the root docker-compose.yml. The root file pins the compose project name: — changing it makes compose recreate the live postgres container. Future compose citizens: vLLM for the fine-tuned model, embedding services. (Migration history: infra → compose first (2026-07-15), web 2026-07-16, then recorder/broker same day — systemd retired.)
The defect model itself runs REMOTELY on MAIL-10 (ssh 10): runs/v05/serve.py on :8100, GPU-backed, started via runs/v05/deploy.sh — prefix linuxbrew on PATH (export PATH=/home/linuxbrew/.linuxbrew/bin:$PATH); non-interactive SSH doesn't get it and nohup uv fails otherwise. Health: curl http://128.2.112.20:8100/health. The bridge POSTs per-layer frames to /infer and records verdicts as events rows (kind='defect_detection'); results are advisory (notify-only) and surface in the Mission Control attention bar.
Never restart the recorder while a print is being recorded. Check /api/health/recording first; for a graceful maintenance window use POST /api/admin/stop-after-build. Logs: docker logs -f agentic-sls-recorder.
The recorder lifecycle self-heals: startup orphan reconciliation, spool recovery (15 s delay so the job detector re-adopts an active build first), unreachable-printer finalization.
-
uv run sls-export— refresh Telemetrysource/+ Database linkage CSV -
uv run sls-summarize-builds— rebuildbuild_summaries - New PrintSessions? rsync
ppak@inova:/home/ppak/SLS4All/PrintSessions/intosls4all/SLS4All-Backup(submodule since 2026-07-16) + the Database repo, thenuv run sls-sync-reference - New builds need a
sls-match-build-sessionspropose/review pass (live session stamping not yet built — see Architecture)
- Root:
uv run pytest -
plugin/:npm test+npm run typecheck+npx tsx tests/smoke.ts(needsDATABASE_URLexported) - Headless agent run:
npx tsx services/harness/run.ts <harness> --prompt '…'
See Inova API Plugin for the build.sh → install.sh → reboot loop. Destructive printer actions (reboot/kill/restart) are user-executed — agents recommend and verify only. One-shot firmware restarts default to sudo reboot over the firmware's soft-restart hooks.
- TypeScript over JavaScript for new JS-adjacent code; Python scripts run via
uv run; CLI tools installed with brew. - Postgres: two-arg
round()needs a::numericcast. - MCP knowledge tools are read-only by construction (forced read-only pool);
agent_actionslogging of every*_setis expected behavior.
-
Old-era telemetry (pre-2026-07-13-restore) lives only in
data/exports/telemetry/*.parquet— summarize from parquet, never--source postgresfor big builds. - Profile names can lie (a "20mJ/mm" profile containing 15) — read the JSON field.
-
Dependency isolation (2026-07-16): no root node_modules or npm workspaces — each node service has its own package.json + lockfile, deps baked into its image via
npm ci, surfaced over the repo bind through an anonymous volume. Dep bumps:docker compose build <svc> && docker compose up -d -V <svc>(recorder: safe moments only). Rootpackage.jsonis a scripts-only stub;npm run typechecksweeps all four services — its FIRST full run caught three latent bugs, including the run.ts model-overwrite behind opencode's null model rows. -
Repo unit files vs installed copies drift:
deploy/systemd/*.serviceis the source; systemd runs the COPY in~/.config/systemd/user/— re-copy + daemon-reload on every unit edit (bit us 2026-07-16; compose services don't have this failure mode). -
opencode quirks: project root = git toplevel (hence the root
.opencode/shim, the one plugin file left at repo root) and it MUST get an explicit--model— the free zen default (big-pickle) errors upstream as of 2026-07-16 and no paid provider is connected. -
Sandboxed shells can't see listening ports —
ss/curlunder the sandbox falsely report healthy services as down; verify health with an unsandboxed curl. -
The docker/mount boot race (root-caused 2026-07-15).
/mnt/storage2mounts with no hard ordering against docker; on the 2026-07-09 boot docker started first, auto-created the postgres bind path on the root filesystem, and postgres initdb'd a fresh cluster there — that was the mysterious "Jul 10 DB reset." All work Jul 10–14 (46-build restore, agent tables,build_summaries) lived in that hidden shadow until the next reboot flipped back to the stale HDD datadir. Recovered 2026-07-15 by rsyncing the shadow cluster back. Guard now in place:/etc/systemd/system/docker.service.d/wait-for-storage2.confsetsRequiresMountsFor=/mnt/storage2. If Postgres ever looks inexplicably stale: check WAL file mtimes in the datadir and look for a shadow under the mountpoint (sudo mount --bind / /mnt/rootfs-peek). -
Wedged watchers (historical, from the pre-systemd
tsx watchera): an orphaned process holding a port while a childless watcher ignores file edits. Diagnose withss -tlnpand compare the owning PID againstps. The systemd migration removed the day-to-day exposure;services/recorder/scripts/run-recorder.shremains for dev only (flock-guarded). - Worktree development during prints: for non-trivial recorder changes while recording, build in a git worktree and merge at a safe moment (worktrees need node_modules + submodule symlinks).
- Storage tiering exists for a reason: Postgres lives on an SMR HDD that cannot absorb 10 Hz inserts — see Data Plane before "simplifying" the spool away.