Skip to content

Operations

Peter Pak edited this page Jul 26, 2026 · 14 revisions

Operations

Deployment model, runbook sequences, and the gotchas that were learned the hard way. (The repo's CLAUDE.md is the compact agent-facing version of this page.)

Services

Systemd is fully retired (2026-07-16) — all six services run in docker compose, one folder per service under services/: postgres, pgadmin (:5050), web (:5173, vite dev, hot-reloads), broker (:3100, services/harness/ — harness CLIs pinned in-image: claude-code 2.1.187, codex 0.144.1, opencode 1.15.12, agy 1.1.1 bind-mounted; auth = bind-mounted host home dirs), recorder (:3000, services/recorder/ — restart: on-failure preserves stop-after-build semantics; never recreate while recording), defect (:3200, services/defect/ — notify-only bridge to the v05 defect model server on the MAIL-10 workstation; safe to restart anytime, losing MAIL-10 never affects recording). The GUI Services page carries a Harness Auth card (versions, last-session outcomes, guided login commands).

Recorder/broker run tsx without watch — services/recorder/src and services/harness/ edits do NOT hot-reload. A deploy is an explicit docker compose restart <recorder|broker> at a safe moment.

Service folders are included by the root docker-compose.yml. The root file pins the compose project name: — changing it makes compose recreate the live postgres container. Future compose citizens: vLLM for the fine-tuned model, embedding services. (Migration history: infra → compose first (2026-07-15), web 2026-07-16, then recorder/broker same day — systemd retired.)

The defect model itself runs REMOTELY on MAIL-10 (ssh 10): runs/v05/serve.py on :8100, GPU-backed, started via runs/v05/deploy.sh — prefix linuxbrew on PATH (export PATH=/home/linuxbrew/.linuxbrew/bin:$PATH); non-interactive SSH doesn't get it and nohup uv fails otherwise. Health: curl http://128.2.112.20:8100/health. The bridge POSTs per-layer frames to /infer and records verdicts as events rows (kind='defect_detection'); results are advisory (notify-only) and surface in the Mission Control attention bar.

Never restart the recorder while a print is being recorded. Check /api/health/recording first; for a graceful maintenance window use POST /api/admin/stop-after-build. Logs: docker logs -f agentic-sls-recorder.

The recorder lifecycle self-heals: startup orphan reconciliation, spool recovery (15 s delay so the job detector re-adopts an active build first), unreachable-printer finalization.

Post-print refresh sequence

  1. uv run sls-export — refresh Telemetry source/ + Database linkage CSV
  2. uv run sls-summarize-builds — rebuild build_summaries
  3. New PrintSessions? rsync ppak@inova:/home/ppak/SLS4All/PrintSessions/ into sls4all/SLS4All-Backup (submodule since 2026-07-16) + the Database repo, then uv run sls-sync-reference
  4. New builds need a sls-match-build-sessions propose/review pass (live session stamping not yet built — see Architecture)

Tests

  • Root: uv run pytest
  • plugin/: npm test + npm run typecheck + npx tsx tests/smoke.ts (needs DATABASE_URL exported)
  • Headless agent run: npx tsx services/harness/run.ts <harness> --prompt '…'

Printer-side deploys

See Inova API Plugin for the build.sh → install.sh → reboot loop. Destructive printer actions (reboot/kill/restart) are user-executed — agents recommend and verify only. One-shot firmware restarts default to sudo reboot over the firmware's soft-restart hooks.

Conventions

  • TypeScript over JavaScript for new JS-adjacent code; Python scripts run via uv run; CLI tools installed with brew.
  • Postgres: two-arg round() needs a ::numeric cast.
  • MCP knowledge tools are read-only by construction (forced read-only pool); agent_actions logging of every *_set is expected behavior.

Gotchas (hard-won)

  • Old-era telemetry (pre-2026-07-13-restore) lives only in data/exports/telemetry/*.parquet — summarize from parquet, never --source postgres for big builds.
  • Profile names can lie (a "20mJ/mm" profile containing 15) — read the JSON field.
  • Dependency isolation (2026-07-16): no root node_modules or npm workspaces — each node service has its own package.json + lockfile, deps baked into its image via npm ci, surfaced over the repo bind through an anonymous volume. Dep bumps: docker compose build <svc> && docker compose up -d -V <svc> (recorder: safe moments only). Root package.json is a scripts-only stub; npm run typecheck sweeps all four services — its FIRST full run caught three latent bugs, including the run.ts model-overwrite behind opencode's null model rows.
  • Repo unit files vs installed copies drift: deploy/systemd/*.service is the source; systemd runs the COPY in ~/.config/systemd/user/ — re-copy + daemon-reload on every unit edit (bit us 2026-07-16; compose services don't have this failure mode).
  • opencode quirks: project root = git toplevel (hence the root .opencode/ shim, the one plugin file left at repo root) and it MUST get an explicit --model — the free zen default (big-pickle) errors upstream as of 2026-07-16 and no paid provider is connected.
  • Sandboxed shells can't see listening ports — ss/curl under the sandbox falsely report healthy services as down; verify health with an unsandboxed curl.
  • The docker/mount boot race (root-caused 2026-07-15). /mnt/storage2 mounts with no hard ordering against docker; on the 2026-07-09 boot docker started first, auto-created the postgres bind path on the root filesystem, and postgres initdb'd a fresh cluster there — that was the mysterious "Jul 10 DB reset." All work Jul 10–14 (46-build restore, agent tables, build_summaries) lived in that hidden shadow until the next reboot flipped back to the stale HDD datadir. Recovered 2026-07-15 by rsyncing the shadow cluster back. Guard now in place: /etc/systemd/system/docker.service.d/wait-for-storage2.conf sets RequiresMountsFor=/mnt/storage2. If Postgres ever looks inexplicably stale: check WAL file mtimes in the datadir and look for a shadow under the mountpoint (sudo mount --bind / /mnt/rootfs-peek).
  • Wedged watchers (historical, from the pre-systemd tsx watch era): an orphaned process holding a port while a childless watcher ignores file edits. Diagnose with ss -tlnp and compare the owning PID against ps. The systemd migration removed the day-to-day exposure; services/recorder/scripts/run-recorder.sh remains for dev only (flock-guarded).
  • Worktree development during prints: for non-trivial recorder changes while recording, build in a git worktree and merge at a safe moment (worktrees need node_modules + submodule symlinks).
  • Storage tiering exists for a reason: Postgres lives on an SMR HDD that cannot absorb 10 Hz inserts — see Data Plane before "simplifying" the spool away.

Clone this wiki locally