Skip to content

Feature/dashboard live - #22

Merged
aam9063 merged 23 commits into
devfrom
feature/dashboard-live
Sep 28, 2026
Merged

aam9063 merged 23 commits into
devfrom
feature/dashboard-live

Conversation

@aam9063

@aam9063 aam9063 commented Sep 28, 2026

Copy link
Copy Markdown
Owner

No description provided.

…rompt v4

Three gaps, one of which blocked the dashboard's decisions screen.

- The orchestrator never persisted interpretations: the interpretation table had
  no writer in the whole app, so the Agent decisions screen had no real data
  source. One row per LLM interpretation now records intent, confidence, the
  structured fields, model, prompt version, tokens, latency and cost. The parser
  path writes nothing (a degraded decision is not an LLM decision), a duplicate
  provider message writes no second row, and a persistence failure is logged
  without breaking the rescue. _persist_inbound now returns the message id.
- OFFER_WITHDRAW was not inferable: the six withdrawal golden rows carried the
  same context as the decline rows, so the state that makes them withdrawals is
  now sent as accepted_offers instead of being guessed from the words 'al final'.
  The fixture gained that key for those rows only, and no label changed.
- interpreter_v4 fixes the prompt's own example: every version since v1 claimed
  'hasta mediodía puedo' fills proposed_start 07:00 as well, while the labels
  expect start=null; the model followed the example and lost 5 of 20 extractions.
  v4 states the rule the data encodes (one boundary fills one field).

Measured with the real provider on 150 golden samples:
  v2 0.9333 / 1.0 / 0.85    v3 0.9867 / 1.0 / 0.75 (times violated)
  v4 0.9933 / 1.0 / 1.0     one failure left: 'cancele el caso de todos'

Also fixes the CI-only test failure: build_runtime(settings) ignored its own
settings for the database and fell back to the ambient .env, so the test
connected to a developer's local Postgres and failed on a clean runner. It now
passes the URL explicitly and the test uses a temp-file database; verified with
the Linux CI image.

317 tests pass, ruff and mypy clean.
The dashboard was a complete UI with no HTTP client: every screen rendered mock
data and the backend exposed only /health, /api/status and the Twilio webhooks.

- Real login: Argon2 hashes for the seeded demo managers, HS256 JWT, and
  role-gated dependencies (401 without a valid token, 403 for the operator-only
  interpretation endpoints). The seeded placeholder hash authenticates nothing.
- Read endpoints for the whole dashboard slice, with response shapes that match
  frontend/src/domain/types.ts field for field (asserted by contract tests, so a
  renamed field breaks the build): locations, shifts, settings, rescues (list and
  detail with timeline, offers and candidates), approvals, conversations and
  their redacted messages, interpretations with filters, and metrics (LLM cost
  per day, p50/p95 latency, low-confidence rate, delivery failures, stuck
  rescues).
- Write endpoints go through the worker: approvals and rescue close enqueue a
  Celery task and answer 202, so the API process still never builds a runtime or
  owns domain logic.
- CORS restricted to the configured dashboard origins.
- Redaction holds: only stored redacted bodies are served and no health detail
  appears in any response.

Verified live against the real stack: login 200, 401 without a token, 403 for a
manager on /api/interpretations, operator 200, wrong password 401, and the full
chain webhook (32 ms) -> worker -> LLM -> persisted interpretation -> served by
the API (OFFER_CONDITIONAL, 0.9, gpt-4o-mini, $0.00038).

375 tests pass, ruff and mypy clean.
The dashboard was a complete UI rendering mock data with no HTTP client. It now
authenticates against the real API and reads the real database, using the seam
that already existed (DashboardDataSource + hooks) instead of rewriting screens.

- Session client (localStorage), an HTTP client that attaches the bearer token
  and turns failures into a typed ApiError/NetworkError, and a 401 path that
  clears the session and returns to the login screen.
- ApiDashboardDataSource: shifts, active rescues, rescue detail, approvals
  (enqueue + invalidate, since decisions are applied by the worker), conversations
  and messages, interpretations, metrics and settings. The location id is
  resolved from GET /api/locations instead of the hardcoded literal that never
  matched the seeded database.
- Login screen and RequireAuth guard, with the demo credentials shown for the
  reviewer, distinct messages for wrong credentials and for an offline API, the
  signed-in manager in the header with a logout action, and deep links preserved
  across login.
- getDataSource() keeps the mock available behind VITE_USE_MOCK (offline demo and
  hermetic tests). The Evals screen still reads mock data and now says so in the
  UI, because /api/evals/runs* was deliberately deferred.
- Dev proxy /api -> localhost:8000 so development needs no CORS, documented
  environment variables, and runbook instructions for logging in.
- Fixes a real defect found on the way: the settings form seeded its draft once,
  so live values arriving later could be overwritten by a save with stale
  defaults.

Verified: 119 frontend tests, build, oxlint and tsc clean; the SPA serves and the
dev proxy reaches the API (401 without a token, as expected); CORS accepts the
configured dashboard origin and rejects others.
Reproduced live: the worker crashed with 'got Future attached to a different
loop' and lost the inbound message. asyncio.run() creates a new loop per task
while the memoized async SQLAlchemy engine holds asyncpg connections bound to the
loop that created them, so a reused pooled connection fails. It passed earlier
only because the pool was still empty, which is exactly why it survived the
suite: intermittent, state-dependent, and invisible until the pool had idle
connections from a dead loop.

app/workers/async_runner.py now owns the rule (one lazily created loop per
process, reused via run_until_complete, recreated only when closed) and every
worker task goes through it. The regression test reproduces the asyncpg binding
mechanics and fails with the old behaviour.

380 tests pass, ruff and mypy clean.
…mation is pending

Two defects found by running the absence flow end to end with the LLM enabled.

- The ABSENCE_REPORTED audit event was written with a synthetic
  'case_<shift>_<timestamp>' id while the case itself got a uuid id, so the
  dashboard's timeline query joined on a case that does not exist and every
  rescue timeline rendered empty. It now records the case it was written for.
- The interpreter was never told that a confirmation was pending, while the
  prompt branches on exactly that: with nothing pending, a bare 'sí' is UNCLEAR
  by design, so a real employee answering SÍ got 'no te he entendido' and the
  absence was never confirmed. The orchestrator now sends pending_confirmation
  (the shift of the OPEN case awaiting confirmation).

The eval suite could not catch the second one: the golden fixture supplies
pending_confirmation, so it measured a context production never produced. A
contract test now pins the context keys the orchestrator sends, and an
end-to-end test drives report -> confirmation -> OFFERING through a fake that
mimics the prompt's procedure.

Verified live: the case advanced OPEN -> OFFERING with 3 offers and the timeline
now shows ABSENCE_REPORTED, RESCUE_OPENED and three OFFER_SENT events.

383 tests pass, ruff and mypy clean.
Timers never fired in the deployed stack, so no rescue ever escalated: two cases
were 20 minutes past their deadline and /scheduler_ran_jobs had never once
appeared in the worker log.

Root cause: the in-memory SimScheduler is per process and the worker runs Celery's
prefork pool (12 children). The task that opens a case registers the deadline in
its own scheduler, while the beat tick lands in whichever child Celery picks, with
its own empty queue. They practically never coincide.

- CeleryScheduler implements the same port as SimScheduler: schedule() becomes one
  deferred Celery task with countdown from the injected clock, so Redis holds the
  timer and any free worker executes it at the right time. Restarting a worker no
  longer loses pending timers, which removes a limitation the project had
  documented as accepted.
- apply_scheduled_job resolves the handler in the executing process; an unknown
  name fails loudly. SimScheduler stays the backend for the eval harness and
  tests, selected by settings.scheduler_backend or an injected scheduler.
- reconcile_stale_cases (beat, every 60s) re-enqueues the deadline of cases that
  still expect action, so a timer lost between a commit and an enqueue recovers
  itself instead of hanging forever.
- Today board: shifts in state 'scheduled' (staffed, untouched) now appear in the
  covered column instead of nowhere, so 'Covered today' no longer reads 0 while
  nine shifts are staffed.

And the defect the fix immediately exposed: with timers firing for the first time,
escalation failed with

  StringDataRightTruncationError: value too long for type character varying(64)
  ('audit_case_bc1809c0..._escalated_1790359021_DEADLINE_REACHED', ...)

because _escalate composed an 81-character audit id and the transaction rolled
back every time, so no rescue had ever escalated in production. The quiet-hours
marker had the same shape at exactly 64 characters. Both use opaque audit_<uuid>
ids now, and the existing width guard — which had only ever exercised the
confirmation path — drives the escalation and deferred-wave paths too.

Verified live: a timer scheduled 40s out appeared in 'celery inspect scheduled'
with its ETA and executed at exactly that second; two overdue cases escalated; the
repeated reconcile sweep changed nothing.

402 tests pass, ruff and mypy clean, frontend 119 tests, build and lint clean.
The dashboard could watch the agent work but not produce the work: demoing the
core flow needed a second phone or ten minutes of waiting for a wave timeout.

- POST /dev/simulator/{employee_id}/messages resolves the employee and enqueues
  the SAME task the Twilio webhook enqueues, so a simulated message is
  indistinguishable in the database from a real one except for its synthetic
  provider id: idempotency, threading, interpretation, auditing and delivery all
  behave identically.
- POST /dev/clock/advance moves a Redis-backed offset every worker process reads
  (deadlines, quiet hours, eligibility) and immediately enqueues the reconcile
  sweep, so an overdue case escalates in seconds instead of waiting for the beat
  tick. GET /dev/clock exposes the virtual time.
- GET /api/employees gives the screen the roster it needs; the Simulator screen
  now shows real employees, their real threads, a send action and the demo clock,
  keeping its phone-frame layout, with the mock still behind VITE_USE_MOCK.
- Both /dev routes are registered only when APP_ENV is local/test/demo and even
  then require a manager JWT: a production deployment does not advertise them.

Verified live: a message sent through the simulator opened a real rescue with the
right shift and offers, and advancing the clock escalated it.

428 tests pass, ruff and mypy clean, frontend 124 tests, build, lint, tsc clean.
Three defects of one family, found by driving the product end to end: the agent
asked something and the answer was ignored.

- 'Which shift?' was a dead end. When an employee has two or more shifts in the
  48-hour window — the common case — the agent asked and nothing consumed the
  reply, so the absence was never reported. Verified live: an employee with a
  shift in progress could not report an absence at all. The orchestrator now
  sends the candidate shifts and, once it has asked, resolves the answer: the
  model returns the chosen shift id, and the deterministic parser matches the
  same candidates in degraded mode. It never guesses, asks once more, then
  redirects.
- The interpreter never received shifts_48h, although the prompt and the golden
  fixture both used it: third instance of the fixture describing a context
  production never produced. Candidates now carry their date and the context
  carries today's date, so 'el de hoy' is resolvable. A context-contract test
  pins the exact key set for the four relevant states.
- An absence that is never confirmed hung forever: the state machine had no
  OPEN + DEADLINE_REACHED transition, so the manager was never told. It now
  escalates and notifies (the product decision), with a new eval scenario and an
  INV6 relaxation that keeps requiring ABSENCE_REPORTED + ESCALATED for that path.

interpreter_v5 adds the shift-choice branch (edited in place: v5 was created in
this same change and never shipped, so inflating a version would be dishonest);
golden gains five rows for the new branch, with only those rows touched.

Measured with the real provider on 155 samples: accuracy 0.9935, health 1.0,
conditional times 0.95 — thresholds met, one failure left.

Verified live: an employee with two upcoming shifts reported an absence and the
case opened for the shift they named, and an unconfirmed absence escalated with a
real WhatsApp to the manager (Twilio answered 201).

435 tests pass, ruff and mypy clean, 27 eval tests green.
DESIGN.md §8 already defines the responsive contract (breakpoints, collapsing
strategy, 44px touch targets, gutter scale); the dashboard did not honour it.

- The header now has a hamburger drawer for phones and tablets. It is lg:hidden,
  not md:hidden: between 768 and 1023 px the main nav is visible but the operator
  group is not, so a md-only drawer left Ops, Agent decisions and Evals
  unreachable on a tablet. The panel lists both groups with the gold active
  indicator, closes on Escape and on selection, keeps the manager and logout
  reachable, and its visible labels are English (they were PRINCIPAL/OPERADOR).
- The two wide tables render as stacked cards below md and keep the table from md
  up in an overflow-x-auto container; every touch control meets 44px through
  pointer-coarse (the wordmark, the demo pill and Log out were 28/36/20px).
- Settings and Evals grids reflow, the chart scales to its container, and the
  screens were audited at 360px for overflow and touch targets.
- Agent decisions shows a plain-English explanation when the operator-only
  endpoint answers 403 instead of rendering a blank page that looks broken.

Verified in a real browser (Playwright + Chromium installed outside the project,
hasTouch on so pointer:coarse matches) at 360, 768 and 1440 px: no horizontal
overflow, every destination reachable, every visible control >= 44px, and the
screens reviewed by eye. The sweep for Spanish UI copy now also flags accented
literals, which is what caught the Spanish group labels.

141 frontend tests, build, lint and tsc clean.
Two things a manager saw in the browser console and in the table.

- The network tab showed four 403s for one screen. A 403 (wrong role) answers the
  same however many times it is asked, so TanStack Query's default three retries
  only produced three extra failing requests and three extra red rows. The retry
  policy now skips 4xx and keeps two attempts for network and 5xx failures, with
  mutations never retried.
- The Agent decisions table rendered the raw API timestamp
  ('2026-09-25T19:27:33.483410+00:00') because the mock had always carried an
  already formatted '15:03'. The conversion moved to domain/decisions.ts with its
  own tests (ISO to HH:MM, an already compact value passes through, an
  unparseable one is shown rather than hidden).

Verified in the browser: the manager now issues one request instead of four and
still sees the explanation for the operator-only view; the operator account sees
the real rows with local times (21:27).

144 frontend tests, build, lint and tsc clean.
…scenario real

Two defects the simulator screen exposed.

- The demo clock read '--:--' because the SPA never reached the API: the Vite dev
  proxy only covered /api, so /dev/* hit the SPA fallback and returned HTML, and
  the deployed Caddyfile did not route /dev/* either — the simulator would have
  been broken in the demo environment too. Both now forward /dev to FastAPI; the
  endpoint is still demo-gated by APP_ENV, so production keeps answering 404.
- 'Load scenario: acceptance race' was a button with no handler. The offer
  payload now carries employeeId, and the control runs the scenario for real: it
  finds the OFFERING rescue with at least two pending offers, fires both
  acceptances concurrently (a sequential pair would not race), refreshes the
  affected queries, and says what to look for. With no eligible rescue it
  explains itself instead of doing nothing.

Verified live end to end: the clock shows the virtual time (21:32), and clicking
the scenario covered the shift with exactly one winner — Marta Lopez ACCEPTED,
the other two offers CANCELLED, case COVERED, shift covered — which is the
concurrency invariant the eval scenario asserts.

Also confirmed while checking the outcome: notifications sent outside a
conversation (the loser's 'already covered', the manager's escalation and
coverage notices) are delivered but never persisted, so the dashboard shows no
trace of them. Recorded as a follow-up.

148 frontend tests, 435 backend tests, ruff, mypy, build, lint and tsc clean.
…oard table

Three things the demo exposed, plus a fidelity fix.

- The Simulator now says what the agent will do: each frame labels the employee
  (on shift now / starts at HH:MM / ended at HH:MM / no shift today), the frames
  that can actually report an absence come first, and when nobody can, the screen
  says why and what to do. Writing into a frame whose shift already ended could
  only ever produce the out-of-scope reply, which made the screen look broken
  while the agent was right.
- The demo clock is reversible and honest: POST /dev/clock/reset (demo-gated,
  JWT) zeroes the offset and enqueues the reconcile sweep, and the control shows
  the virtual time, the offset in plain words and a Reset. A leftover offset had
  reached +18h50m from testing, which moved the worker's 'now' a day ahead: today's
  shifts read as finished, the agent answered out of scope, and stored timestamps
  jumped into the future.
- Today is now a dashboard table (ROLE / WINDOW / EMPLOYEE / STATUS / RESCUE /
  ACTIONS) with the existing tokens, responsive as the other wide screens are.
- Fidelity fix found in the new table: escalated rescues read as 'Uncovered — no
  candidates offered yet', which was worse than the board. A shift whose latest
  case escalated now shows an Escalated badge with the honest caption (the manager
  was notified), a covered shift never reads as escalated, and any shift with a
  case — terminal included — offers View detail.

159 frontend tests, 437 backend tests, ruff, mypy, build, lint and tsc clean.
# Conflicts:
#	backend/app/runtime.py
#	backend/tests/unit/test_runtime.py
…fer text

Two defects the live verification exposed, both in what the agent says.

- Every message the agent cannot act on landed in one generic line: solo
…reload

The user drove the product and could not use it. Five defects, all verified in a
browser walkthrough.

- The Simulator looked broken: the agent takes 10-15 s to answer and nothing
  refreshed, so the reply only appeared after a manual reload. The thread now
  polls after a send (every 3 s, bounded to ~30 s, stopping on the reply, on the
  bound and on unmount), says "the agent is replying…" while waiting and blocks
  duplicate sends. Verified: the confirmation question now lands on its own.
- The manager had nothing to do on an escalated case: the detail screen showed
  the timeline and the candidates with only navigation buttons, while the API
  already exposed the close. It now has Close case (confirmed, queued, 202) and,
  for AWAITING_APPROVAL, approve/reject inline.
- An escalated shift read "Uncovered — No candidates offered yet" on Today,
  because the board only received active cases and the row fell back to "nothing
  was tried". The board now gets the day's cases including terminal ones and
  shows Escalated with "Escalated at HH:MM" and a summary of who was contacted
  and what each answered (spec §6.4), never a countdown on a terminal case.
- docs/demo-script.md is now a walkthrough a newcomer can follow: what each
  screen is for, the exact Simulator steps (including the ~10 s wait), the state
  table, what the manager does on escalation, and the late-acceptance beat.

Deliberately NOT done: cancelling the open offers when a case escalates. That was
my own wrong call in the task document — spec §5 says a late acceptance after the
escalation becomes AWAITING_APPROVAL for the manager, so the offers must stay
open. Reverted with the test that pins it.

175 frontend tests, 444 backend tests, ruff, mypy, oxlint, build and tsc clean.
… reply

The user ran the whole flow and got something confusing: they reported an absence,
answered the confirmation 39 minutes later, and the system produced an approval
request instead of confirming. Three defects behind it.

- **Stale offers were still acceptable.** A candidate's `PENDING` offer whose
  shift had already ended could be "accepted": the live case shows offers from
  shifts three days old being answered, and one of them created a
  `late_acceptance` approval for a finished shift. `_try_accept_offer` now checks
  that the shift has not ended before doing anything with the offer; a stale one
  is cancelled with an `OFFER_STALE` audit and the candidate gets the honest
  close-out. A late acceptance for a shift that is *still ahead* keeps going to
  the manager, as spec §5 requires.
- **A late answer got a generic line, or none at all.** A recent closed case now
  gets its own outcome message (`state_case_escalated`, `state_case_covered`,
  `state_case_closed`) instead of "solo gestiono avisos de ausencia", which was
  the opposite of the truth. Cases older than the window (6 h) stay out of it.
- **Not every reply was stored**, so the dashboard showed the employee's message
  with no answer — the reason the flow looked broken. Every message to an
  employee now lands in that employee's conversation, and the wave offers use the
  same phone-based thread the WhatsApp webhook uses, so there is one thread per
  employee instead of two.

Verified live end to end: an employee answering "SÍ" after their case was closed
now receives "tu aviso del turno de 09:00 a 17:00 quedó cerrado por tu encargado"
and the reply is visible in the conversation.

447 tests pass, ruff and mypy clean.
The user ran the flow and got lost because the screen carried three days of test
leftovers: a stale "[template: offer]" bubble, threads where the agent answered
"no te he entendido" to a question about a shift that had already closed, and a
demo clock left ten hours ahead. None of that was the product misbehaving, but it
is exactly what makes a demo unreadable.

POST /dev/demo/reset (demo-gated, manager JWT) removes the operational artifacts
in foreign-key order (interpretation, message, conversation, approval_request,
offer, audit_event, rescue_case), drops the shifts from today onwards, reseeds the
schedule and returns the demo clock to real time through the same path
POST /dev/clock/reset already used. It reports what it removed, so the UI can be
honest about it.

The Simulator gets a Reset demo data control next to the clock, with a
confirmation that names what disappears and says it cannot be undone, plus the
result line. Verified live against real Postgres: dirtying the environment and
resetting from the UI printed "Deleted 2 interpretations, 6 messages, 4
conversations, 3 offers, 5 audit events, 1 rescues, 143 shifts. demo clock back on
real time." and left the board with every shift Scheduled and no case.

docs/demo-script.md starts with that step now, because leftovers from a previous
run are the main source of confusion.

451 backend tests, 179 frontend tests, ruff, mypy, oxlint, build and tsc clean.
The Simulator rendered `roster.slice(0, 3)`, a leftover from the mock that had
three phones. Because the roster sorts the actionable employees first, the three
visible frames were always the ones already on shift: every employee whose shift
starts later — the calmest way to run the demo, since an in-progress shift gives
only a ten-minute confirmation window — was invisible. The user reported exactly
that ("no hay ninguno que ponga Starts at") while four such frames existed.

The whole roster renders now (cards below md, three across from xl), with the
actionable frames first and the employees without a shift today last. Verified in
a browser at real time: 3 frames say "On shift now" and 4 say "Starts at HH:MM"
(Lucía 13:00, Elena and Iker 15:00, Tomás 18:00).

179 frontend tests, build, lint and tsc clean.
…to miss

The demo clock had been clicked to +40 hours. With the clock that far ahead every
shift of the day reads as finished: the Simulator showed "04:01 (+20 h ahead)",
every frame said "Ended at HH:MM", and the banner admitted nobody could report an
absence. Nothing stopped the drift and nothing made it obvious.

- POST /dev/clock/advance clamps the total offset to a documented ±6 hours and
  answers with the applied offset plus a `clamped` flag, so the drift stops here.
  The old `incrby` was atomic but could not clamp; the read-modify-write is fine
  for the single-operator demo use case and the review notes it.
- From one hour in either direction the small grey note becomes a prominent
  alert that states the consequence plainly — the shift windows are not where
  they look, so people who should be on shift read as finished and vice versa —
  with a Reset clock action inside it.
- The wire type and the hook carry the applied offset and the flag, and the
  screen says when a request was cut.

Verified live: a +20 h request is clamped to 6 h with clamped: true, a second one
stays at 6 h, a normal +1 h is untouched, and at +1 h the banner appears with its
Reset action.

455 backend tests, 182 frontend tests, ruff, mypy, oxlint, build and tsc clean.
The user reported that after writing in the Simulator they had to reload the page
to see the agent's answer — and they were right: the reply poll only started when
the employee already had a conversation.

    onSuccess: (_sid, variables) => {
      if (variables.conversationId) {        // null on first contact
        startReplyPoll(variables.conversationId)

`GET /api/employees` returned the conversation id only when a Conversation row
existed, so after a demo reset every frame was in that state: the poll never
started and the answer only appeared on a manual reload. With a previous thread it
worked, which is why it looked fine in earlier checks.

A thread now exists conceptually from the start: the roster always advertises the
deterministic id the orchestrator already uses (`conv_twilio_<phone>`), and
`GET /api/conversations/{id}/messages` answers an empty list for a real
employee's empty thread instead of a 404 (a normal state, not an error — unknown
ids and numbers that do not exist keep their 404). The frames show an honest
empty state ("No messages yet — write the first one.") and the simulator's 202
carries the conversation id so the poll has somewhere to look immediately.

Verified live on a reset database, the exact reported situation: a frame with no
conversation received the message, showed "the agent is replying…", and the
confirmation question appeared **5 seconds later with no reload**.

457 backend tests, 183 frontend tests, ruff, mypy, oxlint, build and tsc clean.
…hboard

The Evals screen was the last one rendering invented numbers, and it admitted it:
"eval data is not live yet". Behind it, the `eval_run` table had the columns for
exactly this (git sha, trigger, timings, model config, prompt versions, metrics,
invariant violations, passed, baseline, report path) and **zero rows** — no writer
ever existed, and `/api/evals/runs*` was never implemented. Meanwhile the real
evaluation ran from the terminal and its results went nowhere.

- `app/evals/recording.py` persists one row per run: the golden set (metrics,
  verdict from the threshold gate, model, prompt version, report path) and the
  scenario suite (per-scenario pass/fail plus the invariant violations).
  Recording is best effort: it can never break a run, and a failing run is
  recorded as failing rather than dropped.
- `evals/runner.py` records after each run and gains `--scenarios` to run and
  record the scenario suite. Two bugs surfaced while wiring it: the harness
  reports carry `datetime` values that JSON columns reject (normalized now), and
  `Settings` looked for `.env` in the current directory, so a runner started from
  the repository root silently used the default database port instead of the
  configured one — it now resolves `backend/.env` from the module.
- Three operator-only endpoints (`/api/evals/runs`, `/runs/{id}`, `/runs/summary`)
  serve the screen's shape. The threshold travels **inside the run**: the API
  container does not carry `evals/thresholds.yaml`, and a stored verdict should be
  self-explanatory. The verdict describes the newest run that was actually judged,
  so the informational offline baseline never reads as a product failure.
- The screen renders the real summary (verdict, commit, accuracy history with the
  threshold line, model comparison, invariants, per-scenario results) and shows an
  honest empty state naming the command that produces a run.

Verified live with recorded runs: 15/15 scenarios, 0 invariant violations, and the
real model at 0.9871 intent accuracy against the 0.92 gate — shown in the
dashboard, no mock note left.

469 backend tests, 188 frontend tests, ruff, mypy, oxlint, build and tsc clean.
@gitguardian

gitguardian Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

⚠️ GitGuardian has uncovered 2 secrets following the scan of your pull request.

Please consider investigating the findings and remediating the incidents. Failure to do so may lead to compromising the associated services or software components.

🔎 Detected hardcoded secrets in your pull request
GitGuardian id GitGuardian status Secret Commit Filename
37629664 Triggered Generic Password 6f417c8 frontend/src/screens/LoginScreen.test.tsx View secret
37629664 Triggered Generic Password 05a7c8b frontend/src/domain/demoCredentials.ts View secret
🛠 Guidelines to remediate hardcoded secrets
  1. Understand the implications of revoking this secret by investigating where it is used in your code.
  2. Replace and store your secrets safely. Learn here the best practices.
  3. Revoke and rotate these secrets.
  4. If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.

To avoid such incidents in the future consider


🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.

…eclare the demo password once

CI failed on five tests that passed locally, which is the signature of an
ambient dependency.

- `tests/unit/api/test_dev_tools.py`: the `/dev/clock/advance` tests did not use
  the `stub_sweep` fixture, so they enqueued the reconcile sweep for real. On a
  development machine Redis is up and they passed; in CI Celery retried the
  result store for four minutes and the five tests failed with "Retry limit
  exceeded while trying to reconnect to the Celery result store". They stub the
  enqueue now, and `tests/unit/api/conftest.py` adds an autouse fixture that
  records every enqueue for the whole unit API suite — no unit test can reach a
  broker by accident again. Verified with the Linux CI image: 469 passed, and the
  suite went from 225 s to ~105 s because nothing retries a missing Redis.
- The demo password was repeated across tests, docs and two frontend modules. It
  is meant to be public (a fictional venue with fictional data, shown on the
  login screen so a reviewer can get in), so it now has exactly two declared
  homes — `DEMO_PASSWORD` in the seed and in `frontend/src/domain/demoCredentials.ts`
  — and every test imports it instead of repeating the literal. The runbook keeps
  the documented credentials table (that is the point of a demo credential) and
  its curl example no longer embeds the value.

469 backend tests, 188 frontend tests, ruff, mypy, oxlint, build and tsc clean.
@aam9063
aam9063 merged commit ff96627 into dev Sep 28, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant