Skip to content

Dev - #23

Merged
aam9063 merged 36 commits into
mainfrom
dev
Sep 28, 2026
Merged

Dev#23
aam9063 merged 36 commits into
mainfrom
dev

Conversation

@aam9063

@aam9063 aam9063 commented Sep 28, 2026

Copy link
Copy Markdown
Owner

No description provided.

aam9063 and others added 30 commits September 25, 2026 12:22
…time

The LLM layer was implemented and tested but never wired: get_twilio_service()
built the orchestrator without an interpreter, so every live WhatsApp message
was answered by the deterministic parser, and no TracerProvider was ever
installed so nothing reached Langfuse.

- Settings now declares the LLM and tracing configuration (llm_provider,
  model, temperature, tokens, timeout, confidence threshold, prices, provider
  credentials, Langfuse keys) and derives llm_enabled, traces_endpoint,
  tracing_enabled and traces_auth_header. It stays the only reader of the
  environment; the legacy os.getenv helper is gone.
- New app/agent/factory.py builds the Strands model per provider (openai
  default, anthropic, bedrock; OpenAI-compatible gateways via OPENAI_BASE_URL)
  and build_interpreter() fails closed: disabled provider, missing credential
  or missing SDK returns None plus one llm_disabled warning, so a rescue is
  never blocked by LLM configuration.
- configure_tracing() installs an idempotent OTLP/HTTP exporter to Langfuse
  Cloud with Basic auth and injectable exporter/provider for hermetic tests;
  the FastAPI lifespan configures it and flushes on shutdown.
- get_twilio_service() injects the interpreter and logs the active provider;
  the eval runner uses the same factory, so evals and production resolve the
  provider identically.
- Adds openai and the opentelemetry api/sdk/otlp-http dependencies.
… record

- ADR-004 records why OpenAI is the first provider, why OpenAI-compatible
  gateways need no separate code path, and why failure is fail-closed to the
  deterministic parser instead of crashing at boot.
- Runbook gains a configuration section: how to set and verify the provider and
  the Langfuse keys, what the boot log must say, and the incident rows for a
  parser-like answer, unexpected cost and missing traces.
- assumptions.md adds A5 (OpenAI as the first provider) and closes A4 with the
  fact that Langfuse Cloud is now wired code rather than intent.
… log

- scripts/verify_langfuse.py sends one real span through the configured OTLP
  exporter and reads it back from Langfuse's v2 observations API, so key
  problems are diagnosable without deploying (the legacy /api/public/traces
  endpoint answers 410 for organizations created after 2026-09-16).
- The startup log field is 'detail' instead of a 'provider' key holding a full
  'provider=... model=...' string.
Test counts, fail-closed behaviour, the live startup log, the confirmed
Langfuse Cloud round trip and the .env correction, so the feature record shows
what was observed rather than asserted.
Both were invisible to the unit suite because the doubles were more forgiving
than the SDK.

- interpreter: the real client returns the full Interpretation dump, which
  already carries prompt_version, so passing it again raised TypeError and
  every message degraded to the parser through ProviderUnavailableError. Merge
  the payload instead, with a regression test that feeds a full dump.
- llm metering: Strands 1.56 reports EventLoopMetrics.accumulated_usage in
  camelCase and latency in accumulated_metrics['latencyMs'], so the client was
  recording 0 tokens and $0 for every call. Extract tolerantly across versions,
  fall back to wall-clock latency, and keep cached-input tokens visible.

Adds scripts/verify_llm.py (real provider smoke check) and records the evidence
including the latency observation against the advertised p95 target.
The 1.2 s figure in the Ops screen was a mockup number that was never measured
and that the real provider exceeds (1.0-2.8 s per call, including the Strands
agent cycle). The target moves to 2.5 s and the demo p95 shown next to it
becomes 1980 ms, consistent with what the live path reports; the eval harness
already allowed 5 s. The system prompt stays untouched because ~1024 of its
~1150 input tokens are prompt-cache reads, so trimming it would cost
interpretation quality for latency we no longer need.

ADR-004 records the decision and the open deviation found while measuring: the
inbound webhook now awaits the interpretation call, so it takes 1-3 s instead
of the < 200 ms the spec requires for an enqueue-only handler.
…here

Spec §7.5 asks the inbound Twilio webhook to answer in under 200 ms and never
call the LLM inside it. Wiring the provider exposed that the webhook awaited
the interpretation, taking 1-3 s per message.

- app/runtime.py builds the worker's rescue runtime once per process (channel,
  scheduler, workforce, orchestrator, interpreter); the API process no longer
  builds any of it.
- process_inbound_message runs the async handling in the worker with bounded
  retries; run_due_jobs ticks the scheduler and publishes the degraded-status
  snapshot to Redis for the API probe; purge_old_messages is finally scheduled.
- The webhook validates the signature and enqueues, returning TwiML 200 (500 on
  broker rejection so Twilio retries instead of silently dropping a message).
- /api/status reads configuration (new pure is_provider_configured helper) plus
  the published snapshot instead of another process's internals.
- The FastAPI lifespan ticker is gone: beat owns time.
- Tracing bootstrap: the worker installs its own TracerProvider per preforked
  child and flushes on shutdown. Without it the LLM calls that now happen in the
  worker produced no Langfuse traces at all (found live, verified fixed: four
  observations exported for one message).

Verified live: webhook 15-50 ms, worker ~2.9 s per message off the request
path, two children concurrently without loop errors, beat ticking every 5 s.
296 tests pass, ruff and mypy clean.
…ds with prompt v2

The first real-model run exposed three defects, and the most serious one was the
gate itself.

- check_thresholds() compared hardcoded '..._min' keys against a report that
  stores the metrics without the suffix, so every lookup returned None, every
  comparison was skipped and the runner printed 'Thresholds met.' while accuracy
  was 0.86 against a 0.92 minimum. evals/thresholds.yaml was never read at all.
  Thresholds now live in app/evals/thresholds.py, read the YAML (one source of
  truth, min and max directions) and fail closed: an unmapped key, an unknown
  metric or a non-numeric value is a violation, never a silent pass.
- _build_prompt forwarded rescue_id, pending_offers and shifts_48h but dropped
  pending_confirmation, so a bare '1' had no way to be read as ABSENCE_CONFIRM.
  It now renders that key and any remaining non-empty context key, with a
  credential-name filter so a secret can never reach a prompt.
- New interpreter_v2 prompt with an explicit procedure keyed on what is pending,
  numeric replies meaning yes/no only when something is pending, a broader health
  rule and examples for the confusions v1 exposed. The active version is named
  once and recorded on every interpretation.

Measured with the real provider on the 150-sample golden set:
  v1: 0.86 accuracy / 0.9167 health / 0.95 times (2 violations, gate silent)
  v2: 0.9333 accuracy / 1.0 health / 0.85 times (thresholds met, honestly)
The 10 remaining failures are analysed in odd/tasks/interpreter-quality.md; six
of them need the orchestrator to pass an 'already accepted' marker because the
context the harness supplies is identical to cases labelled OFFER_DECLINE.

308 tests pass, ruff and mypy clean. Eval run artifacts are now gitignored.
…rompt v4

Three gaps, one of which blocked the dashboard's decisions screen.

- The orchestrator never persisted interpretations: the interpretation table had
  no writer in the whole app, so the Agent decisions screen had no real data
  source. One row per LLM interpretation now records intent, confidence, the
  structured fields, model, prompt version, tokens, latency and cost. The parser
  path writes nothing (a degraded decision is not an LLM decision), a duplicate
  provider message writes no second row, and a persistence failure is logged
  without breaking the rescue. _persist_inbound now returns the message id.
- OFFER_WITHDRAW was not inferable: the six withdrawal golden rows carried the
  same context as the decline rows, so the state that makes them withdrawals is
  now sent as accepted_offers instead of being guessed from the words 'al final'.
  The fixture gained that key for those rows only, and no label changed.
- interpreter_v4 fixes the prompt's own example: every version since v1 claimed
  'hasta mediodía puedo' fills proposed_start 07:00 as well, while the labels
  expect start=null; the model followed the example and lost 5 of 20 extractions.
  v4 states the rule the data encodes (one boundary fills one field).

Measured with the real provider on 150 golden samples:
  v2 0.9333 / 1.0 / 0.85    v3 0.9867 / 1.0 / 0.75 (times violated)
  v4 0.9933 / 1.0 / 1.0     one failure left: 'cancele el caso de todos'

Also fixes the CI-only test failure: build_runtime(settings) ignored its own
settings for the database and fell back to the ambient .env, so the test
connected to a developer's local Postgres and failed on a clean runner. It now
passes the URL explicitly and the test uses a temp-file database; verified with
the Linux CI image.

317 tests pass, ruff and mypy clean.
The test failed in CI and passed locally: build_runtime(settings) ignored the
settings it was given for the database and fell back to the ambient .env, so the
test connected to the developer's local Postgres on 5433 and found nothing to
connect to on a clean runner (OSError, 111).

The runtime now passes settings.database_url to the engine factory, and the test
builds a temp-file SQLite database with the schema, so it is hermetic on any
machine. Reproduced and verified with the Linux CI image
(ghcr.io/astral-sh/uv:python3.12-bookworm): 308 passed before the fix,
308 passed and no failures after it.
The dashboard was a complete UI with no HTTP client: every screen rendered mock
data and the backend exposed only /health, /api/status and the Twilio webhooks.

- Real login: Argon2 hashes for the seeded demo managers, HS256 JWT, and
  role-gated dependencies (401 without a valid token, 403 for the operator-only
  interpretation endpoints). The seeded placeholder hash authenticates nothing.
- Read endpoints for the whole dashboard slice, with response shapes that match
  frontend/src/domain/types.ts field for field (asserted by contract tests, so a
  renamed field breaks the build): locations, shifts, settings, rescues (list and
  detail with timeline, offers and candidates), approvals, conversations and
  their redacted messages, interpretations with filters, and metrics (LLM cost
  per day, p50/p95 latency, low-confidence rate, delivery failures, stuck
  rescues).
- Write endpoints go through the worker: approvals and rescue close enqueue a
  Celery task and answer 202, so the API process still never builds a runtime or
  owns domain logic.
- CORS restricted to the configured dashboard origins.
- Redaction holds: only stored redacted bodies are served and no health detail
  appears in any response.

Verified live against the real stack: login 200, 401 without a token, 403 for a
manager on /api/interpretations, operator 200, wrong password 401, and the full
chain webhook (32 ms) -> worker -> LLM -> persisted interpretation -> served by
the API (OFFER_CONDITIONAL, 0.9, gpt-4o-mini, $0.00038).

375 tests pass, ruff and mypy clean.
The dashboard was a complete UI rendering mock data with no HTTP client. It now
authenticates against the real API and reads the real database, using the seam
that already existed (DashboardDataSource + hooks) instead of rewriting screens.

- Session client (localStorage), an HTTP client that attaches the bearer token
  and turns failures into a typed ApiError/NetworkError, and a 401 path that
  clears the session and returns to the login screen.
- ApiDashboardDataSource: shifts, active rescues, rescue detail, approvals
  (enqueue + invalidate, since decisions are applied by the worker), conversations
  and messages, interpretations, metrics and settings. The location id is
  resolved from GET /api/locations instead of the hardcoded literal that never
  matched the seeded database.
- Login screen and RequireAuth guard, with the demo credentials shown for the
  reviewer, distinct messages for wrong credentials and for an offline API, the
  signed-in manager in the header with a logout action, and deep links preserved
  across login.
- getDataSource() keeps the mock available behind VITE_USE_MOCK (offline demo and
  hermetic tests). The Evals screen still reads mock data and now says so in the
  UI, because /api/evals/runs* was deliberately deferred.
- Dev proxy /api -> localhost:8000 so development needs no CORS, documented
  environment variables, and runbook instructions for logging in.
- Fixes a real defect found on the way: the settings form seeded its draft once,
  so live values arriving later could be overwritten by a save with stale
  defaults.

Verified: 119 frontend tests, build, oxlint and tsc clean; the SPA serves and the
dev proxy reaches the API (401 without a token, as expected); CORS accepts the
configured dashboard origin and rejects others.
Reproduced live: the worker crashed with 'got Future attached to a different
loop' and lost the inbound message. asyncio.run() creates a new loop per task
while the memoized async SQLAlchemy engine holds asyncpg connections bound to the
loop that created them, so a reused pooled connection fails. It passed earlier
only because the pool was still empty, which is exactly why it survived the
suite: intermittent, state-dependent, and invisible until the pool had idle
connections from a dead loop.

app/workers/async_runner.py now owns the rule (one lazily created loop per
process, reused via run_until_complete, recreated only when closed) and every
worker task goes through it. The regression test reproduces the asyncpg binding
mechanics and fails with the old behaviour.

380 tests pass, ruff and mypy clean.
…mation is pending

Two defects found by running the absence flow end to end with the LLM enabled.

- The ABSENCE_REPORTED audit event was written with a synthetic
  'case_<shift>_<timestamp>' id while the case itself got a uuid id, so the
  dashboard's timeline query joined on a case that does not exist and every
  rescue timeline rendered empty. It now records the case it was written for.
- The interpreter was never told that a confirmation was pending, while the
  prompt branches on exactly that: with nothing pending, a bare 'sí' is UNCLEAR
  by design, so a real employee answering SÍ got 'no te he entendido' and the
  absence was never confirmed. The orchestrator now sends pending_confirmation
  (the shift of the OPEN case awaiting confirmation).

The eval suite could not catch the second one: the golden fixture supplies
pending_confirmation, so it measured a context production never produced. A
contract test now pins the context keys the orchestrator sends, and an
end-to-end test drives report -> confirmation -> OFFERING through a fake that
mimics the prompt's procedure.

Verified live: the case advanced OPEN -> OFFERING with 3 offers and the timeline
now shows ABSENCE_REPORTED, RESCUE_OPENED and three OFFER_SENT events.

383 tests pass, ruff and mypy clean.
Timers never fired in the deployed stack, so no rescue ever escalated: two cases
were 20 minutes past their deadline and /scheduler_ran_jobs had never once
appeared in the worker log.

Root cause: the in-memory SimScheduler is per process and the worker runs Celery's
prefork pool (12 children). The task that opens a case registers the deadline in
its own scheduler, while the beat tick lands in whichever child Celery picks, with
its own empty queue. They practically never coincide.

- CeleryScheduler implements the same port as SimScheduler: schedule() becomes one
  deferred Celery task with countdown from the injected clock, so Redis holds the
  timer and any free worker executes it at the right time. Restarting a worker no
  longer loses pending timers, which removes a limitation the project had
  documented as accepted.
- apply_scheduled_job resolves the handler in the executing process; an unknown
  name fails loudly. SimScheduler stays the backend for the eval harness and
  tests, selected by settings.scheduler_backend or an injected scheduler.
- reconcile_stale_cases (beat, every 60s) re-enqueues the deadline of cases that
  still expect action, so a timer lost between a commit and an enqueue recovers
  itself instead of hanging forever.
- Today board: shifts in state 'scheduled' (staffed, untouched) now appear in the
  covered column instead of nowhere, so 'Covered today' no longer reads 0 while
  nine shifts are staffed.

And the defect the fix immediately exposed: with timers firing for the first time,
escalation failed with

  StringDataRightTruncationError: value too long for type character varying(64)
  ('audit_case_bc1809c0..._escalated_1790359021_DEADLINE_REACHED', ...)

because _escalate composed an 81-character audit id and the transaction rolled
back every time, so no rescue had ever escalated in production. The quiet-hours
marker had the same shape at exactly 64 characters. Both use opaque audit_<uuid>
ids now, and the existing width guard — which had only ever exercised the
confirmation path — drives the escalation and deferred-wave paths too.

Verified live: a timer scheduled 40s out appeared in 'celery inspect scheduled'
with its ETA and executed at exactly that second; two overdue cases escalated; the
repeated reconcile sweep changed nothing.

402 tests pass, ruff and mypy clean, frontend 119 tests, build and lint clean.
The dashboard could watch the agent work but not produce the work: demoing the
core flow needed a second phone or ten minutes of waiting for a wave timeout.

- POST /dev/simulator/{employee_id}/messages resolves the employee and enqueues
  the SAME task the Twilio webhook enqueues, so a simulated message is
  indistinguishable in the database from a real one except for its synthetic
  provider id: idempotency, threading, interpretation, auditing and delivery all
  behave identically.
- POST /dev/clock/advance moves a Redis-backed offset every worker process reads
  (deadlines, quiet hours, eligibility) and immediately enqueues the reconcile
  sweep, so an overdue case escalates in seconds instead of waiting for the beat
  tick. GET /dev/clock exposes the virtual time.
- GET /api/employees gives the screen the roster it needs; the Simulator screen
  now shows real employees, their real threads, a send action and the demo clock,
  keeping its phone-frame layout, with the mock still behind VITE_USE_MOCK.
- Both /dev routes are registered only when APP_ENV is local/test/demo and even
  then require a manager JWT: a production deployment does not advertise them.

Verified live: a message sent through the simulator opened a real rescue with the
right shift and offers, and advancing the clock escalated it.

428 tests pass, ruff and mypy clean, frontend 124 tests, build, lint, tsc clean.
Three defects of one family, found by driving the product end to end: the agent
asked something and the answer was ignored.

- 'Which shift?' was a dead end. When an employee has two or more shifts in the
  48-hour window — the common case — the agent asked and nothing consumed the
  reply, so the absence was never reported. Verified live: an employee with a
  shift in progress could not report an absence at all. The orchestrator now
  sends the candidate shifts and, once it has asked, resolves the answer: the
  model returns the chosen shift id, and the deterministic parser matches the
  same candidates in degraded mode. It never guesses, asks once more, then
  redirects.
- The interpreter never received shifts_48h, although the prompt and the golden
  fixture both used it: third instance of the fixture describing a context
  production never produced. Candidates now carry their date and the context
  carries today's date, so 'el de hoy' is resolvable. A context-contract test
  pins the exact key set for the four relevant states.
- An absence that is never confirmed hung forever: the state machine had no
  OPEN + DEADLINE_REACHED transition, so the manager was never told. It now
  escalates and notifies (the product decision), with a new eval scenario and an
  INV6 relaxation that keeps requiring ABSENCE_REPORTED + ESCALATED for that path.

interpreter_v5 adds the shift-choice branch (edited in place: v5 was created in
this same change and never shipped, so inflating a version would be dishonest);
golden gains five rows for the new branch, with only those rows touched.

Measured with the real provider on 155 samples: accuracy 0.9935, health 1.0,
conditional times 0.95 — thresholds met, one failure left.

Verified live: an employee with two upcoming shifts reported an absence and the
case opened for the shift they named, and an unconfirmed absence escalated with a
real WhatsApp to the manager (Twilio answered 201).

435 tests pass, ruff and mypy clean, 27 eval tests green.
DESIGN.md §8 already defines the responsive contract (breakpoints, collapsing
strategy, 44px touch targets, gutter scale); the dashboard did not honour it.

- The header now has a hamburger drawer for phones and tablets. It is lg:hidden,
  not md:hidden: between 768 and 1023 px the main nav is visible but the operator
  group is not, so a md-only drawer left Ops, Agent decisions and Evals
  unreachable on a tablet. The panel lists both groups with the gold active
  indicator, closes on Escape and on selection, keeps the manager and logout
  reachable, and its visible labels are English (they were PRINCIPAL/OPERADOR).
- The two wide tables render as stacked cards below md and keep the table from md
  up in an overflow-x-auto container; every touch control meets 44px through
  pointer-coarse (the wordmark, the demo pill and Log out were 28/36/20px).
- Settings and Evals grids reflow, the chart scales to its container, and the
  screens were audited at 360px for overflow and touch targets.
- Agent decisions shows a plain-English explanation when the operator-only
  endpoint answers 403 instead of rendering a blank page that looks broken.

Verified in a real browser (Playwright + Chromium installed outside the project,
hasTouch on so pointer:coarse matches) at 360, 768 and 1440 px: no horizontal
overflow, every destination reachable, every visible control >= 44px, and the
screens reviewed by eye. The sweep for Spanish UI copy now also flags accented
literals, which is what caught the Spanish group labels.

141 frontend tests, build, lint and tsc clean.
Two things a manager saw in the browser console and in the table.

- The network tab showed four 403s for one screen. A 403 (wrong role) answers the
  same however many times it is asked, so TanStack Query's default three retries
  only produced three extra failing requests and three extra red rows. The retry
  policy now skips 4xx and keeps two attempts for network and 5xx failures, with
  mutations never retried.
- The Agent decisions table rendered the raw API timestamp
  ('2026-09-25T19:27:33.483410+00:00') because the mock had always carried an
  already formatted '15:03'. The conversion moved to domain/decisions.ts with its
  own tests (ISO to HH:MM, an already compact value passes through, an
  unparseable one is shown rather than hidden).

Verified in the browser: the manager now issues one request instead of four and
still sees the explanation for the operator-only view; the operator account sees
the real rows with local times (21:27).

144 frontend tests, build, lint and tsc clean.
…scenario real

Two defects the simulator screen exposed.

- The demo clock read '--:--' because the SPA never reached the API: the Vite dev
  proxy only covered /api, so /dev/* hit the SPA fallback and returned HTML, and
  the deployed Caddyfile did not route /dev/* either — the simulator would have
  been broken in the demo environment too. Both now forward /dev to FastAPI; the
  endpoint is still demo-gated by APP_ENV, so production keeps answering 404.
- 'Load scenario: acceptance race' was a button with no handler. The offer
  payload now carries employeeId, and the control runs the scenario for real: it
  finds the OFFERING rescue with at least two pending offers, fires both
  acceptances concurrently (a sequential pair would not race), refreshes the
  affected queries, and says what to look for. With no eligible rescue it
  explains itself instead of doing nothing.

Verified live end to end: the clock shows the virtual time (21:32), and clicking
the scenario covered the shift with exactly one winner — Marta Lopez ACCEPTED,
the other two offers CANCELLED, case COVERED, shift covered — which is the
concurrency invariant the eval scenario asserts.

Also confirmed while checking the outcome: notifications sent outside a
conversation (the loser's 'already covered', the manager's escalation and
coverage notices) are delivered but never persisted, so the dashboard shows no
trace of them. Recorded as a follow-up.

148 frontend tests, 435 backend tests, ruff, mypy, build, lint and tsc clean.
…oard table

Three things the demo exposed, plus a fidelity fix.

- The Simulator now says what the agent will do: each frame labels the employee
  (on shift now / starts at HH:MM / ended at HH:MM / no shift today), the frames
  that can actually report an absence come first, and when nobody can, the screen
  says why and what to do. Writing into a frame whose shift already ended could
  only ever produce the out-of-scope reply, which made the screen look broken
  while the agent was right.
- The demo clock is reversible and honest: POST /dev/clock/reset (demo-gated,
  JWT) zeroes the offset and enqueues the reconcile sweep, and the control shows
  the virtual time, the offset in plain words and a Reset. A leftover offset had
  reached +18h50m from testing, which moved the worker's 'now' a day ahead: today's
  shifts read as finished, the agent answered out of scope, and stored timestamps
  jumped into the future.
- Today is now a dashboard table (ROLE / WINDOW / EMPLOYEE / STATUS / RESCUE /
  ACTIONS) with the existing tokens, responsive as the other wide screens are.
- Fidelity fix found in the new table: escalated rescues read as 'Uncovered — no
  candidates offered yet', which was worse than the board. A shift whose latest
  case escalated now shows an Escalated badge with the honest caption (the manager
  was notified), a covered shift never reads as escalated, and any shift with a
  case — terminal included — offers View detail.

159 frontend tests, 437 backend tests, ruff, mypy, build, lint and tsc clean.
# Conflicts:
#	backend/app/runtime.py
#	backend/tests/unit/test_runtime.py
…fer text

Two defects the live verification exposed, both in what the agent says.

- Every message the agent cannot act on landed in one generic line: solo
…reload

The user drove the product and could not use it. Five defects, all verified in a
browser walkthrough.

- The Simulator looked broken: the agent takes 10-15 s to answer and nothing
  refreshed, so the reply only appeared after a manual reload. The thread now
  polls after a send (every 3 s, bounded to ~30 s, stopping on the reply, on the
  bound and on unmount), says "the agent is replying…" while waiting and blocks
  duplicate sends. Verified: the confirmation question now lands on its own.
- The manager had nothing to do on an escalated case: the detail screen showed
  the timeline and the candidates with only navigation buttons, while the API
  already exposed the close. It now has Close case (confirmed, queued, 202) and,
  for AWAITING_APPROVAL, approve/reject inline.
- An escalated shift read "Uncovered — No candidates offered yet" on Today,
  because the board only received active cases and the row fell back to "nothing
  was tried". The board now gets the day's cases including terminal ones and
  shows Escalated with "Escalated at HH:MM" and a summary of who was contacted
  and what each answered (spec §6.4), never a countdown on a terminal case.
- docs/demo-script.md is now a walkthrough a newcomer can follow: what each
  screen is for, the exact Simulator steps (including the ~10 s wait), the state
  table, what the manager does on escalation, and the late-acceptance beat.

Deliberately NOT done: cancelling the open offers when a case escalates. That was
my own wrong call in the task document — spec §5 says a late acceptance after the
escalation becomes AWAITING_APPROVAL for the manager, so the offers must stay
open. Reverted with the test that pins it.

175 frontend tests, 444 backend tests, ruff, mypy, oxlint, build and tsc clean.
… reply

The user ran the whole flow and got something confusing: they reported an absence,
answered the confirmation 39 minutes later, and the system produced an approval
request instead of confirming. Three defects behind it.

- **Stale offers were still acceptable.** A candidate's `PENDING` offer whose
  shift had already ended could be "accepted": the live case shows offers from
  shifts three days old being answered, and one of them created a
  `late_acceptance` approval for a finished shift. `_try_accept_offer` now checks
  that the shift has not ended before doing anything with the offer; a stale one
  is cancelled with an `OFFER_STALE` audit and the candidate gets the honest
  close-out. A late acceptance for a shift that is *still ahead* keeps going to
  the manager, as spec §5 requires.
- **A late answer got a generic line, or none at all.** A recent closed case now
  gets its own outcome message (`state_case_escalated`, `state_case_covered`,
  `state_case_closed`) instead of "solo gestiono avisos de ausencia", which was
  the opposite of the truth. Cases older than the window (6 h) stay out of it.
- **Not every reply was stored**, so the dashboard showed the employee's message
  with no answer — the reason the flow looked broken. Every message to an
  employee now lands in that employee's conversation, and the wave offers use the
  same phone-based thread the WhatsApp webhook uses, so there is one thread per
  employee instead of two.

Verified live end to end: an employee answering "SÍ" after their case was closed
now receives "tu aviso del turno de 09:00 a 17:00 quedó cerrado por tu encargado"
and the reply is visible in the conversation.

447 tests pass, ruff and mypy clean.
The user ran the flow and got lost because the screen carried three days of test
leftovers: a stale "[template: offer]" bubble, threads where the agent answered
"no te he entendido" to a question about a shift that had already closed, and a
demo clock left ten hours ahead. None of that was the product misbehaving, but it
is exactly what makes a demo unreadable.

POST /dev/demo/reset (demo-gated, manager JWT) removes the operational artifacts
in foreign-key order (interpretation, message, conversation, approval_request,
offer, audit_event, rescue_case), drops the shifts from today onwards, reseeds the
schedule and returns the demo clock to real time through the same path
POST /dev/clock/reset already used. It reports what it removed, so the UI can be
honest about it.

The Simulator gets a Reset demo data control next to the clock, with a
confirmation that names what disappears and says it cannot be undone, plus the
result line. Verified live against real Postgres: dirtying the environment and
resetting from the UI printed "Deleted 2 interpretations, 6 messages, 4
conversations, 3 offers, 5 audit events, 1 rescues, 143 shifts. demo clock back on
real time." and left the board with every shift Scheduled and no case.

docs/demo-script.md starts with that step now, because leftovers from a previous
run are the main source of confusion.

451 backend tests, 179 frontend tests, ruff, mypy, oxlint, build and tsc clean.
aam9063 and others added 6 commits September 28, 2026 14:04
The Simulator rendered `roster.slice(0, 3)`, a leftover from the mock that had
three phones. Because the roster sorts the actionable employees first, the three
visible frames were always the ones already on shift: every employee whose shift
starts later — the calmest way to run the demo, since an in-progress shift gives
only a ten-minute confirmation window — was invisible. The user reported exactly
that ("no hay ninguno que ponga Starts at") while four such frames existed.

The whole roster renders now (cards below md, three across from xl), with the
actionable frames first and the employees without a shift today last. Verified in
a browser at real time: 3 frames say "On shift now" and 4 say "Starts at HH:MM"
(Lucía 13:00, Elena and Iker 15:00, Tomás 18:00).

179 frontend tests, build, lint and tsc clean.
…to miss

The demo clock had been clicked to +40 hours. With the clock that far ahead every
shift of the day reads as finished: the Simulator showed "04:01 (+20 h ahead)",
every frame said "Ended at HH:MM", and the banner admitted nobody could report an
absence. Nothing stopped the drift and nothing made it obvious.

- POST /dev/clock/advance clamps the total offset to a documented ±6 hours and
  answers with the applied offset plus a `clamped` flag, so the drift stops here.
  The old `incrby` was atomic but could not clamp; the read-modify-write is fine
  for the single-operator demo use case and the review notes it.
- From one hour in either direction the small grey note becomes a prominent
  alert that states the consequence plainly — the shift windows are not where
  they look, so people who should be on shift read as finished and vice versa —
  with a Reset clock action inside it.
- The wire type and the hook carry the applied offset and the flag, and the
  screen says when a request was cut.

Verified live: a +20 h request is clamped to 6 h with clamped: true, a second one
stays at 6 h, a normal +1 h is untouched, and at +1 h the banner appears with its
Reset action.

455 backend tests, 182 frontend tests, ruff, mypy, oxlint, build and tsc clean.
The user reported that after writing in the Simulator they had to reload the page
to see the agent's answer — and they were right: the reply poll only started when
the employee already had a conversation.

    onSuccess: (_sid, variables) => {
      if (variables.conversationId) {        // null on first contact
        startReplyPoll(variables.conversationId)

`GET /api/employees` returned the conversation id only when a Conversation row
existed, so after a demo reset every frame was in that state: the poll never
started and the answer only appeared on a manual reload. With a previous thread it
worked, which is why it looked fine in earlier checks.

A thread now exists conceptually from the start: the roster always advertises the
deterministic id the orchestrator already uses (`conv_twilio_<phone>`), and
`GET /api/conversations/{id}/messages` answers an empty list for a real
employee's empty thread instead of a 404 (a normal state, not an error — unknown
ids and numbers that do not exist keep their 404). The frames show an honest
empty state ("No messages yet — write the first one.") and the simulator's 202
carries the conversation id so the poll has somewhere to look immediately.

Verified live on a reset database, the exact reported situation: a frame with no
conversation received the message, showed "the agent is replying…", and the
confirmation question appeared **5 seconds later with no reload**.

457 backend tests, 183 frontend tests, ruff, mypy, oxlint, build and tsc clean.
…hboard

The Evals screen was the last one rendering invented numbers, and it admitted it:
"eval data is not live yet". Behind it, the `eval_run` table had the columns for
exactly this (git sha, trigger, timings, model config, prompt versions, metrics,
invariant violations, passed, baseline, report path) and **zero rows** — no writer
ever existed, and `/api/evals/runs*` was never implemented. Meanwhile the real
evaluation ran from the terminal and its results went nowhere.

- `app/evals/recording.py` persists one row per run: the golden set (metrics,
  verdict from the threshold gate, model, prompt version, report path) and the
  scenario suite (per-scenario pass/fail plus the invariant violations).
  Recording is best effort: it can never break a run, and a failing run is
  recorded as failing rather than dropped.
- `evals/runner.py` records after each run and gains `--scenarios` to run and
  record the scenario suite. Two bugs surfaced while wiring it: the harness
  reports carry `datetime` values that JSON columns reject (normalized now), and
  `Settings` looked for `.env` in the current directory, so a runner started from
  the repository root silently used the default database port instead of the
  configured one — it now resolves `backend/.env` from the module.
- Three operator-only endpoints (`/api/evals/runs`, `/runs/{id}`, `/runs/summary`)
  serve the screen's shape. The threshold travels **inside the run**: the API
  container does not carry `evals/thresholds.yaml`, and a stored verdict should be
  self-explanatory. The verdict describes the newest run that was actually judged,
  so the informational offline baseline never reads as a product failure.
- The screen renders the real summary (verdict, commit, accuracy history with the
  threshold line, model comparison, invariants, per-scenario results) and shows an
  honest empty state naming the command that produces a run.

Verified live with recorded runs: 15/15 scenarios, 0 invariant violations, and the
real model at 0.9871 intent accuracy against the 0.92 gate — shown in the
dashboard, no mock note left.

469 backend tests, 188 frontend tests, ruff, mypy, oxlint, build and tsc clean.
…eclare the demo password once

CI failed on five tests that passed locally, which is the signature of an
ambient dependency.

- `tests/unit/api/test_dev_tools.py`: the `/dev/clock/advance` tests did not use
  the `stub_sweep` fixture, so they enqueued the reconcile sweep for real. On a
  development machine Redis is up and they passed; in CI Celery retried the
  result store for four minutes and the five tests failed with "Retry limit
  exceeded while trying to reconnect to the Celery result store". They stub the
  enqueue now, and `tests/unit/api/conftest.py` adds an autouse fixture that
  records every enqueue for the whole unit API suite — no unit test can reach a
  broker by accident again. Verified with the Linux CI image: 469 passed, and the
  suite went from 225 s to ~105 s because nothing retries a missing Redis.
- The demo password was repeated across tests, docs and two frontend modules. It
  is meant to be public (a fictional venue with fictional data, shown on the
  login screen so a reviewer can get in), so it now has exactly two declared
  homes — `DEMO_PASSWORD` in the seed and in `frontend/src/domain/demoCredentials.ts`
  — and every test imports it instead of repeating the literal. The runbook keeps
  the documented credentials table (that is the point of a demo credential) and
  its curl example no longer embeds the value.

469 backend tests, 188 frontend tests, ruff, mypy, oxlint, build and tsc clean.
@gitguardian

gitguardian Bot commented Sep 28, 2026

Copy link
Copy Markdown

⚠️ GitGuardian has uncovered 4 secrets following the scan of your pull request.

Please consider investigating the findings and remediating the incidents. Failure to do so may lead to compromising the associated services or software components.

🔎 Detected hardcoded secrets in your pull request
GitGuardian id GitGuardian status Secret Commit Filename
37629664 Triggered Generic Password 6f417c8 frontend/src/screens/LoginScreen.test.tsx View secret
37629664 Triggered Generic Password 05a7c8b frontend/src/domain/demoCredentials.ts View secret
37629664 Triggered Generic Password ff96627 frontend/src/domain/demoCredentials.ts View secret
37629664 Triggered Generic Password ff96627 backend/app/db/seed.py View secret
🛠 Guidelines to remediate hardcoded secrets
  1. Understand the implications of revoking this secret by investigating where it is used in your code.
  2. Replace and store your secrets safely. Learn here the best practices.
  3. Revoke and rotate these secrets.
  4. If possible, rewrite git history. Rewriting git history is not a trivial act. You might completely break other contributing developers' workflow and you risk accidentally deleting legitimate data.

To avoid such incidents in the future consider


🦉 GitGuardian detects secrets in your source code to help developers and security teams secure the modern development process. You are seeing this because you or someone else with access to this repository has authorized GitGuardian to scan your pull request.

@aam9063
aam9063 merged commit 78345b4 into main Sep 28, 2026
4 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant