Skip to content

Dev - #25

Merged
aam9063 merged 7 commits into
mainfrom
dev
Sep 29, 2026
Merged

Dev#25
aam9063 merged 7 commits into
mainfrom
dev

Conversation

@aam9063

@aam9063 aam9063 commented Sep 29, 2026

Copy link
Copy Markdown
Owner

No description provided.

aam9063 and others added 7 commits September 28, 2026 18:34
… on its own

The dashboard was pull-only: a change produced elsewhere — a candidate accepting,
a case escalating, a rescue closed by the worker — stayed invisible until the
manager refocused, acted or reloaded. Spec §7.5 already defined the endpoint and
Caddy already routed it.

- `app/events.py`: a typed `DashboardEvent` (name, location, rescue, shift,
  origin, at — never a body, phone, name or health flag) and an `EventBus` port
  with a Redis publisher, an async subscriber and a no-op fake. Publishing is
  best effort: a broker outage degrades to the previous pull-only behaviour
  instead of failing a rescue.
- The orchestrator publishes at every transition; the worker runtime injects the
  Redis bus (it was silently using the no-op, which is why the channel was empty
  with a healthy socket).
- `WS /ws/locations/{id}?token=…` validates the manager's JWT and their right to
  that location before accepting (4401/4403), relays the location's channel and
  cleans up on disconnect (4503 when the broker is down). The relay and the
  client listener run together, and the socket is closed before anything is
  cancelled — awaiting or cancelling a blocked `receive()` deadlocks the test
  client, which is how the suite hung the first time.
- The frontend `useLiveEvents` hook maps each event to the query keys the actions
  already invalidate (so nothing is duplicated in the client cache), reconnects
  with capped backoff, clears the session on 4401 and stops on unmount; the shell
  shows a quiet indicator while it is not live, and the dev proxy forwards `/ws`
  with the upgrade.

Verified live end to end: the events reach Redis with the privacy rules intact,
and with two browsers open a rescue produced in one moved the other's Today board
2.5 s later with no reload.

493 backend tests, 196 frontend tests, ruff, mypy, oxlint, build and tsc clean.
Two kinds of message reached a real phone and left nothing behind, so the
dashboard could not tell the story and the audit trail had holes.

- The manager's notices (escalation, coverage) go to a phone, not to an employee
  conversation, so nothing was persisted: the timeline showed a rescue that
  escalated "by itself". They now pass through one helper that sends and records
  `MANAGER_NOTIFIED` on the case with the template key, or
  `MANAGER_NOTIFY_SKIPPED` with the reason when the manager has no reachable
  phone — the attempt is recorded, never a delivery claim. The rescue timeline
  labels both ("Manager notified" / "Manager could not be reached").
- The race loser's "already covered" notice is an employee message and now
  carries the conversation, so it appears in their thread instead of looking
  unanswered.

Verified live: an escalation produced `MANAGER_NOTIFIED {template:
manager_escalated}` on the case.

496 backend tests, 196 frontend tests, ruff, mypy, oxlint, build and tsc clean.
The rotation has fixed windows, so a demo late in the evening found today's
shifts finished and tomorrow's outside the roster's "today": nobody could report
an absence and the board looked dead. That is the trap the user hit ("no hay
ninguno que ponga Starts at").

When the seeded day is dry the seed now anchors two shifts to the current hour
instead — one in progress, one starting soon — assigned to employees the rotation
left free that day, and only for the window that is actually missing (at 03:00
nothing is in progress; at 23:30 the bar close already covers "now" and only the
upcoming one is added). During working hours nothing changes, so the normal case
stays exactly as it was, and the pool is taken from the rotation itself: an
invented employee id would schedule a shift for somebody who does not exist.

Verified live after a reseed: 5 shifts in progress and 2 starting within six
hours. 498 backend tests, ruff and mypy clean.
The last endpoint of spec §7.5 that was missing, and three loose ends.

- `POST /api/shifts/{id}/absence`: the manager marks an absence and the rescue
  opens immediately (no WhatsApp confirmation needed, the manager is the
  authority). It enqueues a Celery task and answers 202 like the other writes;
  404 for an unknown shift or one outside the manager's locations, 409 when the
  shift is already absent or already has a live case. The orchestrator method
  reuses the confirmed-absence path (HRIS marked absent, case opened, deadline
  scheduled, wave sent, manager notified) and audits `ABSENCE_MARKED` with the
  manager as the actor, with a redelivery guard so a retried task cannot open a
  second case.
- `run-due-jobs` is gone from beat: it drove the in-memory scheduler, which the
  broker-owned timers made obsolete, and it always ticked zero. The runtime
  snapshot it used to publish moved to `reconcile-stale-cases` — that snapshot is
  what the degraded banner reads, so removing the tick without moving it would
  have silently stopped an open circuit breaker from reaching /api/status.
- Sentry is initialized when `SENTRY_DSN` is set (API lifespan and worker
  bootstrap), best effort and silent when unset; the DSN is never logged.
- The OTel trace id is stored on each interpretation when a span is active, so
  the decision detail can link to its Langfuse trace instead of always showing
  null.

515 backend tests, ruff and mypy clean.
The floating action said "Report absence" and did nothing — a dead control in
the dashboard, and the reason the new `POST /api/shifts/{id}/absence` was only
reachable from a terminal.

The button now opens a small dialog listing today's still-scheduled shifts (with
role, window and who holds them) and marks the chosen one absent. The copy is
honest about the 202 — "the rescue is opening; the worker applies it in a
moment" — and a refusal reaches the screen with the server's reason (409 when the
shift is already absent or already has a live case, 404 when it is not available
for this location). The board refetches on success and the live channel delivers
whatever else changes.

Verified in the browser: the button opens the dialog with 10 shifts, marking one
was accepted, and Today showed that shift as absent with its rescue searching.

199 frontend tests, build, lint and tsc clean.
@aam9063
aam9063 merged commit b34da44 into main Sep 29, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant