An AI agent that covers last-minute shift absences — so the manager doesn't have to.
Shift Rescue automatically covers same-day absences for shift-based businesses (hospitality, retail, staffing). It works over WhatsApp on top of the existing HR system, and the manager always stays in control.
When someone calls in sick, it is almost always less than two hours before their shift, announced by WhatsApp. The manager — in the middle of opening, receiving stock, setting up the floor — spends 30 to 60 minutes messaging and calling around, without knowing who is available, who is close to their hour limits, or who closed the night before. Uncovered shifts start short-handed: slower service, stressed teams, worse reviews.
A shift gets covered on its own — reliably, measurably and safely.
- An absence arrives by WhatsApp. The agent confirms it with the absent employee in a short conversation.
- Deterministic rules decide who is eligible — availability, hour limits, rest rules, overtime. The LLM only interprets language and drafts replies; it never decides who gets the shift.
- The agent does the legwork: it offers the shift to eligible candidates over WhatsApp and tracks their answers.
- The manager approves overtime, partial coverages and cancellations — and audits everything the agent did in one timeline.
The manager's home screen: today's shifts with role, window, employee and rescue status, plus a live count of active rescues.
Every rescue is a case with a full audit trail: what the agent did and when, which candidates were offered the shift, their scores, and who accepted.
The agent confirms the absence in the employee's own channel, in their language, with explicit Sí/NO confirmation.
The agent offers the open shift to an eligible candidate and handles the answer — full shift, partial coverage, or no.
- Deterministic core, LLM at the edges. Business rules (eligibility, assignment, what needs approval) live in testable code. The LLM interprets language and drafts messages. The LLM proposes; the domain disposes.
- The manager stays in control. The agent does the legwork; humans approve anything sensitive. Every agent action is logged and auditable.
- Built for production failure modes. Duplicate messages, simultaneous replies, providers down, employees typing unexpected things — the system is designed to detect, recover and degrade gracefully.
- Measurable. Agent decisions and eval runs are first-class: everything is traced and reviewed, not vibes.
Currently in development as a portfolio-grade MVP. See
docs/SHIFT_RESCUE_SPEC.mdfor the full specification and docs/system-design.md for the complete system design.
Prerequisites: Docker and uv (Python 3.12), Node 24 and pnpm 11.
# 1. Boot the backend stack (postgres, redis, api, worker, beat)
make up # = docker compose -f infra/docker-compose.yml up -d --build
# 2. Run migrations and seed the demo data
make seed
# 3. Check the API is alive
curl http://localhost:8000/health
# {"status":"ok","service":"shift-rescue-backend","environment":"local"}
# 4. Run the dashboard
make dev-frontend # = cd frontend && pnpm dev → http://localhost:5173Windows note: make is unavailable on plain Windows shells — open each
Makefile target and run the underlying command, or use WSL.
docker compose -f infra/docker-compose.yml --profile observability up -d
# Langfuse on http://localhost:3000 (wired with the evals-observability feature)The demo runs on a single EC2 instance (Docker Compose + Caddy with automatic
TLS) with images from ECR, secrets in SSM Parameter Store and GitHub Actions
authenticated by OIDC. Observability exports to Langfuse Cloud, so no
Langfuse/ClickHouse containers live on the box (see
ADR-003 and docs/runbook.md).
| Component | Where |
|---|---|
| Dashboard | https://<domain>/ (SSR-free SPA behind Caddy) |
| API + Twilio webhooks | https://<domain>/api, https://<domain>/webhooks/twilio/* |
| Workers | EC2 containers worker + beat |
| Data | PostgreSQL 16 + Redis on the instance (EBS volume) |
| Traces | Langfuse Cloud (OTLP with LANGFUSE_*) |
Deploy from GitHub: Actions → Deploy demo → Run workflow (or push to
main). Rollback: re-run the remote script with a previous image tag —
./deploy/remote-deploy.sh <commit-sha>.
Docs: runbook · eval report · demo script · Twilio sandbox setup.
backend/ FastAPI + Celery + SQLAlchemy (async) — the agent service
frontend/ React 19 + Vite — the manager dashboard (kanban + rescue detail)
infra/ docker-compose stack (postgres, redis, api, worker, beat, langfuse*)
docs/ specification, ADRs, assumptions
odd/tasks/ ODD feature documents (one per feature, mirrored in Engram)
| Command | What it does |
|---|---|
make test |
Backend pytest + frontend vitest |
make lint |
ruff + mypy (strict) + oxlint |
make dev-backend |
FastAPI with reload on :8000 |
make dev-frontend |
Vite dev server on :5173 |
Demo credentials (created by make seed): manager@laterraza.demo
(password auth lands with the manager-dashboard auth slice).
- System design — the complete walkthrough: problem, domain, rules, agent, architecture, evaluation, operations, deployment
- Product & engineering specification
- ADR-001: foundation stack
- ADR-002: Strands without the autonomous loop
- ADR-003: demo deployment on AWS
- Evaluation report · Runbook · Demo script
- Twilio sandbox setup · Assumptions log
TBD before public release.



