Skip to content
Merged

Dev #23

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
7a59380
feat(llm): wire the LLM interpreter and Langfuse tracing into the run…
aam9063 Sep 25, 2026
eef1b89
docs(llm): ADR-004 provider selection, runbook configuration and task…
aam9063 Sep 25, 2026
ecd1c34
test(tracing): live Langfuse verification helper and clearer llm_path…
aam9063 Sep 25, 2026
9a75d6c
docs(llm): record verification evidence for the runtime wiring
aam9063 Sep 25, 2026
08df407
fix(llm): two defects the live provider call revealed
aam9063 Sep 25, 2026
c549414
chore(ops): ground the interpreter latency target in measurements
aam9063 Sep 25, 2026
7e978c3
docs: task record for the webhook offload (spec 7.5 deviation)
aam9063 Sep 25, 2026
d7f9986
feat(workers): offload inbound orchestration to Celery and trace it t…
aam9063 Sep 25, 2026
1d38906
fix(evals): make the accuracy gate fail closed and close the threshol…
aam9063 Sep 25, 2026
773503a
docs: task record for interpretation persistence and the accepted-off…
aam9063 Sep 25, 2026
e0d8f1c
Merge pull request #20 from aam9063/feature/llm-runtime-wiring
aam9063 Sep 25, 2026
4ef955a
feat(llm): persist interpretations, pass the accepted-offer marker, p…
aam9063 Sep 25, 2026
a6a4678
fix(test): stop the worker runtime test from using the ambient database
aam9063 Sep 25, 2026
cae43c1
Merge pull request #21 from aam9063/feature/webhook-offload
aam9063 Sep 25, 2026
8542123
feat(api): dashboard REST API with JWT auth (spec §7.5)
aam9063 Sep 25, 2026
6f417c8
feat(web): connect the dashboard to the live API
aam9063 Sep 25, 2026
5bee857
fix(workers): one event loop per worker process
aam9063 Sep 25, 2026
5a589f6
fix(orchestrator): audit the real case id and tell the model a confir…
aam9063 Sep 25, 2026
1e4e15f
feat(workers): timers owned by the broker, plus a reconciliation sweep
aam9063 Sep 25, 2026
36c0eff
feat(api): demo-only simulator endpoints and a shared demo clock
aam9063 Sep 25, 2026
4240100
feat(agent): answer the questions the agent asks, and escalate the ghost
aam9063 Sep 25, 2026
1277032
feat(web): make the dashboard usable on phones and tablets
aam9063 Sep 25, 2026
5ac3b6d
fix(web): do not retry client errors, and format the decision time
aam9063 Sep 25, 2026
cdd2de9
fix(demo): route /dev through the proxy and make the acceptance-race …
aam9063 Sep 25, 2026
2db88e2
feat(web): make the demo self-explanatory and turn Today into a dashb…
aam9063 Sep 25, 2026
6af40c2
Merge remote-tracking branch 'origin/dev' into feature/dashboard-live
aam9063 Sep 28, 2026
3515ab8
fix(agent): answer with the employee's real state, and persist the of…
aam9063 Sep 28, 2026
a9e0b30
feat(web): let the manager act, and show the agent's reply without a …
aam9063 Sep 28, 2026
9429308
fix(agent): stop stale offers from faking approvals, and record every…
aam9063 Sep 28, 2026
dcad17f
feat(api): one-click demo reset
aam9063 Sep 28, 2026
38c0059
fix(web): show every phone frame, not just the first three
aam9063 Sep 28, 2026
c476212
feat(demo): bound the demo clock and make a shifted clock impossible …
aam9063 Sep 28, 2026
626c267
fix(web): the agent's reply shows up on the employee's first message too
aam9063 Sep 28, 2026
8bcb261
feat(evals): record evaluation runs and show the real ones in the das…
aam9063 Sep 28, 2026
05a7c8b
fix(tests): stop the clock tests from talking to a real broker, and d…
aam9063 Sep 28, 2026
ff96627
Merge pull request #22 from aam9063/feature/dashboard-live
aam9063 Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -36,3 +36,6 @@ Desktop.ini
# Local editor state
.idea/
.vscode/
# Eval runs: local artifacts, one per run (sizes and model outputs)
evals/reports/

24 changes: 24 additions & 0 deletions DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -898,3 +898,27 @@ When refining existing screens generated with this design system:

- Starbucks Visa Card / Starbucks-Card (SVC) detailed mockup specs are hinted at by `--svcRoundedCorners` and `--svcShadowFilter` tokens but not fully documented


## Appendix A. Dashboard adoption of the responsive contract (§8)

The Shift Rescue dashboard (frontend) adopts §8 as follows. §8 remains the
contract; this appendix only records what is implemented.

- **Breakpoints.** Tailwind's default scale maps onto §8: `md` (768px) is the
tablet breakpoint, `lg` (1024px) desktop, `xl` (1280px) toward xlarge. No
custom breakpoints and no JavaScript media queries: variants are
CSS-controlled classes (`md:hidden`, `hidden md:flex`, ...).
- **Navigation.** Below the tablet breakpoint the desktop navs
(`hidden md:flex` / `hidden lg:flex`) are replaced by a `md:hidden`
hamburger drawer listing every destination from both groups, with the same
gold active indicator on the House Green band, a gold focus ring,
`Escape`-to-close and close-on-navigate.
- **Wide data.** Tables keep their desktop rendering from `md` up and gain a
stacked card list below `md`, rendered as CSS-controlled siblings
(`md:hidden` / `hidden md:block`); every table container carries
`overflow-x-auto` as a safety net.
- **Touch targets.** Pills and actions reach the 44px floor on touch surfaces
via the `pointer-coarse:` variants (`pointer-coarse:min-h-11`), leaving the
desktop look untouched; drawer items are always 44px (`min-h-11`).
- **Gutters.** 16 -> 24 -> 40px (`px-4` -> `md:px-6` -> `lg:px-10`), matching
the existing header and rescue-detail padding.
38 changes: 24 additions & 14 deletions backend/.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -9,17 +9,23 @@ JWT_SECRET=dev-only-secret
DATABASE_URL=postgresql+asyncpg://shift_rescue:shift_rescue@localhost:5433/shift_rescue
REDIS_URL=redis://localhost:6379/0

# --- LLM (spec §7.2) ---------------------------------------------------------
LLM_PROVIDER=anthropic # anthropic | bedrock | nan
ANTHROPIC_API_KEY=
LLM_MODEL_INTERPRETER=claude-haiku-4-5-20251001
LLM_MODEL_COMPOSER=claude-sonnet-5
LLM_MODEL_SUMMARIZER=claude-sonnet-5
# Optional per-piece provider override (send only the interpreter to NaN):
# LLM_PROVIDER_INTERPRETER=nan
# NAN_API_KEY=
# NAN_BASE_URL=https://api.nan.builders/v1
# LLM_MODEL_INTERPRETER_NAN=nan/deepseek-v4-flash
# --- LLM (spec §6, ADR-004) --------------------------------------------------
# Fail-closed: provider "none", a missing key or a missing provider SDK makes
# the API answer with the deterministic parser and log `llm_disabled` — it
# never blocks a rescue. A stale provider variable is ignored (extra="ignore").
LLM_PROVIDER=openai # openai | anthropic | bedrock | none
OPENAI_API_KEY=... # required for provider=openai
# OPENAI_BASE_URL= # any OpenAI-compatible gateway (e.g. NaN)
ANTHROPIC_API_KEY=... # only for provider=anthropic
AWS_REGION=eu-west-1 # only for provider=bedrock (instance role)
LLM_MODEL_INTERPRETER= # empty = provider default (gpt-4o-mini)
LLM_TEMPERATURE=0.0
LLM_MAX_TOKENS=500
LLM_TIMEOUT_SECONDS=10
LLM_CONFIDENCE_THRESHOLD=0.75
# Optional cost overrides, USD per 1K tokens (0 = provider default):
# LLM_PRICE_INPUT_PER_1K=0.00015
# LLM_PRICE_OUTPUT_PER_1K=0.0006

# --- Twilio WhatsApp (docs/twilio-sandbox-setup.md) --------------------------
TWILIO_ACCOUNT_SID=
Expand All @@ -28,10 +34,14 @@ TWILIO_WHATSAPP_FROM=whatsapp:+14155238886
TWILIO_VALIDATE_SIGNATURE=true

# --- Observability (Langfuse Cloud — no self-hosting, see docs/assumptions.md A4)
OTEL_EXPORTER_OTLP_ENDPOINT=https://cloud.langfuse.com/api/public/otel/v1/traces
LANGFUSE_PUBLIC_KEY=
LANGFUSE_SECRET_KEY=
# Both keys present -> OTLP/HTTP spans to
# <LANGFUSE_HOST>/api/public/otel/v1/traces with Basic auth. Keys absent ->
# tracing is a no-op and the API logs `tracing_disabled`.
LANGFUSE_PUBLIC_KEY=...
LANGFUSE_SECRET_KEY=...
LANGFUSE_HOST=https://cloud.langfuse.com
# Explicit override for any OTLP collector (wins over the derived Langfuse URL):
# OTEL_EXPORTER_OTLP_ENDPOINT=
SENTRY_DSN=

# --- Demo / development ------------------------------------------------------
Expand Down
214 changes: 214 additions & 0 deletions backend/app/agent/factory.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,214 @@
"""LLM provider factory (ADR-004): builds the Strands model and the
`MessageInterpreter` from `Settings`, fail-closed.

Provider SDKs are imported lazily inside the functions, so the API boots
without any provider package installed. `build_interpreter()` never raises on
the API path: when the provider is disabled, a credential is missing or the
SDK is absent it returns `None` and the orchestrator degrades to the
deterministic parser (spec §9.3). Strands is imported only here and in
`app.agent.llm` (ADR-002 boundary); the agent runs WITHOUT tools — the LLM
interprets language, it never mutates state.
"""

from pathlib import Path
from typing import Any

import structlog

from app.agent.interpreter import PROMPT_VERSION, MessageInterpreter
from app.core.config import Settings

logger = structlog.get_logger(__name__)

# USD per 1K tokens (provider defaults; overridable via Settings).
PROVIDER_DEFAULT_MODELS: dict[str, str] = {
"openai": "gpt-4o-mini",
"anthropic": "claude-haiku-4-5",
"bedrock": "eu.anthropic.claude-haiku-4-5-v1:0",
}

PROVIDER_DEFAULT_PRICES: dict[str, dict[str, float]] = {
"openai": {"input": 0.00015, "output": 0.0006},
"anthropic": {"input": 0.0008, "output": 0.004},
"bedrock": {"input": 0.0008, "output": 0.004},
}

# Settings attribute and environment variable holding each provider credential.
_CREDENTIAL_SETTINGS: dict[str, str] = {
"openai": "openai_api_key",
"anthropic": "anthropic_api_key",
}
_CREDENTIAL_ENV_VARS: dict[str, str] = {
"openai": "OPENAI_API_KEY",
"anthropic": "ANTHROPIC_API_KEY",
}


class LLMNotConfiguredError(Exception):
"""Raised when the provider cannot be built (missing credential, unknown
provider). The message names the missing environment variable."""


def _provider(settings: Settings) -> str:
return settings.llm_provider.strip().lower()


def resolve_model_id(settings: Settings) -> str:
"""Configured model id, or the provider default for empty/unknown values."""
if settings.llm_model_interpreter:
return settings.llm_model_interpreter
return PROVIDER_DEFAULT_MODELS.get(_provider(settings), PROVIDER_DEFAULT_MODELS["openai"])


def resolve_price(settings: Settings) -> dict[str, float]:
"""Per-1K USD prices: provider default, overridable per direction when > 0."""
price = dict(
PROVIDER_DEFAULT_PRICES.get(_provider(settings), PROVIDER_DEFAULT_PRICES["openai"])
)
if settings.llm_price_input_per_1k > 0.0:
price["input"] = settings.llm_price_input_per_1k
if settings.llm_price_output_per_1k > 0.0:
price["output"] = settings.llm_price_output_per_1k
return price


def _missing_credential_reason(settings: Settings) -> str:
"""Secret-free reason naming the missing credential variable."""
var = _CREDENTIAL_ENV_VARS.get(_provider(settings))
if var is None:
return f"Unknown llm_provider {_provider(settings)!r}"
return f"{var} is not set — add it to backend/.env"


def is_provider_configured(settings: Settings) -> bool:
"""Pure configuration check (ADR-004): provider enabled **and** its
credential present. Constructs nothing and performs no network call — the
single source of truth for "is the LLM path usable" (spec §9.3).
"""
if not settings.llm_enabled:
return False
provider = _provider(settings)
if provider == "bedrock":
return True # AWS credentials come from the instance role (ADR-004)
attr = _CREDENTIAL_SETTINGS.get(provider)
return bool(attr is not None and getattr(settings, attr))


def _build_openai_model(settings: Settings, model_id: str) -> Any:
from strands.models.openai import OpenAIModel

if not settings.openai_api_key:
raise LLMNotConfiguredError(_missing_credential_reason(settings))
client_args: dict[str, str] = {"api_key": settings.openai_api_key}
if settings.openai_base_url:
client_args["base_url"] = settings.openai_base_url
return OpenAIModel(
model_id=model_id,
params={"max_tokens": settings.llm_max_tokens, "temperature": settings.llm_temperature},
client_args=client_args,
)


def _build_anthropic_model(settings: Settings, model_id: str) -> Any:
# Signature verified for strands 1.56.0: AnthropicModel takes client_args
# (it builds its own AsyncAnthropic client); max_tokens/model_id are
# required config keys, temperature rides in params.
from strands.models.anthropic import AnthropicModel

if not settings.anthropic_api_key:
raise LLMNotConfiguredError(_missing_credential_reason(settings))
return AnthropicModel(
model_id=model_id,
max_tokens=settings.llm_max_tokens,
params={"temperature": settings.llm_temperature},
client_args={"api_key": settings.anthropic_api_key},
)


def _build_bedrock_model(settings: Settings, model_id: str) -> Any:
# Signature verified (strands 1.56.0): BedrockModel takes keyword-only
# region_name plus flat model_config keys (model_id, temperature, max_tokens).
# Credentials come from the instance role; never exercised in unit tests.
from strands.models.bedrock import BedrockModel

return BedrockModel(
model_id=model_id,
temperature=settings.llm_temperature,
max_tokens=settings.llm_max_tokens,
region_name=settings.aws_region,
)


def build_model(settings: Settings) -> Any:
"""Build the Strands model for the configured provider."""
provider = _provider(settings)
model_id = resolve_model_id(settings)
builders = {
"openai": _build_openai_model,
"anthropic": _build_anthropic_model,
"bedrock": _build_bedrock_model,
}
builder = builders.get(provider)
if builder is None:
raise LLMNotConfiguredError(
f"Unknown llm_provider {provider!r} — expected: {', '.join(sorted(builders))}, none"
)
return builder(settings, model_id)


def load_system_prompt(version: str = PROMPT_VERSION) -> str:
"""Interpreter prompt (baked into the image with the app package).

Prompt edits ship as a new versioned file: the version is recorded on every
interpretation, so a quality change is always attributable to a prompt.
"""
return (Path(__file__).parent / "prompts" / f"{version}.md").read_text(encoding="utf-8")


def build_interpreter(settings: Settings) -> MessageInterpreter | None:
"""Fail-closed entry point: `MessageInterpreter` or `None`.

Never raises on the API path; never logs or returns a credential. A `None`
result means the orchestrator answers with the deterministic parser.
"""
if not is_provider_configured(settings):
reason = (
"provider is disabled"
if not settings.llm_enabled
else _missing_credential_reason(settings)
)
logger.warning("llm_disabled", reason=reason)
return None
try:
model = build_model(settings)
except (ImportError, ModuleNotFoundError):
logger.warning("llm_disabled", reason="provider SDK is not installed")
return None

from strands import Agent

from app.agent.llm import StrandsLLMClient
from app.agent.schemas import Interpretation

model_id = resolve_model_id(settings)
client = StrandsLLMClient(
agent_factory=lambda: Agent(
model=model,
system_prompt=load_system_prompt(),
structured_output_model=Interpretation,
callback_handler=None,
),
timeout_seconds=settings.llm_timeout_seconds,
model_id=model_id,
price_per_1k=resolve_price(settings),
)
return MessageInterpreter(
llm=client,
prompt_version=PROMPT_VERSION,
confidence_threshold=settings.llm_confidence_threshold,
)


def describe_provider(settings: Settings) -> str:
"""One-line, secret-free provider/model description for structured logs."""
return f"provider={_provider(settings)} model={resolve_model_id(settings)}"
17 changes: 13 additions & 4 deletions backend/app/agent/interpreter.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,8 @@
from app.agent.schemas import Interpretation
from app.ports import LLMClient

PROMPT_VERSION = "interpreter_v5"

FALLBACK = Interpretation(intent="UNCLEAR", confidence=0.0)


Expand All @@ -24,7 +26,7 @@ def __init__(
self,
llm: LLMClient,
*,
prompt_version: str = "interpreter_v1",
prompt_version: str = PROMPT_VERSION,
confidence_threshold: float = 0.75,
max_retries: int = 1,
) -> None:
Expand All @@ -42,13 +44,20 @@ async def interpret(self, message_body: str, context: dict[str, Any]) -> Interpr
attempt_context = {**attempt_context, "validation_error": last_error}
try:
raw = await self._llm.interpret(message_body, attempt_context)
return Interpretation(
**raw, prompt_version=self.prompt_version
).model_copy()
# A real LLMClient returns the whole structured payload, which
# already carries prompt_version; the interpreter owns it, so it
# must overwrite rather than duplicate the keyword argument.
return Interpretation(**{**raw, "prompt_version": self.prompt_version})
except ValidationError as error:
last_error = str(error)
except Exception as error:
# Provider outage/timeout: degrade to the parser (spec §9.3).
raise ProviderUnavailableError(str(error)) from error

return FALLBACK

@property
def last_usage(self) -> dict[str, Any] | None:
"""Usage dict of the wrapped client's last call, or None if unreported."""
usage = getattr(self._llm, "last_usage", None)
return usage if isinstance(usage, dict) else None
Loading
Loading