Feature/webhook offload - #21
Merged
Merged
Conversation
…here Spec §7.5 asks the inbound Twilio webhook to answer in under 200 ms and never call the LLM inside it. Wiring the provider exposed that the webhook awaited the interpretation, taking 1-3 s per message. - app/runtime.py builds the worker's rescue runtime once per process (channel, scheduler, workforce, orchestrator, interpreter); the API process no longer builds any of it. - process_inbound_message runs the async handling in the worker with bounded retries; run_due_jobs ticks the scheduler and publishes the degraded-status snapshot to Redis for the API probe; purge_old_messages is finally scheduled. - The webhook validates the signature and enqueues, returning TwiML 200 (500 on broker rejection so Twilio retries instead of silently dropping a message). - /api/status reads configuration (new pure is_provider_configured helper) plus the published snapshot instead of another process's internals. - The FastAPI lifespan ticker is gone: beat owns time. - Tracing bootstrap: the worker installs its own TracerProvider per preforked child and flushes on shutdown. Without it the LLM calls that now happen in the worker produced no Langfuse traces at all (found live, verified fixed: four observations exported for one message). Verified live: webhook 15-50 ms, worker ~2.9 s per message off the request path, two children concurrently without loop errors, beat ticking every 5 s. 296 tests pass, ruff and mypy clean.
…ds with prompt v2 The first real-model run exposed three defects, and the most serious one was the gate itself. - check_thresholds() compared hardcoded '..._min' keys against a report that stores the metrics without the suffix, so every lookup returned None, every comparison was skipped and the runner printed 'Thresholds met.' while accuracy was 0.86 against a 0.92 minimum. evals/thresholds.yaml was never read at all. Thresholds now live in app/evals/thresholds.py, read the YAML (one source of truth, min and max directions) and fail closed: an unmapped key, an unknown metric or a non-numeric value is a violation, never a silent pass. - _build_prompt forwarded rescue_id, pending_offers and shifts_48h but dropped pending_confirmation, so a bare '1' had no way to be read as ABSENCE_CONFIRM. It now renders that key and any remaining non-empty context key, with a credential-name filter so a secret can never reach a prompt. - New interpreter_v2 prompt with an explicit procedure keyed on what is pending, numeric replies meaning yes/no only when something is pending, a broader health rule and examples for the confusions v1 exposed. The active version is named once and recorded on every interpretation. Measured with the real provider on the 150-sample golden set: v1: 0.86 accuracy / 0.9167 health / 0.95 times (2 violations, gate silent) v2: 0.9333 accuracy / 1.0 health / 0.85 times (thresholds met, honestly) The 10 remaining failures are analysed in odd/tasks/interpreter-quality.md; six of them need the orchestrator to pass an 'already accepted' marker because the context the harness supplies is identical to cases labelled OFFER_DECLINE. 308 tests pass, ruff and mypy clean. Eval run artifacts are now gitignored.
The test failed in CI and passed locally: build_runtime(settings) ignored the settings it was given for the database and fell back to the ambient .env, so the test connected to the developer's local Postgres on 5433 and found nothing to connect to on a clean runner (OSError, 111). The runtime now passes settings.database_url to the engine factory, and the test builds a temp-file SQLite database with the schema, so it is hermetic on any machine. Reproduced and verified with the Linux CI image (ghcr.io/astral-sh/uv:python3.12-bookworm): 308 passed before the fix, 308 passed and no failures after it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.