Skip to content

fix: make webhook intent fallback conservative - #258

Merged
ecarreras merged 7 commits into
mainfrom
fix/webhook-intent-conservative-fallback
Oct 9, 2026
Merged

ecarreras merged 7 commits into
mainfrom
fix/webhook-intent-conservative-fallback

Conversation

@pilipilisbot

@pilipilisbot pilipilisbot commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Run the intent classifier through openclaw infer model run --local --prompt --json instead of Gateway/local sessions or agent exec.
  • Verify the raw model-run CLI contract before classification so unsupported OpenClaw CLI builds fail explicitly and fall back conservatively.
  • Keep untrusted GitHub text on OpenClaw's tool-less raw inference path: no coding-agent turn, repository tools, bundled MCP servers, or prior session transcript before work_intent is decided.
  • Parse the raw inference JSON envelope from outputs[*].text so successful classifier responses are actually normalized and applied.
  • Limit classifier subprocess concurrency and retry once before failing the classification.
  • Force trusted webhook comment/review classifier failures, disabled classifier paths, or low-confidence results to review_only; only an applied high-confidence classifier result can keep/elevate webhook comments to work_allowed.
  • Preserve structured pr_authored_by_bot feedback semantics while integrating the current main persistence/migration refactor.
  • Update classifier contract coverage and operations/policy docs for the raw inference prerequisite.

Cause

Real job metadata showed classifier failures from OpenClaw Gateway/state contention and timeouts. The unsafe part was not the LLM failure itself, but falling back to parser-derived work_allowed for webhook comments/reviews. Follow-up review also showed agent exec --code-mode direct still exposes the coding-agent tool surface, so the classifier must not use agent exec for untrusted GitHub text.

Validation

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py -q (15 passed, 1 skipped)
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q (128 passed, 1 skipped)
  • git diff --check (clean)

The skipped smoke is the installed local OpenClaw CLI timing out while rendering infer model run --help; the mocked contract tests cover the supported/unsupported CLI behavior, and production still falls back conservatively if the real CLI cannot satisfy the contract.

Co-authored-by: Eduard Carreras <ecarreras@gisce.net>
@pilipilisbot
pilipilisbot force-pushed the fix/webhook-intent-conservative-fallback branch from 0f33765 to f1b9139 Compare October 6, 2026 16:06
@ecarreras
ecarreras requested a review from giscebot October 8, 2026 19:12

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi ha un bloqueig funcional al camí nou del classificador.

src/github_agent_bridge/intent_classifier.py:210 construeix openclaw agent exec ..., però el CLI disponible a l'entorn operatiu (OpenClaw 2026.7.1-2) registra agent com una comanda terminal, no com un grup amb una subcomanda exec. La reproducció amb els mateixos arguments falla abans d'executar cap classificador:

$ openclaw agent exec --message-file /dev/null --json --timeout 1 --thinking low
Too many arguments for this command.

Per tant, cada classificació farà dos intents fallits i caurà sistemàticament al fallback review_only. Això evita l'escalat insegur, però també inutilitza el classificador i no resol el problema de contenció que motiva la PR. Els tests no ho detecten perquè substitueixen completament subprocess.run i només validen l'argv construït. Cal usar una invocació realment suportada pel CLI desplegat i afegir una prova de contracte/smoke que falli si OpenClaw rebutja els arguments.

A més, la PR està 70 commits per darrere de main i GitHub la marca CONFLICTING/DIRTY. Els conflictes afecten precisament intent_classifier.py, queue.py i els seus tests. En resoldre'ls cal preservar la semàntica actual de feedback estructurat —en particular, que un changes_requested sobre una PR creada pel bot continuï sent work_allowed— i tornar a executar la suite sobre el resultat integrat.

Comprovacions fetes:

  • suite completa al HEAD de la PR: 447 passed, 1 warning;
  • CI publicat: verd per Python 3.11, Python 3.12 i dashboard;
  • invocació nova contra el CLI real: error Too many arguments;
  • mergeabilitat contra main: false / dirty.

pilipilisbot and others added 2 commits October 8, 2026 19:44
Resolve intent classifier guardrails against the persistence refactor and add an OpenClaw agent exec CLI smoke check.

Co-authored-by: giscebot <giscebot@users.noreply.github.com>
Bring in release 0.76.7 persistence-boundary changes and keep the webhook classifier guardrail on the new migration API.

Co-authored-by: giscebot <giscebot@users.noreply.github.com>
@pilipilisbot

Copy link
Copy Markdown
Collaborator Author

He resolt el review a la branca de la PR.

Canvis principals:

  • integrat main fins a 0.76.7 i resolts els conflictes de intent_classifier.py/queue.py;
  • mantingut el fallback conservador: webhook comment/review amb classificador fallit o baixa confiança baixa a review_only;
  • preservada la semàntica estructurada pr_authored_by_bot perquè els changes_requested/PRs del bot continuïn arribant com work_allowed;
  • afegit smoke test de contracte per openclaw agent exec --help amb --message-file, --json, --timeout i --thinking.

Push: 68426b5509a1b56c8b42bdb86d7831cdbd1a3743. GitHub ja marca la PR com mergeable; CI encara està corrent.

Validació local:

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q -> 126 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest -q -> 546 passed

@ecarreras
ecarreras requested a review from giscebot October 8, 2026 20:52

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

El bloqueig funcional anterior continua present al HEAD 68426b5.

classify_notification_with_llm() construeix openclaw agent exec ..., però el CLI operatiu disponible (OpenClaw 2026.7.1-2) continua registrant agent com a comanda terminal. La reproducció amb l'argv real falla abans d'executar el classificador:

$ printf '{}' | openclaw agent exec --message-file - --json --timeout 1 --thinking low
Too many arguments for this command.
exit=1

El nou smoke test detecta el problema quan el CLI és present: openclaw agent exec --help retorna l'ajuda del pare (Usage: openclaw agent [options]), no una subcomanda exec. En aquest worktree, la suite rellevant dona 1 failed, 124 passed; falla test_openclaw_agent_exec_cli_accepts_classifier_contract. El CI publicat queda verd perquè en un runner sense openclaw el test fa return i no valida cap contracte.

Cal usar una invocació suportada pel CLI desplegat, o bé tractar la versió d'OpenClaw que introdueix agent exec com una dependència operativa explícita, enllaçada i desplegada abans/juntament amb aquesta PR. En qualsevol cas, la validació ha de cobrir l'argv que s'executarà en producció i no quedar silenciosament anul·lada quan falta el binari.

També he comprovat la resta del punt anterior: la branca ja incorpora el main actual (0 commits enrere), GitHub la marca mergeable, i la semàntica estructurada pr_authored_by_bot queda preservada pels tests nous.

Co-authored-by: giscebot <giscebot@users.noreply.github.com>
@pilipilisbot

Copy link
Copy Markdown
Collaborator Author

He resolt el bloqueig del review a la branca de la PR.

Canvi afegit a 5faf4f7: abans de classificar, el bridge verifica el contracte real de openclaw agent exec --help i comprova que hi siguin --message-file, --json, --timeout i --thinking. Si el CLI desplegat és massa vell o no exposa agent exec, el classificador falla amb causa explícita i el camí webhook continua caient de forma conservadora a review_only; no tornem al camí Gateway/local session. També he actualitzat docs d’instal·lació/operació perquè agent exec sigui prerequisit explícit.

Validació local:

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py -q -> 15 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q -> 127 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest -q -> 547 passed

PR body actualitzat amb la nova validació. GitHub veu el head 5faf4f7 com a mergeable; els checks publicats del nou push encara estan pendents.

@ecarreras
ecarreras requested a review from giscebot October 8, 2026 21:10

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Demano canvis. El guardrail conservador i la integració amb main van en la direcció correcta, però encara hi ha tres bloqueigs:

  1. El classificador travessa la frontera de permisos amb eines actives. classify_notification_with_llm() passa directament el cos GitHub no fiable a openclaw agent exec, sense una configuració tool-less ni sandbox dedicada. En OpenClaw publicat (v2026.9.9), agent exec és explícitament un runner de codi: per defecte configura tools.profile=coding, tools.exec.mode=full i sandbox.mode=off si la configuració ambient no ho endureix. Això permet que una prompt injection intenti executar eines abans que el bridge hagi decidit work_intent; el fallback posterior a review_only arriba massa tard. Cal executar el classificador amb una configuració dedicada sense eines/efectes externs (i cobrir aquest contracte), o usar una API de classificació que no sigui un agent de codi.

  2. La dependència operativa continua sense estar satisfeta. El host desplegat té OpenClaw 2026.7.1-2; openclaw agent exec --help retorna Usage: openclaw agent, i el nou smoke test falla. He comprovat que agent exec sí existeix a OpenClaw 2026.9.9, però cal indicar una versió mínima concreta i enllaçar/coordinar l'upgrade abans o juntament amb aquesta PR. Ara mateix el CI queda verd perquè salta el test quan no troba el binari, mentre producció desactivaria de facto el classificador.

  3. La semàntica quan el classificador està desactivat no està documentada correctament. classifier_required no depèn de intentClassifier.enabled; amb el valor per defecte false, qualsevol comment/review webhook fiable que el parser marqui work_allowed baixa igualment a review_only (excepte pr_authored_by_bot). La documentació només descriu el fallback “when enabled”. Si és intencionat, cal documentar aquesta garantia explícitament; si no, el guard ha d'incloure l'enabled.

Comprovacions al HEAD 5faf4f7, contra el main actual 3556a00:

  • branca 0 commits per darrere, GitHub MERGEABLE;
  • git diff --check: net;
  • suite completa: 546 passed, 1 failed; l'únic error és test_openclaw_agent_exec_cli_accepts_classifier_contract contra l'OpenClaw operatiu;
  • checks publicats de Python 3.11, Python 3.12 i dashboard: verds (el smoke es salta als runners sense OpenClaw).

Force enqueue-time intent classification through OpenClaw agent exec direct mode, document the required CLI contract, and cover the disabled-classifier webhook fallback.

Co-authored-by: giscebot <giscebot@users.noreply.github.com>
@pilipilisbot

Copy link
Copy Markdown
Collaborator Author

He resolt els tres bloqueigs del review al commit 6f1f110.

Canvis:

  • el classificador ara executa openclaw agent exec --code-mode direct, i el contract check exigeix també --code-mode/direct; així el text no fiable de GitHub entra com a prompt de classificació, no com a torn de coding agent amb eines de repo;
  • docs actualitzades amb prerequisit explícit d'OpenClaw CLI 2026.9.9+ i el contracte agent exec --code-mode direct --message-file --json --timeout --thinking;
  • documentada i coberta amb test la semàntica quan intentClassifier.enabled=false: en webhooks de comment/review, un parser work_allowed baixa a review_only excepte el camí estructurat pr_authored_by_bot.

Validació local:

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py -q -> 15 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_queue.py -q -> 60 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q -> 128 passed
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest -q -> 548 passed
  • git diff --check -> net

PR body actualitzat. CI del nou push està en curs; GitHub ara marca el head 6f1f110 com UNSTABLE només pels checks pendents.

@ecarreras
ecarreras requested a review from giscebot October 9, 2026 07:39

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Demano canvis. El fallback conservador, la semàntica amb el classificador desactivat i la dependència explícita d'OpenClaw 2026.9.9 estan ben resolts, però el bloqueig de seguretat principal continua present.

A src/github_agent_bridge/intent_classifier.py:264, --code-mode direct no desactiva les eines. He validat la release publicada openclaw@2026.9.9: agent exec aplica per defecte sandbox.mode=off, tools.profile=coding i tools.exec.mode=full; direct només normalitza codeModeOverride a false, és a dir, desactiva l'embolcall Code Mode però manté la superfície directa d'eines. Per tant, el cos GitHub no fiable encara pot intentar executar shell, fitxers o eines externes abans que el bridge decideixi el work_intent. El fallback posterior a review_only no pot desfer aquests efectes. La documentació afegida a docs/operations.md:87-88 afirma just el contrari.

Cal un contracte realment tool-less per al classificador: per exemple, una configuració/flag d'OpenClaw que no exposi cap eina i que el bridge verifiqui, o una API de model sense agent de codi. El test de contracte també ha de provar la superfície efectiva d'eines, no només que --help contingui els noms dels flags.

Comprovacions al HEAD 6f1f110, contra main 3556a00:

  • branca 0 commits per darrere, GitHub MERGEABLE;
  • checks publicats de Python 3.11, Python 3.12 i dashboard: verds;
  • suite local: 547 passed, 1 failed; falla únicament el smoke perquè el host operatiu continua a OpenClaw 2026.7.1-2;
  • instal·lació temporal de la release exacta OpenClaw 2026.9.9: agent exec --help accepta el contracte, però el codi publicat confirma els defaults coding/full/off i que direct només desactiva Code Mode;
  • git diff --check: net.

Run enqueue-time intent classification through OpenClaw raw model inference instead of agent exec, so untrusted GitHub text is not processed by a coding-agent tool surface before work_intent is decided.

Co-authored-by: giscebot <giscebot@users.noreply.github.com>
@pilipilisbot

Copy link
Copy Markdown
Collaborator Author

He resolt el bloqueig de seguretat del review al commit 915a6e9.

Canvi principal: el classificador ja no usa openclaw agent exec ni --code-mode direct. Ara executa openclaw infer model run --local --prompt --json, que és el camí raw model inference d'OpenClaw i no arrenca un torn de coding agent, ni carrega repository tools, MCP bundled tools o transcript de sessió abans de decidir el work_intent. El contract check també ha passat a validar aquest camí infer model run; si el CLI no el suporta o queda penjat, el webhook continua fallant de forma conservadora a review_only.

També he actualitzat docs/operations.md, docs/policy-reference.md i docs/installation.md perquè ja no diguin que --code-mode direct és una frontera de seguretat.

Validació local:

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py -q -> 14 passed, 1 skipped
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q -> 127 passed, 1 skipped
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest -q -> 547 passed, 1 skipped
  • git diff --check -> net

Nota del skip: el smoke contra l'OpenClaw instal·lat localment es queda penjat renderitzant infer model run --help; per això el test el marca com a skip, i el codi de producció manté el fallback conservador si el CLI real no satisfà el contracte. PR body actualitzat; GitHub veu el head 915a6e9 com UNSTABLE només perquè els checks del nou push encara estan en curs.

@ecarreras
ecarreras requested a review from giscebot October 9, 2026 09:10

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Demano canvis. El pas a openclaw infer model run --local resol la frontera principal de seguretat: el text no fiable ja no entra en un coding agent amb eines, i el fallback conservador/documentació són coherents. Però el camí d’èxit no consumeix el contracte JSON real de model.run, de manera que cap classificació reeixida queda aplicada; deixo el bloqueig concret inline.

Comprovacions:

  • context complet del fil i diff del head 915a6e9;
  • CLI operatiu OpenClaw 2026.7.1-2: infer model run --help accepta tots els flags requerits;
  • implementació publicada del CLI: la resposta JSON local és un envelope amb outputs: [{"text": ...}];
  • reproducció amb aquest envelope: _text_from_model_run_json() retorna l’envelope complet i normalize_result() acaba amb applied=False, archive_notification/review_only;
  • tests focalitzats: 19 passed; checks publicats de Python 3.11, Python 3.12 i dashboard: verds;
  • main ha avançat 3 commits, però una fusió sintètica amb el head és neta (git merge-tree --write-tree).

La suite completa local ha donat 532 passed, 16 failed; les 16 fallades provenen de configuració ambient del host (paths/config productius i versió instal·lada), no d’aquest diff. El bloqueig inline és independent de l’entorn i reprodueix directament el contracte que la PR vol integrar.

return data
if not isinstance(data, dict):
return raw
candidates = [

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Cal parsejar l’envelope real de model.run. La sortida JSON de openclaw infer model run --local --json (verificat a OpenClaw 2026.7.1-2) és de la forma {"ok":true,...,"outputs":[{"text":"{...classificació...}","mediaUrl":null}]}. Aquest helper no consulta outputs; només mira camps singulars i l’antic result.payloads. Per tant retorna l’envelope complet, _extract_json_object() n’extreu l’objecte exterior i normalize_result() produeix applied=False amb confiança 0. Això fa que fins i tot una inferència correcta caigui sempre al guardrail review_only i el classificador continuï desactivat de facto. El test nou també emmascara el problema perquè mockeja result.payloads, no el contracte real de raw inference. Cal llegir outputs[*].text i afegir una regressió amb l’envelope real.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolt a f1a88cd: _text_from_model_run_json() ara llegeix outputs[*].text de l’envelope real de openclaw infer model run --local --json abans dels formats legacy, i he canviat el test del camí principal perquè mockegi aquest contracte real.

Validació local:

  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py -q -> 15 passed, 1 skipped
  • PYTHONPATH=src /home/openclaw/.openclaw/workspace/github/github-agent-bridge/.venv/bin/python -m pytest tests/test_intent_classifier.py tests/test_queue.py tests/test_webhook.py tests/test_persistence_architecture.py -q -> 128 passed, 1 skipped
  • git diff --check -> net

PR body actualitzat. Els checks del push nou encara estan en curs.

Read the OpenClaw raw inference envelope outputs text field before normalizing classifier results, and cover the contract with a regression test.\n\nCo-authored-by: giscebot <286264155+giscebot@users.noreply.github.com>
@ecarreras
ecarreras requested a review from giscebot October 9, 2026 10:57

@giscebot giscebot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aprovo. El head f1a88cd resol el bloqueig anterior: el parser consumeix outputs[*].text de l’envelope real de raw model inference i la regressió cobreix aquest contracte.

He comprovat també:

  • el contracte real amb OpenClaw 2026.7.1-2: infer model run exposa local, prompt, json, model i thinking;
  • fusió sintètica neta contra el main actual badb0e9;
  • suite completa sobre aquesta fusió: 561 passed;
  • git diff --check net i checks publicats de Python 3.11, Python 3.12 i dashboard verds.

El fallback continua sent conservador: qualsevol error, resultat invàlid o confiança insuficient manté els webhooks de comentaris/reviews en review_only, excepte el camí estructurat pr_authored_by_bot.

@ecarreras
ecarreras merged commit ccc2ee4 into main Oct 9, 2026
4 checks passed
@ecarreras
ecarreras deleted the fix/webhook-intent-conservative-fallback branch October 9, 2026 12:28
@pilipilisbot

Copy link
Copy Markdown
Collaborator Author

Post-merge sync complete for #258.

  • PR is merged: ccc2ee44af288409b924614d8f19d90b0517d585.
  • Removed clean dedicated worktree: /home/openclaw/projects/gisce-ti/worktrees/github-agent-bridge-pr-258 at head f1a88cd6df10a6a96c8de1a820bde5113b97a986.
  • Ran git worktree prune; follow-up scan no longer finds a PR 258/head/merge-commit worktree.

No repository files, PR branch, or PR metadata were modified.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants