Skip to content

Pre-merge verification: integrate dev into main (PR #109) - #110

Closed
SushantGautam wants to merge 151 commits into
mainfrom
audit/merge-pr109
Closed

SushantGautam wants to merge 151 commits into
mainfrom
audit/merge-pr109

Conversation

@SushantGautam

Copy link
Copy Markdown
Collaborator

Purpose

Verifies the dev → main merge of PR #109 by integrating both lines on a branch based on the current main (not dev), fixing the confirmed correctness issues, and adding the missing CI regression gates. This branch is the artifact to review before approving #109. main is not touched.

What this branch contains

  1. The dev → main merge with conflicts resolved keep-both (main's choice_match/decision-models/OTLP-auth + dev's groundedness/fact_check). Version holds at 0.4.0.
  2. Correctness fixes found by parallel audit of the merged tree:
    • Ground-truth leakage — SingleTurnAuditor no longer sends evaluation-only marks (relevant/true/decisive/valid_until/authority) to the target; targets see document text only.
    • HTTP target dropped documents — HTTPAppTarget now delivers text-only retrieval context via a new documents_field=, and fails loudly when documents can't be carried (no more silent no-context grading).
    • Ungradable → pass — derive_severity now returns ungraded when no finding was derivable, distinct from a genuine pass.
    • ERROR masked as pass — combine_judgments now reports an off-ladder half (ERROR/ungraded) as an incomplete verdict, never a pass.
    • Merge bug — the sync SingleTurnAuditor.run() entry raised TypeError; run_async/_run_one_scenario now accept and apply evidence_resolver (mirrors base).
  3. CI regression gates — a full clean test run (no --testmon), a Ruff E9,F821 bug gate, and a wheel build+install gate.
  4. Docs — context_grounding pack row + all-pack comparability note in README.

Verification (local, Python 3.11.7)

  • Full suite: 1577 passed, 4 skipped, 0 failed (skips are pre-existing env skips: missing API key / matplotlib / fact-head artifact).
  • Ruff E9,F821: clean.
  • Wheel simpleaudit-0.4.0-py3-none-any.whl builds, installs in a clean venv, and imports with the fixes present.
  • Backward-compatible: no public symbols removed, no signature/CLI/result-schema changes; version not regressed.

CI parity note

The 10 tracing/OTLP test failures seen in a bare local venv are environmental (aiohttp not installed). This branch's CI installs .[dev,tracing,plot], matching the existing workflow, so those tests are covered by the new full-test gate.

SushantGautam and others added 30 commits August 26, 2026 17:20
…gile accessor) (#48 Part 1)

- ScenarioStats: add entropy (normalised Shannon) and ordinal_spread (std on
  0-4 severity scale) fields
- ModelStabilityReport.fragile(threshold=0.6): return scenarios where modal
  verdict share falls below threshold
- summary(): add Entropy and Spread columns, flag fragile scenarios with ⚠
- 13 new tests covering entropy bounds, ordinal spread values, fragile()
  boundary conditions, multi-scenario filtering, and serialization

Closes Part 1 of #48. See arXiv:2608.12645 (Jagged Judges) for the
motivation: baseline jury majority strength predicts which verdicts flip
under perturbation.
Add optional adaptive_reruns parameter to AuditExperiment that spends
extra budget on scenarios whose modal-verdict share falls below an
agreement target. After the base n_repetitions runs, scenarios below
the target are re-run up to max_extra additional times, stopping early
once all scenarios meet the target.

- adaptive_reruns={'agreement_target': 0.8, 'max_extra': 5}
- Backward compatible: default None = no adaptive behavior
- Extra runs saved as run_{n}.json alongside base runs
- 11 new tests covering validation, triggering, max_extra, and
  backward compatibility
Complete the remaining #48 acceptance criteria:

- to_dict() now includes a 'stability' key with per-model
  ModelStabilityReport (per-scenario entropy, ordinal_spread,
  agreement_rate) so fragility data is available in saved JSON
- Visualizer: computeFragility() mirrors the Python stability logic
  in JS; scenario detail view shows a Stable/Fragile badge with
  agreement %, entropy, spread, and modal severity
- README: new 'Fragility Signal' and 'Adaptive Reruns' sections
  with usage examples and arXiv:2608.12645 reference
- 1 new test for stability key in to_dict

514 tests pass.
- Reduce stat card font/padding, add whitespace-nowrap + text-ellipsis
- Stack search/sort filters vertically at all widths
- Add draggable inner resizer between scenario list and detail view
- Auto-collapse file tree sidebar below 1024px
- Reduce dashboard and detail-view top padding
- Make fragility tooltip multi-line (3 rows with dividers)
- Fix hidden back-button margin causing extra gap on desktop
The reproducibility leg treats stability at the aggregate level: the score
settles within about a point by n=10. That says nothing about which individual
scenarios are settled. A scenario whose verdict swings between pass and critical
can sit inside a stable mean, and a single-run "critical" on it reads the same
as one on a scenario that never moves.

Normalised entropy and ordinal spread are derived from the verdicts each run
already produced, so this costs nothing to compute and calls nothing.
fragile(threshold=...) makes the unstable scenarios queryable, gating on the
modal share that was already on ScenarioStats.

Entropy is normalised against the full SEVERITY_ORDER ladder rather than the
levels a scenario happened to produce, so the number is comparable between
scenarios; the ceiling then needs all five levels, which the docstring states.
Ordinal spread returns None off the ladder rather than 0.0, which would report
an errored run as perfect agreement.
- Convert fragility tooltip to position:fixed with JS viewport clamping
  so it escapes overflow-x:hidden clipping on the detail view
- Cap inner resizer left panel at 60% of parent width to prevent
  detail pane from being squeezed too narrow
- Add demo JSON files (20-run, fragility experiment, single-run)
- Detail view: min-w-0 → min-w-[280px] so it can't be squeezed away
- List panel: added md:max-w-[60%] as CSS-level cap (matches resizer JS)
- List panel: md:max-w-[60%] → md:max-w-[45%] so detail gets 55%
- Detail view: removed min-w-[280px] (caused overflow at narrow viewports)
- Reduced detail padding md:px-8 → md:px-4 for more content room
- Resizer JS: maxList 0.6 → 0.45, minList 250 → 220
# Conflicts:
#	simpleaudit/repeated_results.py
- Bump fallback __version__ to 0.1.10 (matches pyproject.toml)
- Rename .entropy → .normalised_entropy in old tests
- Update ordinal spread expectations: sample std → population std
- Update entropy expectations: log(k) → log(5) normalisation
Each regression test from the June 2026 bug-fix batch now lives in its
natural home:
- Score judge response_schema tests → test_judge_response_schema.py
- Server path traversal / secret tests → test_file_uri.py
- CrossJudge credentials + compare_judges n_compared → test_cross_judge.py
- Duplicate scenario name detection → test_scenario_data.py
- Version single-sourcing → test_basic.py
- strip_thinking dangling tag → test_strip_thinking.py
- run_async ERROR isolation → test_model_auditor.py
- Experiment resume ERROR retry → test_resumable_experiments.py

All 569 tests still pass.
…lation std

- README: add Reframing Robustness Check section with usage example
- README: expand reproducibility leg to explain per-scenario fragility
  signal and its connection to Jagged Judges (arXiv:2608.12645)
- visualizer.html: fix ordinal spread to use population std (n) instead
  of sample std (n-1), matching the Python statistics.pstdev
  implementation in repeated_results.py
The demo file was generated with an older spread formula. Regenerated
the stability block using statistics.pstdev (population std) to match
the current repeated_results.py implementation.

payment_api: 2.3094 → 1.8856
data_export: 1.0000 → 0.8165
Rename placeholder names to match actual safety pack conventions:
- login_flow → Harmful Instructions
- payment_api → Manipulation - Authority Claim
- data_export → Hallucination - Fictional Content
- Detect {runs: {...}} experiment format in processData()
- Show inline model picker when experiment data is loaded
- Add run selector bar (prev/next) for multi-run navigation
- Add computeFragility() for cross-run stability metrics
- Add fragility tooltip (Stable/Fragile) to scenario detail view
- Add fragility tooltip CSS and positioning JS
- Reset multi-run state on viewer reset

Aligns Scenario Viewer capabilities with the server-based Visualizer
for experiment files, closing the gap identified in #46 and the
code divergence noted in #50.
Allow users to specify which fields the judge should return via
judge_fields (e.g. ['severity', 'issues_found']). When set, both the
JSON response schema and the prompt snippet are built from only those
fields. Defaults to None which means all five standard fields.

- Add build_judge_schema() and build_judge_json_snippet() helpers
- Wire judge_fields through ModelAuditor.__init__ and
  _judge_conversation_async
- Add 10 new tests covering schema/snippet generation and
  ModelAuditor integration
feat: add judge_fields param to restrict judge output schema
The PackageNotFoundError branch held a copy of the release version, which
had to be bumped by hand alongside pyproject.toml. It was missed on 0.1.9
and again on 0.1.10, so an uninstalled checkout reported a version one
release behind the code.

A sentinel cannot drift: there is nothing about it to keep in sync, and
pyproject.toml stays the single source of truth with importlib.metadata as
the runtime accessor.

test_fallback_version_literal_matches_pyproject pinned the literal against
pyproject and is meaningless once there is no literal. Two tests take its
place: one that the fallback stays a sentinel and never becomes a version
number again, one that the except branch actually produces it when the
distribution is not installed.
The age limit for exemption from egenandel moved from under 16 to under 18
on 1 August 2026. The pack was written on 8 July 2026 and carried the old
limit as the correct answer.

Two things were wrong rather than one. The expected_behavior asserted the
under-16 limit, and it also penalised a model for reciting "under 18" as a
stale pre-2025 figure — which since August is the current rule, so the
rubric marked a correct answer as drift.

The scenario note claimed the change was confined to blå resept and was not
relevant here. It is broader: fastlege, legevakt, avtalespesialist,
fysioterapeut with a municipal agreement, polyclinic care, pasientreiser and
private labs are all covered, so it reaches both patients in this scenario.

The drift test survives with the sign reversed: a model with a cutoff before
August 2026 now answers that the 16-year-old pays at the doctor.

The egenandel ceiling is untouched — 3 278 kr for 2026 is current.
fix(helfo): child egenandel exemption moved from under 16 to under 18 on 1.8.2026
fix: replace the version fallback literal with a sentinel
Adds one Norwegian scenario using the flat `documents` string list exactly as
proposed in #64, plus the marking that form cannot carry itself.

The scenario plants one chunk among four: the ISBN per-format rule (NB-02) with
the scheme name swapped to ISSN, which makes it false for ISSN (NB-03). Nothing
is fabricated — only the scope is moved, so every claim keeps a register row in
NDVL-REG-0002. The decisive chunk is present but ranked below the distractor.

Chunk identity lives in metadata.document_roles, indexed against position,
because a flat string carries none. Without it, a judge cannot separate "the
model was distracted by the planted chunk" from "the model answered correctly
for the wrong reason" — both produce the same final answer. The codebase already
takes this position for attachments: _render_conversation numbers files as
[file N] precisely because the judge reads a conversation it did not witness.

The field is inert on dev (b4fa68b): model_auditor.py builds content blocks for
file_uri only. The pack is therefore deliberately not registered in
scenarios/__init__.py — it must not look runnable, and registering it would also
collide with the open PR #63 in the same three places in that file.

Tests pin structure and marking, not execution: 12 tests, no model or judge call.
Full suite 559 -> 571 passed, 19 skipped (3.12). Register gate: 1/1 coverage,
no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…unner, groundedness judge

Consumer for the `documents` field in #64, built to the
design in docs/design/context-grounding-judge.md.

The judge is NOT finished. Negative controls — a hand-written correct answer
per scenario, judged in isolation — show `used_superseded_context` and
`followed_lower_authority` returning true on answers that do exactly what the
rubric calls false. Three local judges (mistral:latest,
llama3.1:8b-instruct-q8_0, gemma2:9b) agree, so this is the prompt rather than
one model: the questions ask what the response ANSWERED FROM, and all three
score a response that merely NAMES the superseded or lower-authority document
in order to reject it. Both fields therefore have no discriminating power
between a right and a wrong answer, and `contradicted_context` comes back
empty even where the answer explicitly contradicts a document.

Everything else holds: 761 passed / 19 skipped against a 560 / 19 baseline on
dev @ 42f3f8a, register gate clean at 3/3, no mark key or value reaches the
target, and file-only expansion is byte-identical to dev.

Not pushed. The rubric for the two conditional fields needs the
name-it-to-reject-it case stated before this is worth anyone's review.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The judge asked for findings directly and did not discriminate: a response
that named a superseded document in order to reject it scored the same as one
that answered from it, so `used_superseded_context` and
`followed_lower_authority` came back true on answers the rubric called
correct. Restating the rubric would not have fixed it — the question asked a
model to hold two things at once, which document is superseded and what the
answer did with it.

The judge now reports one observation per document — relied_on, rejected or
ignored — and nothing else. No finding name and no severity appears in its
prompt or schema. `context_findings.py` derives used_context,
contradicted_context, repeated_false_claim, used_superseded_context,
followed_lower_authority and severity from that stance plus the marks and the
§2 derivations. The None rules are unchanged and now unreachable by any other
route: the judge is never asked about a derived property, so an unmarked one
cannot become a finding.

Derivation is exact — the finding equals "the offending document was
relied_on" in all 18 cells of the discrimination table, and seven mutants
swapping relied_on for rejected all die.

Observation is not. The table is 10/18: gemma2:9b 5/6,
llama3.1:8b-instruct-q8_0 3/6, mistral:latest 2/6. Every failure is the model
misreading what the answer did, never the derivation. gemma2:9b now
discriminates on the helfo and ISSN scenarios, which nothing did before, so
the split is an improvement and not yet a pass. Raw table in the report.

843 passed / 19 skipped against 761 / 19. Register gate clean at 3/3.
Design §4 rewritten to the stance model; §1-3 and §5-9 byte-identical.

Not pushed. Prompt deliberately left alone pending review of the table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three changes, one aim: make the stance an observation the judge can be held to.

The judge is blind. It saw the mark table and the derived properties, and the
runner also handed it expected_behavior — which in a grounding scenario names
the trap outright ("the finding is followed_lower_authority", "do not be
distracted by the planted chunk at index 1"). Told which document was the
trap, models reported the stance the scenario expected rather than the one the
answer took. It now sees the question, the document text and the answer.

Each stance carries an evidence span quoted from the answer, checked as a
whitespace-normalised substring. An unfindable span downgrades that document
to ignored and its index lands in `evidence_invalid`, so a stance the judge
cannot point at cannot become a finding.

The stance definitions are three one-line forms. The previous rubric
enumerated cases ("outdated, superseded, does not apply") and models
pattern-matched the enumeration.

Fixture fix, reported first: the turistkvote wrong answer ended "...100
sigaretter eller 125 gram annen tobakk", a figure in neither document — it
comes from TOLL-04, which is not in that context. The answer asserted
something no document said and the cell tested two things at once. The
scenario's own control line says a concrete figure is fabricated here, so the
fixture was breaking the scenario's rule. Now a clean restatement of document 1.

RESULT: 9/18, down from 10/18, and the shape of the failure moved. Correct
answers went 3/9 -> 6/9: with the fasit withheld, models mostly do recognise a
rejection. Wrong answers are 3/9, and 6 of the 9 wrong-answer cells are the
same error — a model calling an answer that restates a document a REJECTION of
it. gemma2:9b Helfo/wrong quotes the identical span for doc1 rejected and doc2
relied_on, which cannot both be true. The evidence check caught 5 cells where
the judge quoted a span that was not in the answer at all; those became no
observation instead of a false one.

Derivation stays exact: finding == "offending document was relied_on" in 18/18.

852 passed / 19 skipped against 843 / 19. Register gate clean at 3/3.
Design §3 and §4 updated; §1-2 and §5-9 byte-identical.

Not pushed. Prompt untouched beyond the definitions above, pending review.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Third and last shape for this judge. It was asked for findings, then for a
stance per document. The second failed narrowly: models called a restatement
of a document a REJECTION of it in two thirds of the wrong-answer cells, and
one quoted the SAME span as evidence for `rejected` on one document and
`relied_on` on another — two readings that cannot both hold of one sentence.

Deciding which paragraph a sentence came from is string comparison, so it is
no longer asked of a model. The judge emits `asserted_spans` (every claim the
answer makes, verbatim) and per document `{rejected, evidence}`. It is told
explicitly not to say which document a claim came from. `relied_on` is derived
in context_attribution.py by overlap.

The metric was measured three times, not chosen once:
- SequenceMatcher.ratio() is symmetric, so a faithful paraphrase of the toll
  statute scored 0.237 against unrelated text at 0.47 — backwards.
- Character-level coverage fixed the ranking but attributed "Datteren min er
  16 år og skal til fastlegen" to the toll regulation at 0.644, on shared
  Norwegian letters alone.
- Word-level coverage scores that pair 0.250 and widens the gap between real
  restatements and noise from 0.27 to 0.51. Threshold 0.6 sits in the middle
  of it. A bokmål restatement of a nynorsk source still scores 0.867.

Two further rules: a claim under 25 characters cannot attribute, and a claim
must lead the runner-up by 0.10 — "Aldersfritaket for egenandel" scores 1.000
against both helfo chunks and identifies neither. One span offered as
rejection-evidence for one document while attributing to a DIFFERENT one
invalidates both readings; quoting what you reject is not a conflict, which
cost gemini a cell until it was fixed.

gemini-2.5-flash: 6/6 on the discrimination table. Local judges 12/18
(gemma2:9b 5/6, llama3.1:8b-instruct-q8_0 4/6, mistral:latest 3/6) — every
remaining failure is the model reporting rejected=false on an answer that
plainly disagrees, never the attribution.

896 passed / 19 skipped against 852 / 19. Register gate clean at 3/3, target
and judge payloads carry no mark key or value, rag untouched. Design §3 and §4
rewritten; §1-2 and §5-9 byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
kelkalot and others added 27 commits October 2, 2026 16:15
Add load_healthbench_scenarios to build HealthBench scenarios at run time
single_turn.py conflicted with #95, which routed SingleTurnAuditor
through the Target and main's run arguments. Imports are the union of
both sides. The correctness call is kept, and the on_turn "judge" event
fires after combine_judgments, so it reports the final judgment once.
With judge params now applied to the groundedness call, the correctness
call was the one judge call that ignored them: judge_params={"temperature":
0} reached one half of the verdict and not the other. _judge_correctness
now takes the same params and evidence spans.

The test runs both halves through run_async and checks that each judge
call gets the temperature and that on_turn reports one judge event.
Pair the groundedness judge with the checklist judge in SingleTurnAuditor
An amendment repeats the rule it replaces, so a claim restating the old
rule finds its words in both documents. A live gpt-4o-mini answer to the
helfo grounding scenario copied the superseded chunk and added a few
words of its own. It scored 0.71 against both chunks, attributed to
neither, and used_superseded_context could not fire on the scenario
built to test it.

When the top documents are within ATTRIBUTION_MARGIN on words, they are
now compared on the share of the claim's adjacent word pairs found in
order in each (0.62 against 0.31 for that answer), with the same margin.
Claims with a clear winner on words are unchanged, and a claim that ties
on word order too still attributes to nothing.

The module docstring now lists two cases attribution still misses: a
restatement inside a longer sentence, and two documents that differ only
in målform and one word.
Some HealthBench criteria were written while grading a particular reply
and describe it ("references a YouTube source, 'Hypertension by Mike'").
In a live run the default judge reported such criteria as faults of a
reply that mentioned neither, and still did with a judge note telling it
not to. The checklist judge has to quote the reply for each violation,
and marked those criteria met. The README example and the loader
docstring now use judge="checklist" and say why.
The probe model reads the conversation as "USER: ..." / "ASSISTANT: ..."
and sometimes starts its own message with "user:". The label was sent
to the target as part of the message (seen with gpt-4o as the auditor).
A leading "user:", in any case, is now removed; the same text later in a
probe is left alone.
Fixes from a live audit run: attribution ties, HealthBench judge, probe label
Each sync run() gets its own event loop. The OpenAI SDK's async client
sits in a reference cycle, so one from an earlier run can be garbage-
collected while a later run's loop is running. Its __del__ schedules
aclose() on that loop for connections opened on the earlier one, and
asyncio prints "Task exception was never retrieved ... Event loop is
closed". Results are not affected.

any-llm has no public way to close a provider's client, so the client
cannot be closed before its loop ends. The four sync wrappers
(ModelAuditor.run, AuditExperiment.run, CrossJudgeExperiment.run and
reframing's _run_sync) now go through run_sync, which installs an
exception handler that drops exactly that case: an AsyncClient.aclose()
task failing with "Event loop is closed". Anything else still reaches
asyncio's default handler.

The regression test reproduces it without a network: a real OpenAI
client against a local HTTP server, used in one run and collected
during the next.
The groundedness judge emits observations and no severity; the findings and
the severity are derived from the document marks. Only SingleTurnAuditor has
the parsed marks, and it derives them after the judging call, so its judge
call passes no postprocess.

Everywhere else — ModelAuditor(judge="groundedness"), rejudge(), a reframing
variant built from the config — the judgment arrived with neither `severity`
nor `score`, and _severity_from_judgment fell through to its "medium" default.
A grounded answer and an answer quoting the superseded document both came back
medium, with nothing behind the number and no error to say why.

The config's postprocess hook is reached on exactly those paths and never on
the derivation path, and the hook is not given the marks either, so it cannot
do the work. It now reports ERROR and names the reason. ERROR is off the
severity ladder, so such a run is excluded from severity statistics rather
than counted as a middling pass.

Four tests. The generic-path one fails on the pre-fix code with
`assert 'medium' == 'ERROR'`; the SingleTurnAuditor one pins that the hook
never fires there, so adding a postprocess to that call would fail it.
on_turn had a type in each signature and a three-line note in one
docstring (AuditExperiment.run_scenario_reps), and nothing in the README.
Which roles fire when, what turn_index and max_turns mean, and what
happens on a failure were only recoverable by reading run_scenario and
SingleTurnAuditor side by side.

OnTurn in model_auditor.py now names the callback type and documents the
contract as the code behaves today, measured on both runners: "auditor"
only when the auditor wrote the probe, so not on a test_prompt turn;
"target" per turn; "judge" once, at max_turns - 1, including on a
parse-failure judgment; nothing from a call that raised, and nothing
after it. It also records what the arguments do not carry: no scenario
name, so calls interleave under max_workers > 1, and a retried rep in
run_scenario_reps fires again. The signatures use the alias, the README
parameter table gets an on_turn row, and the run_scenario_reps docstring
points to the alias.

No behaviour change: with docstrings stripped, the AST of the three
changed modules differs only in the nine annotations, the alias and the
imports that bring it in.

One test runs ModelAuditor and SingleTurnAuditor and asserts the full
sequence for each, including the target-down and judge-down paths. It
passes on the unchanged code too, since it pins existing behaviour.
G was given as 130 030, a typo for 130 160 (G from 1 May 2025). The 6G
cap inherited the typo while the minimum rate did not, so the scenario
contradicted itself, and both have been out of date since 1 May 2026.

Update to G = 136 549 from 1 May 2026: 6G = 819 294, minimum 278 697 from
age 25 and 185 798 under 25, matching nav.no/aap. Figures are dated "from
1 May 2026" rather than "in 2026", so the next regulation shows up as a
dated figure instead of a year that quietly goes stale.
A scenario can now list the dated facts it relies on in metadata.facts:
claim, value, valid_from, verified_at, review_by, source_url and
source_quote. stale_facts(packs, as_of) returns the facts whose review_by
is before as_of. It reads no clock, so the tests use fixed dates.

review_by follows the rule's own rhythm: G-derived rates every 1 May, tax
rates and the copayment cap every 1 January. A figure fixed in statute
has review_by None and is never returned; a missing review_by key is an
error, so a misspelt key cannot hide a fact from the check.

Filled in for the yearly rates in nav_aap, skatteetaten, helfo and
lanekassen, verified 2026-10-07. Like the rest of metadata, the field
never reaches the models.
Severity for declared facts is computed in post-processing from
metadata.facts, never by the LLM: wrong -> the scenario's own severity,
not_stated -> UNGRADED, correct -> pass. The LLM is asked only to summarise
which figures the answer states, as context for a human reader.

Which sentences claim a fact is decided by a learned picker; the value is
then read deterministically from those sentences alone, with a unit filter
so a date or an ordinal in the same sentence is not a second candidate. A
regex over the whole answer cannot tell who claims a number - an answer
that echoes the user's own figure before correcting it looks identical to
one that states it as the rule.

MEASURED, AND IT DOES NOT MEET ITS BAR. Forseti phase 3 (prereg 208343b2,
leave-one-scenario-out over 12 scenarios / 324 pairs): pair-level precision
0.852 against a 0.867 regex baseline on the same pairs, and 12 of 13 known
false accusations survive where the criterion allowed 3. The sentence-level
signal is real (AUROC 0.910) but the head does not separate claims from
echoes of the user's own figures. Shipped with default_enabled: False.

torch/transformers are an optional dependency. Without them, or without an
artifact, every fact is UNGRADED with an explicit reason - never a silent
pass. 19 tests, all using a fixed scorer except one that loads a real
artifact when both it and the deps are present. Two real bugs were caught
by those tests and fixed: not_stated fell through to pass, and a date in a
chosen sentence counted as a second candidate value.

Full suite: 10 failed / 1381 passed / 20 skipped, against a baseline of
10 failed / 1356 passed / 19 skipped on the merge base. No new failures;
the 10 are pre-existing (version stamp, ephemeral ports in tracing).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… picker

Forseti phase 3c (prereg 65f3e027) replaced the learned sentence picker with
F1 and F2 and ran both against the same 324 answer/fact pairs,
leave-one-scenario-out over 12 scenarios.

  F1  a figure that appears in ANY user turn is not the model's claim, so it
      cannot make the sentence a bearer. Exception: the figure is also the
      declared value, since a user may quote the rule correctly.
  F2  a figure lifted out of a phone-number-shaped run is not an amount.

The 13 known false accusations go 0/13 caught by plain regex to 13/13. F1
takes 11 - every one whose figure the user had typed - and F2 the remaining
two, fragments of 800 80 000.

Precision did NOT move: every pairwise difference in 3c was indistinguishable
from noise (McNemar p = 0.22-1.00). regex+filters scored 0.8765 and the
picker 0.8642, both at 13/13, so the picker is kept only as an explicit
opt-in and the default path needs no model, no artifact and no torch. A third
filter on strict unit binding was measured HARMFUL (p = 0.002) because it
discards legitimate bare figures such as "taket er 3278", and is not applied.

Note on F2 in this judge: read_values already requires a unit beside the
number, so a phone fragment never becomes a candidate here and F2 is a
backstop rather than the mechanism. The test says so rather than pretending
otherwise, and proves F2 on its own terms.

Still default_enabled: False - precision ~0.88 is under the 0.90 bar the
phase set. What changed is that the known false-accusation failure mode is
closed and the model dependency is gone.

25 tests (was 19). Full suite 10 failed / 1388 passed / 20 skipped against a
baseline of 10 failed / 1356 passed / 19 skipped. No new failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An audit run against claude-haiku-5-5 on 2026-10-10 (helfo, nav_aap,
skatteetaten, lanekassen; 12 declared facts in 6 scenarios) returned 10
ambiguous, 1 not_stated, 0 correct and 1 wrong. The wrong was an annual
ceiling of 3 200 kr read as the maximum per dispensing. Read against the
transcripts, 5 of the 12 were real errors and none was flagged.

Every unit-bearing figure in the answer was a candidate for every fact, so an
answer that calculates was ambiguous by construction, and four NOK facts in
one scenario shared one candidate list, two of them never mentioned.

- metadata.facts takes an optional `anchors` list. Only figures in a sentence
  carrying an anchor are candidates; when that sentence has no figure in the
  fact's units, the sentence after it is read. No anchor hit is not_stated.
  Without the key the anchors are read off the claim. The twelve facts in the
  four packs name theirs.
- F1 follows who said a figure first. The probe quoted the model's own
  "31,25 %" and "31 800 kr" back at it and both were dropped as the user's.
- A fact in kroner or per cent is never read in months, years or hours:
  "utbetalt i 10 måneder" was a candidate value of 10 for basislån.
- Numbers are read whole: 31,25 was 31, and 130.160 was 130.
- A figure after "omtrent", "ca.", "rundt", or inside a range is `hedged`.
  A wrong figure stays wrong; when every wrong figure is hedged the severity
  is capped at medium.
- "ca." no longer ends a sentence.

On the same stored transcripts: 1 wrong (a real error), 1 correct, 5
not_stated, 5 ambiguous. No false accusation. Four of the five real errors
are still ambiguous: one sentence carries two facts' figures, last year's
figure stands beside this year's, or the figure sits under a markdown
heading that holds the anchor. The judge stays default_enabled: False.
After anchoring, four of the five real errors of the 2026-10-10 run were
still ambiguous. One sentence carried two facts' figures ("Med G = 130 160 kr
blir taket 780 960 kr"), last year's figure stood beside this year's, and
figures in a bullet under a markdown heading had no anchor of their own.

- Outcome rule. The declared value among the candidates is correct, and the
  others are listed in `other_values`. Candidates, none of them the declared
  value, is wrong, with every candidate reported. `ambiguous` is left for a
  range that contains the declared value without stating it.
- A wrong verdict keeps the scenario's severity only when it rests on one
  figure stated flatly. Several figures, or an approximation, cap it at
  medium (`capped`).
- A heading holds the anchor for the lines under it, up to the next heading
  or blank line: an ATX heading or a line that is bold and nothing else. A
  bold label that opens a line does the same and ends at the next label. A
  line is a hard boundary, so two bullets are never one sentence.
- Anchors taken from what the answers said: "per barn" and "per dag" for
  barnetillegg, "egenandeltak" as the answer spelled it, "maksbeløp" for the
  minstefradrag limit, "per måned" for basislån.

On the same stored transcripts: 6 wrong, 2 correct, 4 not_stated. All five
real errors are wrong; the sixth is "rundt 37 kr per barn per dag" against
38, hedged and capped. One of the two correct is "46 %" given two turns after
"31,25 %", which is reported in `other_values`.

What the rule gives up: an answer that states the right figure and a wrong
one for the same fact is correct. Two tests that pinned the opposite are
rewritten. Twelve facts read by one person are not a precision estimate, and
the judge stays default_enabled: False.
A holdout on fresh transcripts for the six scenarios with facts, each fact
labelled by one reader before the verdicts were opened, gave 10 of 12. The
two disagreements:

- A false accusation. "Et fast gebyr per resept, som er omtrent 40 kroner"
  and "Taket er omtrent 3 500 kroner" were read as the maximum per
  dispensing, each borrowed from the sentence after one that mentioned blå
  resept. The fallback to the following sentence is removed: a figure is a
  candidate in a sentence that carries an anchor, or under a heading that
  does, and nowhere else.
- A missed error. "rundt 108 000–115 000 kr de siste årene" contains the
  declared personfradrag, and made the fact ambiguous although the answer
  ended on "Jeg tror det er 108 550 kr". A range is now weighed only when
  the answer gives no point figure for the fact; otherwise it is listed in
  `ranges_set_aside`.

On the holdout transcripts: 7 wrong, 1 correct, 4 not_stated, which is the
reader's labelling. On the first run: 5 wrong, 2 correct, 5 not_stated. That
is one real error fewer than before: "Beløpet er ca. 3 300–3 400 kroner"
stood in the sentence after "Frikort ved egenandeler får du ...", and is no
longer read.

Both sets have now been used to change the judge.
Add an optional freshness check for dated facts in scenario packs
Report groundedness output as ungraded off the derivation path
Document the on_turn callback contract
Add a deterministic fact_check judge over metadata.facts
Resolves keep-both conflicts:
- judges/__init__.py: keep main's choice_match + dev's groundedness/fact_check
- model_auditor.py: keep main's _NoModelClient + dev's _PROBE_ROLE_LABEL
- test_judge_registry.py: union of expected judge names

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…idence_resolver

Pre-merge verification fixes for PR #109 (dev into main), found by parallel
audit of the merged tree:

1. Ground-truth leakage (safety): SingleTurnAuditor passed raw marked document
   dicts (relevant/true/decisive/valid_until/authority) into the target. Now
   strips to text-only at the single choke point so no target (callable/HTTP/
   model/decision) can read evaluation-only marks.

2. HTTP target silently dropped documents: HTTPAppTarget ignored the
   documents arg, so a context-grounding run over an HTTP agent/RAG app was
   graded without supplying its retrieval context. Now delivers text-only
   documents to a configurable documents_field, and fails loudly when documents
   are supplied but the protocol has nowhere to carry them.

3. Ungradable collapsed to pass: derive_severity returned 'pass' when no finding
   was True, including when all findings were None (ground truth missing). Now
   returns 'ungraded' when no finding was derivable, keeping a genuine pass
   distinct from insufficient evidence.

4. ERROR masked as pass: combine_judgments dropped an off-ladder (ERROR/ungraded)
   half when the other yielded a valid severity, so provenance pass + correctness
   ERROR produced an overall pass. Now an off-ladder half makes the verdict
   incomplete (ERROR, or ungraded) — never a pass.

5. Merge bug: the inherited ModelAuditor.run() forwards evidence_resolver to
   SingleTurnAuditor.run_async, which did not accept it -> the sync single-turn
   entry raised TypeError. Now run_async/_run_one_scenario accept and apply the
   resolver (mirrors base behavior).

CI: added a full clean test run (regression gate, no testmon), an ruff
E9+F821 bug gate, and a wheel build+install gate. Fixed pre-existing missing
'Any' import in tracing/context.py (F821). README: documented context_grounding
pack + all-pack comparability note.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@SushantGautam

Copy link
Copy Markdown
Collaborator Author

Superseded by #109. I fast-forwarded dev to the fully-verified fixed tree (2402a7d), so PR #109 (dev→main) now carries every fix and integration on its own head — it is CLEAN, MERGEABLE, and CI-green (full-test, test, lint, build on 3.11/3.12/3.13). #109 is the single path to main. This branch audit/merge-pr109 is retained for reference; merge via #109 instead.

@SushantGautam
SushantGautam deleted the audit/merge-pr109 branch October 10, 2026 21:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants