Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 24 additions & 11 deletions docs/adr/0001-episode-semantics-boundaries.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,11 +46,12 @@ explicit for v0.7 research:
distinct-username objective for multi-user probing), not the longest possible
time span. Tie-breaking and whether a candidate may be extended without
increasing its score are v0.7 research questions.
3. **Non-overlapping episode** is currently enforced by segment construction
and one selection per segment. It does not yet describe a global candidate
ranking problem where two windows compete for shared events. v0.7 must state
whether exclusion is by event IDs, time intervals, or a rule-specific signal
budget.
3. **Non-overlapping episode** has two candidate-v1 boundaries. Search windows
may share event IDs and time intervals. Selection compatibility separately
requires a later window to start more than one rule window after the prior
window ends. At publication, selected episodes inside a segment must have
disjoint event-ID sets; the oracle validator fails closed on reuse. This does
not define cross-rule evidence reuse or a future production signal budget.
4. **Cooldown merge** currently means that adjacent signals with a gap less
than or equal to the rule window remain in one candidate segment. A gap
larger than the rule window starts another segment. This is an adjacency
Expand All @@ -76,7 +77,7 @@ that makes each boundary observable. At minimum it should cover:
| Continuous bridge | Two dense bursts connected by background gaps at or below the rule window remain one baseline segment. |
| Threshold edge | A candidate exactly at threshold is eligible; one event below threshold is not. |
| Maximal-window tie | Equal-score windows have a documented deterministic tie-break. |
| Shared evidence | Overlapping candidate windows make the non-overlap unit explicit. |
| Shared evidence | Overlapping search candidates expose shared event IDs; selected episodes materialize each event ID at most once, and reuse fails closed. |
| Cooldown boundary | A gap exactly at the rule window and one second beyond it test the inclusive boundary. |
| Bimodal background | Two dense peaks are surrounded by lower-rate background, with expected baseline output recorded separately from the research candidate output. |

Expand Down Expand Up @@ -160,11 +161,23 @@ The earlier candidate wins for both forward and reversed candidate input. This
accepts the final tie-break and candidate-order invariance for that bounded
control; it does not settle the separate shared-evidence policy.

Together, the two fixtures and focused tie control accept the recovery,
null-control, and deterministic tie-break hypotheses for their bounded cases
only. Shared evidence, background calibration, and a production complexity
design still require independent evidence. `Detector::analyze()` and
`loglens.report.v3` remain unchanged.
A focused shared-evidence control uses six events at offsets 0 through 5
seconds. With threshold five, enumeration produces `line:1` through `line:5`,
`line:1` through `line:6`, and `line:2` through `line:6`; `line:2` through
`line:5` belong to all three search candidates. Candidate v1 selects the
six-event maximal window and materializes each event ID once. A separate
negative control gives two distinct selected candidates the same event-ID set;
oracle validation rejects it as `selected episodes reuse event evidence`.
Mutation testing confirms that disabling this guard makes the focused test
fail. This accepts event IDs as the candidate-v1 publication non-overlap unit,
not as a cross-rule or production policy.

Together, the two fixtures and focused tie/shared-evidence controls accept the
recovery, null-control, deterministic tie-break, and publication
single-consumption hypotheses for their bounded cases only. Background and
alert-volume calibration plus a production complexity design still require
independent evidence. `Detector::analyze()` and `loglens.report.v3` remain
unchanged.

## Alternatives considered

Expand Down
21 changes: 18 additions & 3 deletions tests/test_episode_candidate_core.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,14 +60,29 @@ def test_candidate_cooldown_requires_more_than_one_rule_window(self) -> None:
],
)

def test_overlapping_candidates_materialize_one_maximal_window(self) -> None:
def test_shared_evidence_is_search_membership_not_double_consumption(self) -> None:
candidates = enumerate_candidate_windows(
make_events([0, 1, 2, 3, 4, 5]), 5, 600
)
selected = select_window_separated_candidates(candidates, 600)
materialized_event_ids = [
event_id for candidate in selected for event_id in candidate.event_ids
]

self.assertEqual(len(selected), 1)
self.assertEqual(selected[0].event_ids, tuple(f"line:{i}" for i in range(1, 7)))
self.assertEqual(
[candidate.event_ids for candidate in candidates],
[
tuple(f"line:{i}" for i in range(1, 6)),
tuple(f"line:{i}" for i in range(1, 7)),
tuple(f"line:{i}" for i in range(2, 7)),
],
)
self.assertEqual(
set.intersection(*(set(candidate.event_ids) for candidate in candidates)),
{f"line:{i}" for i in range(2, 6)},
)
self.assertEqual(materialized_event_ids, [f"line:{i}" for i in range(1, 7)])
self.assertEqual(len(materialized_event_ids), len(set(materialized_event_ids)))

def test_equal_score_windows_prefer_chronological_key(self) -> None:
candidates = enumerate_candidate_windows(
Expand Down
20 changes: 20 additions & 0 deletions tests/test_episode_candidate_validation.py
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,26 @@ def test_validator_rejects_selected_candidate_without_episode(self) -> None:
with self.assertRaisesRegex(ValueError, "selected candidates"):
validate_oracle(self.fixture, self.baseline, oracle)

def test_validator_rejects_selected_episodes_that_reuse_event_evidence(
self,
) -> None:
oracle = evaluate_fixture(self.fixture, self.baseline)
segment = oracle["segments"][0]
first_episode, second_episode = segment["selected_episodes"]
candidates = {
candidate["candidate_id"]: candidate
for candidate in segment["candidate_windows"]
}
first_candidate = candidates[first_episode["candidate_id"]]
second_candidate = candidates[second_episode["candidate_id"]]
for field in ("event_ids", "first_seen", "last_seen"):
second_candidate[field] = copy.deepcopy(first_candidate[field])
second_episode[field] = copy.deepcopy(first_episode[field])
second_episode["finding_id"] = first_episode["finding_id"]

with self.assertRaisesRegex(ValueError, "reuse event evidence"):
validate_oracle(self.fixture, self.baseline, oracle)

def test_validator_rejects_event_candidate_reference_drift(self) -> None:
oracle = evaluate_fixture(self.fixture, self.baseline)
oracle["segments"][0]["event_decisions"][0]["candidate_ids"] = []
Expand Down
Loading