Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
100 changes: 100 additions & 0 deletions docs/email-security/automation.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ is one sensor, not ten thousand.
|---|---|
| `EMAIL_MESSAGE` | Once per message, at ingest. Carries the whole parsed model — headers, sender, recipients, body, links, attachments, authentication, hops — plus the enrichments and the verdict. It is the record that this mail arrived |
| `EMAIL_VERDICT` | On **every** verdict decision: the rule pack's own, at ingest right after the `EMAIL_MESSAGE` (`revision/seq: 0`, `revision/mode: auto`), and then once per override afterwards (`seq: 1…`, mode `analyst`, `ai` or `detonation`) |
| `EMAIL_ANALYSIS_COMPLETE` | When the initial analysis window closes, including messages with nothing pending. Carries terminal outcomes, final verdict snapshot and timing |
| `EMAIL_ACTION` | On every remediation outcome, including failures and skips, **and on every raw-message download** (`action: get_eml`), served or refused. Who asked, what was attempted, what happened |
| `EMAIL_USER_REPORT` | When a message reaches the abuse mailbox and becomes a report |
| `EMAIL_INGEST_ERROR` | When a message could not be fetched or processed. Coverage honesty: failures are visible, never silent |
Expand Down Expand Up @@ -57,6 +58,21 @@ The MDM is deliberately *not* repeated here: it is already in the immutable
`EMAIL_MESSAGE`, and copying it into every verdict change would multiply a year
of telemetry by how often people change their minds.

### `EMAIL_ANALYSIS_COMPLETE`

This event uses the same identity and `revision` paths as `EMAIL_VERDICT`, so
triage reads `event/revision/verdict`, not the MDM's `event/verdict/verdict`.
`event/results` maps each armed analysis (`detonation`, `attachment_scan`) to
`completed`, `changed_verdict`, `skipped`, `shed`, `failed`, or `timed_out`.
`event/completed_at` and `event/timing` describe the completed window. Empty
results mean no delayed work was needed. The completion payload does not include
the seq-0 `analysis` snapshot or `after_complete`.

Start an AI or analyst triage workflow on this event when it needs the initial
analysis results. Continue handling later `EMAIL_VERDICT` escalations: analyses
that finish beyond the deadline carry `event/after_complete: true`. Completion
never means that a failed or timed-out analysis cleared the message.

### A message that joins a campaign late

Clustering runs while a message is being ingested, so two copies of one attack
Expand Down Expand Up @@ -285,3 +301,87 @@ Two conventions make this pleasant to keep in git:

Onboarding a new tenant is then: subscribe the extension, write the secret, write
the provider record, apply the policy directory, run the connection test.

## Triage after initial analysis

This platform D&R detection reports suspicious or malicious messages after the
initial evidence window closes. The final verdict is a snapshot at completion;
read `results` when your triage needs to distinguish an examined message from a
deadline, capacity refusal or analysis failure. Completion delivery is at least
once. Use `completion_id` as the workflow's idempotency key when dispatching
external work. Initial historical backfill emits no completion event.

```yaml
# Detect
event: EMAIL_ANALYSIS_COMPLETE
op: and
rules:
- op: exists
path: event/revision/verdict
truthy: true
- op: or
rules:
- op: is
path: event/revision/verdict
value: malicious
case sensitive: false
- op: is
path: event/revision/verdict
value: suspicious
case sensitive: false
```

```yaml
# Respond
- action: report
name: email-analysis-triage
suppression:
max_count: 1
period: 720h
is_global: true
keys:
- email-analysis-triage
- '{{ .event.completion_id }}'
```

The example suppresses repeated reports for the same completion for 30 days
across sensors within the organization. After the suppression period expires,
the same identity can report again. External workflows that require durable
idempotency should retain their own completion identities for their retry horizon.

The initial `EMAIL_MESSAGE` remains useful for immediate containment and
content rules. Waiting for completion is a workflow choice; it does not prevent
the existing ingest-time automations from containing an already malicious message.

## Alerting on provider delivery delays

This example reports a provider notification lag over five minutes. The
existence guard excludes messages whose notification time is unknown. Adjust
the threshold to your own operating expectations; this is an example rule,
not a built-in alert.

```yaml
# Detect
op: and
rules:
- op: is
path: routing/event_type
value: EMAIL_ANALYSIS_COMPLETE
- op: exists
path: event/timing/provider_lag_ms
- op: is greater than
path: event/timing/provider_lag_ms
value: 300000
```

```yaml
# Respond
- action: report
name: email-analysis-provider-lag
```

Compare `provider_lag_ms` with `queue_ms` and `processing_ms` to separate delay
before notification from delay inside processing. The same pattern can alarm
on another present timing field. `analysis_ms` includes the delayed analysis
window; `end_to_end_ms` ends at the initial verdict. Check `clock_skew` before
interpreting clamped measurements.
89 changes: 84 additions & 5 deletions docs/email-security/pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,10 +7,10 @@ message from the moment your provider says it exists to the moment somebody
decides what to do about it, and it is explicit about which parts happen in one
pass and which parts can happen later.

The short version: **everything that produces the first verdict happens
synchronously, in one pass, per message.** There is no queue of half-judged mail
and no second job that fills in the answer. Anything that changes a verdict
afterwards is recorded as a *revision*, not as a late arrival.
The first verdict uses the evidence available within the bounded initial pass.
Delayed attachment scans and link detonation can add evidence afterwards; a
changed verdict is recorded as a *revision*. Wait for `EMAIL_ANALYSIS_COMPLETE`
when your workflow needs the initial analysis window to close.

## The stages

Expand Down Expand Up @@ -112,7 +112,7 @@ the message at the provider.
| | |
|---|---|
| **Synchronous, one pass per message** | Fetch, parse, every enrichment including attachment explosion, matching, scoring, the verdict, campaign clustering, persistence, and the emission of `EMAIL_MESSAGE` followed by the engine's own `EMAIL_VERDICT` (`revision/seq: 0`) |
| **Later, and recorded as such** | Verdict revisions, remediation outcomes, campaign membership added when a later message joins the cluster |
| **Later, and recorded as such** | Link detonation, deferred attachment scans, analysis completion, verdict revisions, remediation outcomes, campaign membership added when a later message joins the cluster |

A revision does **not** rewrite `EMAIL_MESSAGE`. The original event stands as the
record of what the engine decided at ingest, a revision row is appended with its
Expand All @@ -135,6 +135,85 @@ revision: a message nobody has overridden still reports zero revisions.
the gaps named. A slow enrichment degrades one signal. A blocking one
degrades coverage, which is the thing you bought.

## Knowing when initial analysis has finished

The initial `EMAIL_VERDICT` (`revision/seq: 0`) includes
`analysis: {pending: [...], complete: false}` when delayed work is outstanding.
The closed set of pending kinds is `detonation` and `attachment_scan`. With
nothing outstanding, it carries `pending: []` and `complete: true`.

`attachment_scan` is also pending when your organization's custom YARA rules were
not yet loaded while the message was processed, for example just after a deploy or
restart. The attachments are rescanned with your rules once they load, whatever the
verdict or direction, and a match revises the verdict. If the rules cannot be loaded
before the deadline, the result is `timed_out`, never a clean `completed`.

`EMAIL_ANALYSIS_COMPLETE` closes that initial window for every message emitted
through the live lane, including re-drives that emit and messages with no delayed
work. Initial historical backfill emits no `EMAIL_*` events and has no completion
event; a later live notification can promote that message into the live lane. It carries the final
`revision/verdict`, `revision/score`, `revision/seq`, and known message identity
fields, plus `results` and `timing`. Use this event to start triage that needs the
initial batch of evidence. A completion is a processing fact, **not a safety
verdict**. Its `completion_id` is the message's `msg_uuid` and stays the same
for that immutable snapshot. Delivery is at least once: retries after an interruption can repeat
the snapshot. Use `completion_id` for idempotent triage or the
[completion suppression example](automation.md#triage-after-initial-analysis).

| Result | Meaning |
|---|---|
| `completed` | The analysis finished without changing the verdict |
| `changed_verdict` | The analysis committed a verdict revision |
| `skipped` | The analysis was no longer applicable |
| `shed` | Capacity admission refused the work |
| `failed` | The analysis failed |
| `timed_out` | The initial analysis deadline expired before a terminal result |

Only armed kinds appear in `results`; `{}` means no delayed work was needed.
Pending work is durable. A collector restart or lost in-memory task cannot leave
the window open forever: unresolved work becomes `timed_out` at the configured
deadline, which defaults to 20 minutes. Recovery publishes queued completion
snapshots after a service interruption.

Evidence that arrives after this boundary can still change the verdict. Its
`EMAIL_VERDICT` carries `after_complete: true`; it does not rewrite the completion
snapshot. Keep a revision handler alongside completion-based triage for those
later changes. A later escalation re-runs post-verdict rules and automations;
a downgrade does not automatically restore mail. An explicit operator re-judge
that escalates a message also runs responses against its newly judged evidence,
including a re-judge of historical mail. Re-judge responses run asynchronously
after the stored correction: a changed count does not mean containment has
finished. Dry runs do not act; initial historical ingestion still does not run
automations.

## Processing latency

The MDM's `timestamps` includes optional `notified`, the time the provider's
notification reached LimaCharlie. Historical or periodically discovered mail
may have no notification time. The seq-0 verdict and completion event include
absolute `sent`, `received`, `notified`, `ingested`, `decided`, and `completed`
instants where applicable, with these integer millisecond measurements:

| Timing field | Interval |
|---|---|
| `provider_lag_ms` | Provider received → notification reached LimaCharlie |
| `queue_ms` | Notification reached LimaCharlie → processing began |
| `processing_ms` | Processing began → initial verdict decided |
| `end_to_end_ms` | Provider received → initial verdict decided |
| `analysis_ms` | Initial verdict decided → initial analysis window completed |

`provider_lag_ms` and `queue_ms` are **absent** when `notified` is unknown. A
measured zero is present as `0`. Negative intervals clamp to zero and set
`clock_skew: true`; investigate clock differences before treating those zeros as
fast processing. `sent` comes from the sender's untrusted Date header and is
never used for these calculations. In coalesced Gmail notifications, `notified`
is the batch's receiver observation, not a claim of one notification per message.

The message drawer shows this timeline and each analysis outcome. The API and
`limacharlie mailsec message get <UUID>` expose the same `analysis`
and `timing` state. See [Events & Automation](automation.md#triage-after-initial-analysis)
for completion triage and provider-delay alerts.

## Rules are organization-owned

The organization’s enabled `dr-mail` records are the complete rule set. Defaults
Expand Down
37 changes: 36 additions & 1 deletion docs/email-security/rule-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ have different wrappers and validation rules.
| `dr-mail`, either phase | `sender/email/domain/root` |
| `dr-mail`, `post_verdict` only | `verdict/verdict` |
| `dr-general` on `EMAIL_MESSAGE` | `event/sender/email/domain/root` |
| `dr-general` on `EMAIL_VERDICT` | `event/revision/verdict` |
| `dr-general` on `EMAIL_VERDICT` or `EMAIL_ANALYSIS_COMPLETE` | `event/revision/verdict` |
| Inside `scope` with `path: links` | `href_url/domain/root` |

The MDM is the root of a mail rule. Do not add `mdm/` or `event/`. The Hive
Expand Down Expand Up @@ -214,6 +214,7 @@ D&R rule on `EMAIL_MESSAGE`. Message-list filtering does not accept `mail_type`.
|---|---|---|
| `sent` | timestamp | When set |
| `received` | timestamp | Always |
| `notified` | timestamp | When the provider notification reached LimaCharlie; absent for mail without a known notification |
| `ingested` | timestamp | Always |

### Headers
Expand Down Expand Up @@ -689,3 +690,37 @@ lookup failure must not be treated as evidence against a message.

When present, `code` is the stable identifier for a recovered failure. Prefer it
over matching the human-readable `message`, which may change.

## Analysis status and timing on platform events

These fields belong to `EMAIL_VERDICT` and `EMAIL_ANALYSIS_COMPLETE`, rather
than the MDM rule root. A `dr-mail` scoring rule cannot wait for completion;
use a platform `dr-general` rule on the emitted completion event.

| Path | Available on | Meaning |
|---|---|---|
| `event/completion_id` | `EMAIL_ANALYSIS_COMPLETE` | Stable completion identity, equal to `msg_uuid`; use it to suppress retried delivery |
| `event/analysis/pending` | Seq-0 `EMAIL_VERDICT` | Array of `detonation` and/or `attachment_scan`; empty when none outstanding |
| `event/analysis/complete` | Seq-0 `EMAIL_VERDICT` | Whether there was no outstanding work in that snapshot |
| `event/results/<kind>` | `EMAIL_ANALYSIS_COMPLETE` | `completed`, `changed_verdict`, `skipped`, `shed`, `failed`, or `timed_out` |
| `event/completed_at` | `EMAIL_ANALYSIS_COMPLETE` | When the initial analysis window was durably closed |
| `event/revision/verdict`, `event/revision/score`, `event/revision/seq` | Both | Initial decision or final completion snapshot |
| `event/after_complete` | Later `EMAIL_VERDICT` | True when a revision was decided after the completion boundary |
| `event/timing/received`, `ingested`, `decided` | Both | Required absolute processing instants |
| `event/timing/sent`, `notified` | Both, when known | Sender Date header and notification arrival; sent is untrusted |
| `event/timing/completed` | Completion | Absolute completion instant |
| `event/timing/provider_lag_ms` | Both, when notified known | Received → notified, integer ms |
| `event/timing/queue_ms` | Both, when notified known | Notified → ingested, integer ms |
| `event/timing/processing_ms` | Both | Ingested → initial decided, integer ms |
| `event/timing/end_to_end_ms` | Both | Received → initial decided, integer ms |
| `event/timing/analysis_ms` | Completion | Initial decided → completed, integer ms |
| `event/timing/clock_skew` | When true | At least one negative interval was clamped to zero |

Completion delivery is at least once; key response suppression on
`event/completion_id`, as shown in the [triage example](automation.md#triage-after-initial-analysis).
Completion covers emitted live messages, including emitting re-drives. Initial
historical backfill emits no `EMAIL_*` events and no completion event.

Missing optional intervals are absent, never fabricated zero. A completion with
`failed`, `shed`, or `timed_out` results does not classify the message as benign.
See [completion triage and delay rules](automation.md#triage-after-initial-analysis).
Loading