Skip to content
70 changes: 70 additions & 0 deletions docs/email-security/detections.md
Original file line number Diff line number Diff line change
Expand Up @@ -285,13 +285,83 @@ attachment threats, suspicious content, detonation evidence and graymail.
They are ordinary D&R rules over the Message Data Model, not a separate engine.
See [Mail Rules](custom-rules.md) for the format, IaC and explicit restoration.

### Callback phishing and HTML smuggling

Two attack classes are invisible to a scanner that only looks at links and known-bad
files, so the defaults read them from structure.

**Callback phishing** (telephone-oriented attack delivery) is an invoice, renewal or
"your device is infected" notice whose only action is a phone number. The defaults
read the number as a fact ([PhoneNumbers](rule-reference.md#phonenumbers)) from the
body, from the text of attached images, from numbers in a PDF, and from attached
messages, and combine it with a call to action, billing vocabulary, how short the
message is, and whether the sender looks odd (a free-mail address, a young domain, a
failing DMARC result, a Reply-To elsewhere). A legitimate vendor's receipt carries the
same words and a support number, which is why a sender oddity is required and an
established sender is never read as a lure. The callback rules describe one
observation and do not add up: a message that trips all of them scores as the
strongest. A PDF's wording is not available to rules, so the PDF rule judges a short
PDF by its shape and its numbers; numbers in a PDF are recognised for North American
formats only.

**HTML smuggling** is a web page, often an `.html` or `.svg` attachment, that builds
the real payload in the victim's browser. The parser scans every HTML-like attachment
and HTML body in full and reports encoded data, the type that data decodes to,
decoding and download primitives, redirects and password forms
([HTMLIndicators](rule-reference.md#htmlindicators)). Defaults flag a page that
decodes encoded data and saves it, a page whose encoded data is an archive or
executable, a page that builds its own decoder, an HTML sign-in page delivered as a
file, a tiny redirect page, an SVG that carries script, and a message body that runs
a decoder. A single-file report or export tool that embeds data and offers a download
button matches the same facts as a smuggling page and is scored as suspicious, not
malicious, unless its data also decodes to a recognisable payload. Credential-page
and tiny-redirect defaults require a browser-file extension; source templates and
files with unconventional names can fall outside those two checks.

In managed pack `0.6.0`, these 22 new rules carry explicit severity. Older managed
rules currently use the informational fallback. Severity is independent of the
verdict, so a malicious verdict from an older rule can still have informational
severity.

Display-name brand impersonation ("PayPal Support" over an unrelated address),
advance-fee and extortion text, voicemail and fax lures, free-hosting and
open-redirector links, internationalised look-alike domains, OneNote files, locked
PDFs with the password in the message, and web pages hidden inside archives from a
stranger are covered by further defaults. **Email Security → Rules** shows every rule's
conditions and false-positive notes.

A verdict's `engine_version` is a SHA-256 fingerprint of the scoring rules,
resolved thresholds, exclusions, VIPs, threat-feed references and clustering policy,
and linked parsing/enrichment library build.
Changing rule content or scoring policy changes the fingerprint. It identifies
the decision configuration; it is not a promise that an external lookup feed
or other message enrichment is unchanged.

### Outbound PII detections

The optional email DLP pack adds outbound detections for validated payment card
numbers, IBANs and US Social Security numbers, plus a bulk detection when any
one kind has at least ten distinct values. Bulk counts are per kind: four cards,
four IBANs and four SSNs do not meet the bulk threshold. Bulk detections fire
alongside the matching single-kind detection. These are platform D&R rules on
`EMAIL_MESSAGE`, separate from the engine verdict; installing them does not
change the verdict or automatically move mail.

The rules read [PIIFindings](rule-reference.md#piifindings), which stores counts
only. A detection still carries the originating email event, whose body can
contain the actual sensitive values; the count facts do not redact that body.
Plan detection access and outputs accordingly. IBANs commonly occur on ordinary
invoices, and dashed SSN-shaped internal IDs can match. Tune the optional rules
for your organization rather than treating a match as proof of malicious intent.
The detector covers message text, OCR and attached messages, but does not inspect
text inside ordinary document, spreadsheet or PDF files. Counts remain lower
bounds for these excluded sources even when `enrichments/pii/truncated` is absent.
Known incomplete extraction or parsing, and inspection limits, set that flag;
the counts still describe only the available text.
Deferred attachment scans refresh the stored facts but do not replay the initial
`EMAIL_MESSAGE` evaluation. A DLP match therefore describes evidence available
when that event was emitted, rather than every later attachment result.

## Link detonation

Static link features answer what a URL *looks* like. Detonation answers where it
Expand Down
208 changes: 208 additions & 0 deletions docs/email-security/rule-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,12 @@ Scoring classes require `weight` from 1 to 100; graymail records must omit it.
| What did attachment analysis actually inspect? | Scope `attachments`; inspect `explode/scanners` before interpreting scanner-specific results |
| Is this a known sender? | `enrichments/sender_profile/prevalence` (`none`, `new`, `rare`, `common`) |
| Is the sender impersonating an organization? | `enrichments/lookalike/org_domain_distance`, `enrichments/lookalike/vip_hit` |
| Is the display name a well-known brand over an address that is not the brand's? | `enrichments/lookalike/display_name_brand` |
| Does it contain validated payment cards, IBANs or US Social Security numbers? | `enrichments/pii/card_numbers`, `ibans`, `us_ssns`; counts only, with `truncated` for incomplete inspection. See [PIIFindings](#piifindings) |
| Does the message tell the reader to call a number (callback phishing)? | `enrichments/phone_numbers/body` and `enrichments/phone_numbers/attachments`; read `call_to_action`, `lure_terms`, `toll_free`. See [PhoneNumbers](#phonenumbers) |
| Is this attachment a web page that builds a file in the browser (HTML smuggling)? | Scope `attachments`; read `html/base64_bytes`, `html/payload_types`, `html/blob_download`, `html/atob`. See [HTMLIndicators](#htmlindicators) |
| Is a PDF short, locked, or carrying phone numbers? | Scope `attachments`; read `explode/pdf`. See [PDFInfo](#pdfinfo) |
| How much of the message does the reader actually see? | `body/current_thread/visible_chars` |
| What happened after a link was fetched? | `enrichments/detonation`; see [Link detonation](detections.md#link-detonation) |
| Was parsing or analysis incomplete? | `_meta/truncations`, `_meta/errors`, `_meta/explode_timeout`, `body/truncated` |

Expand Down Expand Up @@ -315,6 +321,7 @@ the pipeline includes the whole body in `current_thread` for an unverified reply
| `inner_text` | string | Non-empty |
| `display_text` | string | Non-empty |
| `charset` | string | Non-empty |
| `indicators` | [HTMLIndicators](#htmlindicators) | When the scan found something |

### PlainBody

Expand All @@ -329,8 +336,15 @@ the pipeline includes the whole body in `current_thread` for an unverified reply
| `text` | string | Non-empty |
| `renderings` | array of [ThreadRendering](#threadrendering) | Non-empty |
| `visible_text` | string | Non-empty |
| `visible_chars` | integer | Non-empty |
| `links` | array of [Link object](#link) | Non-empty |

`visible_chars` is the number of non-whitespace characters in `visible_text`. It
measures how much the reader is shown, and blank-line padding cannot inflate it.
A mail rule cannot compute a length itself because its regular expressions are
limited to short repeat counts, so use this field for "the message is a two-line
note" conditions.

### ThreadRendering

| Field | Type | Presence |
Expand Down Expand Up @@ -399,8 +413,67 @@ and a decoded destination; when both links are emitted, it is set on both.
| `tlsh` | string | Non-empty |
| `magic_type` | string | Non-empty |
| `is_inline` | boolean | Non-empty |
| `html` | [HTMLIndicators](#htmlindicators) | When the part is HTML-like |
| `explode` | [Explode](#explode) | When set |

`html` is computed from the attachment's own bytes when the message is parsed, so
it does not depend on attachment analysis. It is present exactly when the part
was recognised as HTML-like (an HTML or SVG file, or any part whose first bytes
open like a web page, whatever it is called). An absent block means "not a web
page", never "a clean web page".

### HTMLIndicators

Structural facts from a bounded scan of an HTML-like document: an attachment, or
the HTML body. HTML smuggling is a structure, not a string: a large blob of
encoded data, a few lines of script that decode it, and a browser call that
saves the result as a download. No single field below is a finding, because HTML
exports embed images as base64 and ordinary pages call `atob`. Combine them.
Keyword fields are matched on a normalised view of the document (lower case,
whitespace and quotes removed), so spacing and case do not matter, but splitting a
word across string pieces (`'at'+'ob'`) defeats them. The encoded data and its
decoded type are reported for that reason.

| Field | Type | Presence |
|---|---|---|
| `scripts` | integer | Non-empty |
| `event_handlers` | boolean | Non-empty |
| `base64_bytes` | integer | Non-empty |
| `base64_max_run` | integer | Non-empty |
| `numeric_array_bytes` | integer | Non-empty |
| `payload_types` | array of string | Non-empty |
| `atob` | boolean | Non-empty |
| `eval` | boolean | Non-empty |
| `from_char_code` | boolean | Non-empty |
| `unescape` | boolean | Non-empty |
| `document_write` | boolean | Non-empty |
| `blob_download` | boolean | Non-empty |
| `download_attr` | boolean | Non-empty |
| `auto_click` | boolean | Non-empty |
| `js_redirect` | boolean | Non-empty |
| `meta_refresh` | boolean | Non-empty |
| `password_input` | boolean | Non-empty |
| `remote_form_action` | boolean | Non-empty |
| `truncated` | boolean | Non-empty |

- `scripts` counts `<script` elements. `event_handlers` means an inline `on*`
attribute (`onload`, `onerror`, `onclick`, and similar) appears. Both are
ordinary in web pages and meaningful in an SVG image or a message body.
- `base64_bytes` is the total length of base64 runs of at least 256 characters;
`base64_max_run` is the longest single run; `numeric_array_bytes` is the
longest run of only digits and commas (a payload written as a decimal array).
- `payload_types` are the file types the first encoded runs decode to, in the
same vocabulary as `magic_type` (`Zip archive`, `PE executable`, ...). Empty
means none decoded to a recognised format, not that none is a payload.
- `blob_download` means the page builds a `Blob` and hands it to the browser to
save. `download_attr` means a `download` attribute or property is set, and
`auto_click` that script clicks or dispatches an event.
- `js_redirect` and `meta_refresh` mean the page navigates itself.
`password_input` means a password field; `remote_form_action` means a form
posts to an absolute URL.
- `truncated` means the scan stopped at its byte limit (16 MiB); everything else
describes the prefix.

### Explode

| Field | Type | Presence |
Expand All @@ -411,6 +484,7 @@ and a decoded destination; when both links are emitted, it is set on both.
| `vba` | [VBAInfo](#vbainfo) | When set |
| `qr` | array of [QRCode](#qrcode) | Non-empty |
| `ocr_excerpt` | string | Non-empty |
| `pdf` | [PDFInfo](#pdfinfo) | When set |
| `yara_matches` | array of string | Non-empty |
| `archive` | [ArchiveInfo](#archiveinfo) | When set |
| `flavors` | array of string | Non-empty |
Expand All @@ -419,6 +493,33 @@ and a decoded destination; when both links are emitted, it is set on both.
| `truncated` | boolean | Non-empty |
| `truncation_reasons` | array of string | Non-empty |

### PDFInfo

What the analyzer's PDF scanner reported about one file. Unlike the other
findings, it does not aggregate over the tree: a PDF inside an archive carries
its own block on its own child. The analyzer does not return a PDF's text, so
this is everything a rule gets about a PDF's content.

| Field | Type | Presence |
|---|---|---|
| `pages` | integer | Non-empty |
| `words` | integer | Non-empty |
| `images` | integer | Non-empty |
| `encrypted` | boolean | Non-empty |
| `needs_password` | boolean | Non-empty |
| `embedded_files` | integer | Non-empty |
| `links` | integer | Non-empty |
| `phones` | array of string | Non-empty |

`encrypted` includes owner-password restrictions on a readable PDF.
`needs_password` means opening requires a password; the analyzer stops without
examining its content. A locked PDF is unexamined, not empty. An owner password
alone does not make a PDF opaque.
`phones` holds the telephone numbers the analyzer read from the text layer, digits
only, exactly as it reported them (separators and any leading `+` are dropped), at
most 16. They are raw evidence: read
[`enrichments/phone_numbers`](#phonenumbers) instead, which validates them.

### VBAInfo

| Field | Type | Presence |
Expand Down Expand Up @@ -510,6 +611,42 @@ and a decoded destination; when both links are emitted, it is set on both.
| `thread_verification` | [ThreadVerification](#threadverification) | When set |
| `password_in_body` | boolean | Non-empty |
| `detonation` | [Detonation](#detonation) | When set |
| `phone_numbers` | [PhoneNumbers](#phonenumbers) | When a number was found |
| `pii` | [PIIFindings](#piifindings) | When a validated value was found or inspection was truncated |

### PIIFindings

Counts of distinct validated values across the subject, plain body, HTML text,
attachment OCR and attached messages. Quoted and hidden body text count because
that text was sent too. Repeating a value in two body renderings counts once.
The facts contain no values, masked values, prefixes or hashes.

| Field | Type | Presence |
|---|---|---|
| `card_numbers` | integer | Always inside `pii`, including zero |
| `ibans` | integer | Always inside `pii`, including zero |
| `us_ssns` | integer | Always inside `pii`, including zero |
| `truncated` | boolean | Only when inspection was incomplete |

Payment cards require a supported issuer prefix and length and a valid Luhn
checksum. IBANs require registered country length and structure plus valid mod-97
check digits. US Social Security numbers require valid area, group and serial
shapes; common published placeholders are excluded. Dashed `NNN-NN-NNNN` values
need no label, so similarly shaped internal IDs can match. Spaced or contiguous
forms need a preceding SSN or social-security label. These validators recognize
plausible values; they cannot establish that an account or identity exists.

Scanning is bounded to 512 KiB per text, 2 MiB total and 1,000 distinct values per
kind. `truncated: true` makes counts lower bounds, including an all-zero block
when nothing was found before a limit, or attachment extraction was unavailable.
Known parser limits and unavailable extraction inside attached messages also set
the flag. The entire `pii` block is omitted when available-text inspection
completes without a finding. Text inside ordinary document, spreadsheet and PDF
files is not inspected by this detector; OCR and attached-message text are covered.
Counts remain lower bounds for excluded file text even when `truncated` is absent.
Missing findings are not assurance that all attachments were examined.

See [outbound PII detections](detections.md#outbound-pii-detections).

### SenderProfile

Expand Down Expand Up @@ -575,6 +712,77 @@ lookup failure must not be treated as evidence against a message.
| `vip_hit` | string | Non-empty |
| `org_domain_distance` | integer | When set |
| `brand_domain_distance` | integer | When set |
| `display_name_brand` | string | Non-empty |

`display_name_brand` names a well-known brand (for example `paypal`, `microsoft`,
`docusign`) when the sender's display name is that brand plus only role words such
as Support, Security or Team, and the sender's domain does not belong to the brand.
A personal name beside the brand ("John Smith via PayPal"), another company, or the
brand's own mail (regional domains included) never sets it. Consumer mailbox
domains such as `outlook.com` and `gmail.com` are never a brand's own. It needs no
organization configuration.

### PhoneNumbers

The telephone numbers a message asks its reader to use. Callback phishing carries
no link and often no attachment payload: the harm is the phone call, so the number
is a fact of its own. Numbers are normalised (`+18885550142`), validated (a
numbering-plan shape, not an order number), distinct and bounded.

| Field | Type | Presence |
|---|---|---|
| `body` | [PhoneSource](#phonesource) | When a number was found |
| `attachments` | [PhoneSource](#phonesource) | When a number was found |
| `pdf` | [PhoneSource](#phonesource) | When a PDF carried a valid number |
| `pdf_documents` | array of [PDFPhoneDocument](#pdfphonedocument) | When a root PDF carried a valid number |

`body` reads the newest segment the sender wrote (quoted history never supplies a
number). `attachments` unions the text recovered from images by OCR, the numbers
the PDF scanner read, and the body of attached messages: its number fields
(`count`, `numbers`, `toll_free`) are a union, while its context fields
(`call_to_action`, `toll_free_call_to_action`, `lure_terms`, `text_chars`) all
come from the one text that looks most like a callback lure, never a mixture of
unrelated texts. `pdf` holds only the numbers read from PDF text layers, so a rule
about a PDF is not satisfied by a number in an image. A PDF contributes numbers
only, with no `call_to_action` or `lure_terms`, and only North American numbers can
be recognised from it. `pdf_documents` binds validated numbers and shape to each
root PDF separately; a number from one file cannot satisfy a rule about another
file's page or word count. Absent means no number was found in the sources that were
available: an image that was not OCRed says nothing.

### PDFPhoneDocument

One root PDF attachment, up to 128 records. Small candidates are retained first
when the limit is reached. OCR, attached-message numbers and cross-document joins
do not supply these facts.

| Field | Type | Presence |
|---|---|---|
| `count` | integer | Always; distinct validated numbers in this PDF |
| `pages` | integer | Always; analyzer-reported page count |
| `words` | integer | Always; analyzer-reported word count |

### PhoneSource

| Field | Type | Presence |
|---|---|---|
| `count` | integer | Non-empty |
| `numbers` | array of string | Non-empty |
| `toll_free` | boolean | Non-empty |
| `call_to_action` | boolean | Non-empty |
| `toll_free_call_to_action` | boolean | Non-empty |
| `lure_terms` | integer | Non-empty |
| `text_chars` | integer | Non-empty |

`numbers` keeps at most five. `call_to_action` is true when a number sits within
about 80 characters of a verb that tells the reader to use it (call, dial,
contact, reach, helpline...). `toll_free_call_to_action` is true when the same
number is toll-free and has that call to action; `toll_free` and `call_to_action`
alone can belong to two different numbers. `lure_terms` counts distinct billing and support
words in the same text (purchase, subscription, invoice, refund, renew, charged,
antivirus, ...); one is ordinary commerce, four beside a phone number is the
callback-lure shape. The vocabulary is English. `text_chars` is the length of the
text examined.

### Detonation

Expand Down
Loading