From dd68cd9fa6e46f27dacd6c0e5d34c4163596b0af Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Wed, 30 Sep 2026 23:23:06 +0000 Subject: [PATCH 1/7] Email Security: document callback-phishing and HTML-smuggling facts Adds HTMLIndicators, PDFInfo, PhoneNumbers/PhoneSource, visible_chars and lookalike.display_name_brand to the rule reference, plus a detections section on how the default rules read callback phishing and HTML smuggling. Co-Authored-By: Claude Sonnet 5.5 --- docs/email-security/detections.md | 38 +++++++ docs/email-security/rule-reference.md | 145 ++++++++++++++++++++++++++ 2 files changed, 183 insertions(+) diff --git a/docs/email-security/detections.md b/docs/email-security/detections.md index 0fa15ac40..c61815fb8 100644 --- a/docs/email-security/detections.md +++ b/docs/email-security/detections.md @@ -285,6 +285,44 @@ attachment threats, suspicious content, detonation evidence and graymail. They are ordinary D&R rules over the Message Data Model, not a separate engine. See [Mail Rules](custom-rules.md) for the format, IaC and explicit restoration. +### Callback phishing and HTML smuggling + +Two attack classes are invisible to a scanner that only looks at links and known-bad +files, so the defaults read them from structure. + +**Callback phishing** (telephone-oriented attack delivery) is an invoice, renewal or +"your device is infected" notice whose only action is a phone number. The defaults +read the number as a fact ([PhoneNumbers](rule-reference.md#phonenumbers)) from the +body, from the text of attached images, from numbers in a PDF, and from attached +messages, and combine it with a call to action, billing vocabulary, how short the +message is, and whether the sender looks odd (a free-mail address, a young domain, a +failing DMARC result, a Reply-To elsewhere). A legitimate vendor's receipt carries the +same words and a support number, which is why a sender oddity is required and an +established sender is never read as a lure. The callback rules describe one +observation and do not add up: a message that trips all of them scores as the +strongest. A PDF's wording is not available to rules, so the PDF rule judges a short +PDF by its shape and its numbers; numbers in a PDF are recognised for North American +formats only. + +**HTML smuggling** is a web page, often an `.html` or `.svg` attachment, that builds +the real payload in the victim's browser. The parser scans every HTML-like attachment +and HTML body in full and reports encoded data, the type that data decodes to, +decoding and download primitives, redirects and password forms +([HTMLIndicators](rule-reference.md#htmlindicators)). Defaults flag a page that +decodes encoded data and saves it, a page whose encoded data is an archive or +executable, a page that builds its own decoder, an HTML sign-in page delivered as a +file, a tiny redirect page, an SVG that carries script, and a message body that runs +a decoder. A single-file report or export tool that embeds data and offers a download +button matches the same facts as a smuggling page and is scored as suspicious, not +malicious, unless its data also decodes to a recognisable payload. + +Display-name brand impersonation ("PayPal Support" over an unrelated address), +advance-fee and extortion text, voicemail and fax lures, free-hosting and +open-redirector links, internationalised look-alike domains, OneNote files, locked +PDFs with the password in the message, and web pages hidden inside archives from a +stranger are covered by further defaults. **Email Security → Rules** shows every rule's +conditions and false-positive notes. + A verdict's `engine_version` is a SHA-256 fingerprint of the scoring rules, resolved thresholds, exclusions, VIPs, threat-feed references and clustering policy, and linked parsing/enrichment library build. diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index e8cb83b3a..028b3edc5 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -75,6 +75,11 @@ Scoring classes require `weight` from 1 to 100; graymail records must omit it. | What did attachment analysis actually inspect? | Scope `attachments`; inspect `explode/scanners` before interpreting scanner-specific results | | Is this a known sender? | `enrichments/sender_profile/prevalence` (`none`, `new`, `rare`, `common`) | | Is the sender impersonating an organization? | `enrichments/lookalike/org_domain_distance`, `enrichments/lookalike/vip_hit` | +| Is the display name a well-known brand over an address that is not the brand's? | `enrichments/lookalike/display_name_brand` | +| Does the message tell the reader to call a number (callback phishing)? | `enrichments/phone_numbers/body` and `enrichments/phone_numbers/attachments`; read `call_to_action`, `lure_terms`, `toll_free`. See [PhoneNumbers](#phonenumbers) | +| Is this attachment a web page that builds a file in the browser (HTML smuggling)? | Scope `attachments`; read `html/base64_bytes`, `html/payload_types`, `html/blob_download`, `html/atob`. See [HTMLIndicators](#htmlindicators) | +| Is a PDF short, locked, or carrying phone numbers? | Scope `attachments`; read `explode/pdf`. See [PDFInfo](#pdfinfo) | +| How much of the message does the reader actually see? | `body/current_thread/visible_chars` | | What happened after a link was fetched? | `enrichments/detonation`; see [Link detonation](detections.md#link-detonation) | | Was parsing or analysis incomplete? | `_meta/truncations`, `_meta/errors`, `_meta/explode_timeout`, `body/truncated` | @@ -315,6 +320,7 @@ the pipeline includes the whole body in `current_thread` for an unverified reply | `inner_text` | string | Non-empty | | `display_text` | string | Non-empty | | `charset` | string | Non-empty | +| `indicators` | [HTMLIndicators](#htmlindicators) | When the scan found something | ### PlainBody @@ -329,8 +335,15 @@ the pipeline includes the whole body in `current_thread` for an unverified reply | `text` | string | Non-empty | | `renderings` | array of [ThreadRendering](#threadrendering) | Non-empty | | `visible_text` | string | Non-empty | +| `visible_chars` | integer | Non-empty | | `links` | array of [Link object](#link) | Non-empty | +`visible_chars` is the number of non-whitespace characters in `visible_text`. It +measures how much the reader is shown, and blank-line padding cannot inflate it. +A mail rule cannot compute a length itself because its regular expressions are +limited to short repeat counts, so use this field for "the message is a two-line +note" conditions. + ### ThreadRendering | Field | Type | Presence | @@ -399,8 +412,67 @@ and a decoded destination; when both links are emitted, it is set on both. | `tlsh` | string | Non-empty | | `magic_type` | string | Non-empty | | `is_inline` | boolean | Non-empty | +| `html` | [HTMLIndicators](#htmlindicators) | When the part is HTML-like | | `explode` | [Explode](#explode) | When set | +`html` is computed from the attachment's own bytes when the message is parsed, so +it does not depend on attachment analysis. It is present exactly when the part +was recognised as HTML-like (an HTML or SVG file, or any part whose first bytes +open like a web page, whatever it is called). An absent block means "not a web +page", never "a clean web page". + +### HTMLIndicators + +Structural facts from a bounded scan of an HTML-like document: an attachment, or +the HTML body. HTML smuggling is a structure, not a string: a large blob of +encoded data, a few lines of script that decode it, and a browser call that +saves the result as a download. No single field below is a finding, because HTML +exports embed images as base64 and ordinary pages call `atob`. Combine them. +Keyword fields are matched on a normalised view of the document (lower case, +whitespace and quotes removed), so spacing and case do not matter, but splitting a +word across string pieces (`'at'+'ob'`) defeats them. The encoded data and its +decoded type are reported for that reason. + +| Field | Type | Presence | +|---|---|---| +| `scripts` | integer | Non-empty | +| `event_handlers` | boolean | Non-empty | +| `base64_bytes` | integer | Non-empty | +| `base64_max_run` | integer | Non-empty | +| `numeric_array_bytes` | integer | Non-empty | +| `payload_types` | array of string | Non-empty | +| `atob` | boolean | Non-empty | +| `eval` | boolean | Non-empty | +| `from_char_code` | boolean | Non-empty | +| `unescape` | boolean | Non-empty | +| `document_write` | boolean | Non-empty | +| `blob_download` | boolean | Non-empty | +| `download_attr` | boolean | Non-empty | +| `auto_click` | boolean | Non-empty | +| `js_redirect` | boolean | Non-empty | +| `meta_refresh` | boolean | Non-empty | +| `password_input` | boolean | Non-empty | +| `remote_form_action` | boolean | Non-empty | +| `truncated` | boolean | Non-empty | + +- `scripts` counts ` Date: Wed, 30 Sep 2026 23:36:22 +0000 Subject: [PATCH 2/7] Email Security docs: PDF-only phone source and toll-free call to action Co-Authored-By: Claude Sonnet 5.5 --- docs/email-security/rule-reference.md | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index 028b3edc5..d08e5c8e7 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -694,10 +694,16 @@ numbering-plan shape, not an order number), distinct and bounded. |---|---|---| | `body` | [PhoneSource](#phonesource) | When a number was found | | `attachments` | [PhoneSource](#phonesource) | When a number was found | +| `pdf` | [PhoneSource](#phonesource) | When a PDF carried a valid number | `body` reads the newest segment the sender wrote (quoted history never supplies a number). `attachments` unions the text recovered from images by OCR, the numbers -the PDF scanner read, and the body of attached messages. A PDF contributes numbers +the PDF scanner read, and the body of attached messages: its number fields +(`count`, `numbers`, `toll_free`) are a union, while its context fields +(`call_to_action`, `toll_free_call_to_action`, `lure_terms`, `text_chars`) all +come from the one text that looks most like a callback lure, never a mixture of +unrelated texts. `pdf` holds only the numbers read from PDF text layers, so a rule +about a PDF is not satisfied by a number in an image. A PDF contributes numbers only, with no `call_to_action` or `lure_terms`, and only North American numbers can be recognised from it. Absent means no number was found in the sources that were available: an image that was not OCRed says nothing. @@ -710,12 +716,15 @@ available: an image that was not OCRed says nothing. | `numbers` | array of string | Non-empty | | `toll_free` | boolean | Non-empty | | `call_to_action` | boolean | Non-empty | +| `toll_free_call_to_action` | boolean | Non-empty | | `lure_terms` | integer | Non-empty | | `text_chars` | integer | Non-empty | `numbers` keeps at most five. `call_to_action` is true when a number sits within about 80 characters of a verb that tells the reader to use it (call, dial, -contact, reach, helpline...). `lure_terms` counts distinct billing and support +contact, reach, helpline...). `toll_free_call_to_action` is true when the same +number is toll-free and has that call to action; `toll_free` and `call_to_action` +alone can belong to two different numbers. `lure_terms` counts distinct billing and support words in the same text (purchase, subscription, invoice, refund, renew, charged, antivirus, ...); one is ordinary commerce, four beside a phone number is the callback-lure shape. The vocabulary is English. `text_chars` is the length of the From 9497c3c7e2e23eef3fd20a0c25e7c87fbfeaab82 Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Thu, 1 Oct 2026 01:52:34 +0000 Subject: [PATCH 3/7] Describe PII count facts and optional outbound DLP detections --- docs/email-security/detections.md | 20 ++++++++++++++++ docs/email-security/rule-reference.md | 33 +++++++++++++++++++++++++++ 2 files changed, 53 insertions(+) diff --git a/docs/email-security/detections.md b/docs/email-security/detections.md index c61815fb8..b79a89cfe 100644 --- a/docs/email-security/detections.md +++ b/docs/email-security/detections.md @@ -330,6 +330,26 @@ Changing rule content or scoring policy changes the fingerprint. It identifies the decision configuration; it is not a promise that an external lookup feed or other message enrichment is unchanged. +### Outbound PII detections + +The optional email DLP pack adds outbound detections for validated payment card +numbers, IBANs and US Social Security numbers, plus a bulk detection when any +one kind has at least ten distinct values. Bulk counts are per kind: four cards, +four IBANs and four SSNs do not meet the bulk threshold. Bulk detections fire +alongside the matching single-kind detection. These are platform D&R rules on +`EMAIL_MESSAGE`, separate from the engine verdict; installing them does not +change the verdict or automatically move mail. + +The rules read [PIIFindings](rule-reference.md#piifindings), which stores counts +only. A detection still carries the originating email event, whose body can +contain the actual sensitive values; the count facts do not redact that body. +Plan detection access and outputs accordingly. IBANs commonly occur on ordinary +invoices, and dashed SSN-shaped internal IDs can match. Tune the optional rules +for your organization rather than treating a match as proof of malicious intent. +The detector covers message text, OCR and attached messages, but does not inspect +text inside ordinary document, spreadsheet or PDF files. Incomplete inspection +is marked `enrichments/pii/truncated`; the counts then describe only what was seen. + ## Link detonation Static link features answer what a URL *looks* like. Detonation answers where it diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index d08e5c8e7..98756d5d3 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -76,6 +76,7 @@ Scoring classes require `weight` from 1 to 100; graymail records must omit it. | Is this a known sender? | `enrichments/sender_profile/prevalence` (`none`, `new`, `rare`, `common`) | | Is the sender impersonating an organization? | `enrichments/lookalike/org_domain_distance`, `enrichments/lookalike/vip_hit` | | Is the display name a well-known brand over an address that is not the brand's? | `enrichments/lookalike/display_name_brand` | +| Does it contain validated payment cards, IBANs or US Social Security numbers? | `enrichments/pii/card_numbers`, `ibans`, `us_ssns`; counts only, with `truncated` for incomplete inspection. See [PIIFindings](#piifindings) | | Does the message tell the reader to call a number (callback phishing)? | `enrichments/phone_numbers/body` and `enrichments/phone_numbers/attachments`; read `call_to_action`, `lure_terms`, `toll_free`. See [PhoneNumbers](#phonenumbers) | | Is this attachment a web page that builds a file in the browser (HTML smuggling)? | Scope `attachments`; read `html/base64_bytes`, `html/payload_types`, `html/blob_download`, `html/atob`. See [HTMLIndicators](#htmlindicators) | | Is a PDF short, locked, or carrying phone numbers? | Scope `attachments`; read `explode/pdf`. See [PDFInfo](#pdfinfo) | @@ -608,6 +609,38 @@ most 16. They are raw evidence: read | `password_in_body` | boolean | Non-empty | | `detonation` | [Detonation](#detonation) | When set | | `phone_numbers` | [PhoneNumbers](#phonenumbers) | When a number was found | +| `pii` | [PIIFindings](#piifindings) | When a validated value was found or inspection was truncated | + +### PIIFindings + +Counts of distinct validated values across the subject, plain body, HTML text, +attachment OCR and attached messages. Quoted and hidden body text count because +that text was sent too. Repeating a value in two body renderings counts once. +The facts contain no values, masked values, prefixes or hashes. + +| Field | Type | Presence | +|---|---|---| +| `card_numbers` | integer | Always inside `pii`, including zero | +| `ibans` | integer | Always inside `pii`, including zero | +| `us_ssns` | integer | Always inside `pii`, including zero | +| `truncated` | boolean | Only when inspection was incomplete | + +Payment cards require a supported issuer prefix and length and a valid Luhn +checksum. IBANs require registered country length and structure plus valid mod-97 +check digits. US Social Security numbers require valid area, group and serial +shapes; common published placeholders are excluded. Dashed `NNN-NN-NNNN` values +need no label, so similarly shaped internal IDs can match. Spaced or contiguous +forms need a preceding SSN or social-security label. These validators recognize +plausible values; they cannot establish that an account or identity exists. + +Scanning is bounded to 512 KiB per text, 2 MiB total and 1,000 distinct values per +kind. `truncated: true` makes counts lower bounds, including an all-zero block +when nothing was found before a limit. The entire `pii` block is omitted when +inspection completes without a finding. Text inside ordinary document, spreadsheet +and PDF files is not inspected by this detector; OCR and attached-message text +are covered. Missing findings are not assurance that all attachments were examined. + +See [outbound PII detections](detections.md#outbound-pii-detections). ### SenderProfile From 0a0e27fa079ef5bf99a32d148db0bd198634829a Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Thu, 1 Oct 2026 03:10:00 +0000 Subject: [PATCH 4/7] Clarify delayed attachment evidence in outbound PII detections --- docs/email-security/detections.md | 3 +++ docs/email-security/rule-reference.md | 2 +- 2 files changed, 4 insertions(+), 1 deletion(-) diff --git a/docs/email-security/detections.md b/docs/email-security/detections.md index b79a89cfe..1b014402d 100644 --- a/docs/email-security/detections.md +++ b/docs/email-security/detections.md @@ -349,6 +349,9 @@ for your organization rather than treating a match as proof of malicious intent. The detector covers message text, OCR and attached messages, but does not inspect text inside ordinary document, spreadsheet or PDF files. Incomplete inspection is marked `enrichments/pii/truncated`; the counts then describe only what was seen. +Deferred attachment scans refresh the stored facts but do not replay the initial +`EMAIL_MESSAGE` evaluation. A DLP match therefore describes evidence available +when that event was emitted, rather than every later attachment result. ## Link detonation diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index 98756d5d3..37f5e8b96 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -635,7 +635,7 @@ plausible values; they cannot establish that an account or identity exists. Scanning is bounded to 512 KiB per text, 2 MiB total and 1,000 distinct values per kind. `truncated: true` makes counts lower bounds, including an all-zero block -when nothing was found before a limit. The entire `pii` block is omitted when +when nothing was found before a limit, or attachment extraction was unavailable. The entire `pii` block is omitted when inspection completes without a finding. Text inside ordinary document, spreadsheet and PDF files is not inspected by this detector; OCR and attached-message text are covered. Missing findings are not assurance that all attachments were examined. From a31b652d359a2f02856f55c63f1d523df202e84e Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Thu, 1 Oct 2026 04:23:11 +0000 Subject: [PATCH 5/7] Clarify managed severity coverage and browser attachment gates --- docs/email-security/detections.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/email-security/detections.md b/docs/email-security/detections.md index 1b014402d..eb8bb982b 100644 --- a/docs/email-security/detections.md +++ b/docs/email-security/detections.md @@ -314,7 +314,14 @@ executable, a page that builds its own decoder, an HTML sign-in page delivered a file, a tiny redirect page, an SVG that carries script, and a message body that runs a decoder. A single-file report or export tool that embeds data and offers a download button matches the same facts as a smuggling page and is scored as suspicious, not -malicious, unless its data also decodes to a recognisable payload. +malicious, unless its data also decodes to a recognisable payload. Credential-page +and tiny-redirect defaults require a browser-file extension; source templates and +files with unconventional names can fall outside those two checks. + +In managed pack `0.6.0`, these 22 new rules carry explicit severity. Older managed +rules currently use the informational fallback. Severity is independent of the +verdict, so a malicious verdict from an older rule can still have informational +severity. Display-name brand impersonation ("PayPal Support" over an unrelated address), advance-fee and extortion text, voicemail and fax lures, free-hosting and From 5cb7a83140533838260d3e1502318ae61648563e Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Thu, 1 Oct 2026 04:29:45 +0000 Subject: [PATCH 6/7] Document locked PDF and per-document phone facts accurately --- docs/email-security/rule-reference.md | 24 +++++++++++++++++++++--- 1 file changed, 21 insertions(+), 3 deletions(-) diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index 37f5e8b96..3c54d003a 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -506,12 +506,15 @@ this is everything a rule gets about a PDF's content. | `words` | integer | Non-empty | | `images` | integer | Non-empty | | `encrypted` | boolean | Non-empty | +| `needs_password` | boolean | Non-empty | | `embedded_files` | integer | Non-empty | | `links` | integer | Non-empty | | `phones` | array of string | Non-empty | -`encrypted` marks a PDF that needs a password. The analyzer stops there, so every -other field is zero for a locked PDF: a locked PDF is unexamined, not empty. +`encrypted` includes owner-password restrictions on a readable PDF. +`needs_password` means opening requires a password; the analyzer stops without +examining its content. A locked PDF is unexamined, not empty. An owner password +alone does not make a PDF opaque. `phones` holds the telephone numbers the analyzer read from the text layer, digits only, exactly as it reported them (separators and any leading `+` are dropped), at most 16. They are raw evidence: read @@ -728,6 +731,7 @@ numbering-plan shape, not an order number), distinct and bounded. | `body` | [PhoneSource](#phonesource) | When a number was found | | `attachments` | [PhoneSource](#phonesource) | When a number was found | | `pdf` | [PhoneSource](#phonesource) | When a PDF carried a valid number | +| `pdf_documents` | array of [PDFPhoneDocument](#pdfphonedocument) | When a root PDF carried a valid number | `body` reads the newest segment the sender wrote (quoted history never supplies a number). `attachments` unions the text recovered from images by OCR, the numbers @@ -738,9 +742,23 @@ come from the one text that looks most like a callback lure, never a mixture of unrelated texts. `pdf` holds only the numbers read from PDF text layers, so a rule about a PDF is not satisfied by a number in an image. A PDF contributes numbers only, with no `call_to_action` or `lure_terms`, and only North American numbers can -be recognised from it. Absent means no number was found in the sources that were +be recognised from it. `pdf_documents` binds validated numbers and shape to each +root PDF separately; a number from one file cannot satisfy a rule about another +file's page or word count. Absent means no number was found in the sources that were available: an image that was not OCRed says nothing. +### PDFPhoneDocument + +One root PDF attachment, up to 128 records. Small candidates are retained first +when the limit is reached. OCR, attached-message numbers and cross-document joins +do not supply these facts. + +| Field | Type | Presence | +|---|---|---| +| `count` | integer | Always; distinct validated numbers in this PDF | +| `pages` | integer | Always; analyzer-reported page count | +| `words` | integer | Always; analyzer-reported word count | + ### PhoneSource | Field | Type | Presence | From 55f719160db01f166c8ded0bee61f33069b0b910 Mon Sep 17 00:00:00 2001 From: Maxime Lamothe-Brassard Date: Thu, 1 Oct 2026 09:19:09 +0000 Subject: [PATCH 7/7] Clarify PII counts remain lower bounds for excluded file text --- docs/email-security/detections.md | 6 ++++-- docs/email-security/rule-reference.md | 11 +++++++---- 2 files changed, 11 insertions(+), 6 deletions(-) diff --git a/docs/email-security/detections.md b/docs/email-security/detections.md index eb8bb982b..541b7a1d8 100644 --- a/docs/email-security/detections.md +++ b/docs/email-security/detections.md @@ -354,8 +354,10 @@ Plan detection access and outputs accordingly. IBANs commonly occur on ordinary invoices, and dashed SSN-shaped internal IDs can match. Tune the optional rules for your organization rather than treating a match as proof of malicious intent. The detector covers message text, OCR and attached messages, but does not inspect -text inside ordinary document, spreadsheet or PDF files. Incomplete inspection -is marked `enrichments/pii/truncated`; the counts then describe only what was seen. +text inside ordinary document, spreadsheet or PDF files. Counts remain lower +bounds for these excluded sources even when `enrichments/pii/truncated` is absent. +Known incomplete extraction or parsing, and inspection limits, set that flag; +the counts still describe only the available text. Deferred attachment scans refresh the stored facts but do not replay the initial `EMAIL_MESSAGE` evaluation. A DLP match therefore describes evidence available when that event was emitted, rather than every later attachment result. diff --git a/docs/email-security/rule-reference.md b/docs/email-security/rule-reference.md index 3c54d003a..c1e255fd0 100644 --- a/docs/email-security/rule-reference.md +++ b/docs/email-security/rule-reference.md @@ -638,10 +638,13 @@ plausible values; they cannot establish that an account or identity exists. Scanning is bounded to 512 KiB per text, 2 MiB total and 1,000 distinct values per kind. `truncated: true` makes counts lower bounds, including an all-zero block -when nothing was found before a limit, or attachment extraction was unavailable. The entire `pii` block is omitted when -inspection completes without a finding. Text inside ordinary document, spreadsheet -and PDF files is not inspected by this detector; OCR and attached-message text -are covered. Missing findings are not assurance that all attachments were examined. +when nothing was found before a limit, or attachment extraction was unavailable. +Known parser limits and unavailable extraction inside attached messages also set +the flag. The entire `pii` block is omitted when available-text inspection +completes without a finding. Text inside ordinary document, spreadsheet and PDF +files is not inspected by this detector; OCR and attached-message text are covered. +Counts remain lower bounds for excluded file text even when `truncated` is absent. +Missing findings are not assurance that all attachments were examined. See [outbound PII detections](detections.md#outbound-pii-detections).