diff --git a/docs/cloud-security/code-security/autofix.md b/docs/cloud-security/code-security/autofix.md index 56d4b32d9..fc0fbfc5b 100644 --- a/docs/cloud-security/code-security/autofix.md +++ b/docs/cloud-security/code-security/autofix.md @@ -16,7 +16,9 @@ AutoFix needs: AutoFix to open fix pull requests**. For any other App, add the permissions under the App's **Permissions & events** on GitHub, then have an organization owner approve the change on the installation page; -- an enabled code-scanning policy that selects the repository. +- an enabled code-scanning policy that selects the repository; +- the `cloudsec.respond` permission for whoever asks for the fix. `cloudsec.set` + does not include it. See [Permissions](containment-setup.md#permissions). There is no separate policy switch. Once the permissions are granted, the **Code security** Overview tab stops showing **Automated fixes** as needing @@ -37,8 +39,14 @@ as `available`. limacharlie cloudsec code autofix fnd_2290bab86c1b4d0374d1e2666f64aeca ``` -The request is accepted immediately. The pull request appears a few minutes -later, on a branch named `limacharlie/autofix/-` (for +Each request creates a remediation run of type `open_fix_pr`. A person or API key +with `cloudsec.respond` can request it. The same identity is recorded as both +requester and approver, and the pull request comes back to the run through an +authenticated callback. Asking again before the pull +request is open returns the same run. See +[Remediation runs](containment-setup.md#remediation-runs). + +The pull request appears a few minutes later, on a branch named `limacharlie/autofix/-` (for example `limacharlie/autofix/npm-babel-core` for `@babel/core`). There is at most one open AutoFix pull request per repository and package. @@ -46,7 +54,10 @@ You can only ask for a fix by finding: the package and target version come from LimaCharlie's own scan, never from the request. The finding closes once the pull request is merged and the next scan of the -default branch no longer sees the vulnerable version. +default branch no longer sees the vulnerable version. A merged pull request +alone is not a fix. The run becomes `verified` only when every in-scope +deployment also runs the fixed build. When no deployment is in scope, it ends +`expired` with `pr_merged_unverifiable`. ## What gets edited @@ -91,10 +102,14 @@ Maven versions inherited from a parent POM, and pip pins other than `==`, `===`, ## When no pull request appears -The request is accepted before the work runs, so a refusal is not returned by -the CLI or the console. It is reported as a `cloudsec.code_autofix_refused` -operational event in the organization's event stream, with the reason in its -`error` field. Operational events are off by default: turn them on with +The request is accepted before the work runs, so most refusals are not +returned by the CLI or the console. The run records the reason in its +`failure_reason` (`limacharlie cloudsec remediation get `). It is also +reported as a `cloudsec.code_autofix_refused` operational event in the +organization's event stream, with a sentence in `error`, a reason code in +`reason`, and the run in `remediation_id`. The table below lists the codes that +appear in `error`. The `reason` field and the run's `failure_reason` use the +[fix pull request codes](reasons.md#fix-pull-requests-and-autofix). Operational events are off by default: turn them on with `ops_events: true` in the [`emission` policy](../configuration.md#emission-the-event-feed). Common reasons: @@ -110,3 +125,15 @@ Common reasons: | `autofix_pr_already_open` | A pull request for that package is already open. | | `autofix_budget_exhausted` | The daily limit of 20 AutoFix requests was reached. Requests count when they start, even if they later fail. The limit resets at midnight UTC. | | `autofix_job_failed`, `autofix_pr_failed` | The job or the pull request creation failed. Try again later. | + +Some requests are refused immediately, with an HTTP error: + +| Error | Meaning | +|---|---| +| `403 missing_permission` | You lack `cloudsec.respond`. | +| `422 action_unavailable` | The finding is not a dependency finding AutoFix can raise, or fix pull requests are not enabled for your organization. | +| `429 capacity` | Your organization already has 100 active remediation runs, or this finding has 10. | +| `503 disabled` | Remediation is not enabled for your organization yet. | + +See [Unknown, partial and refusal reasons](reasons.md#remediation-runs) for +every code. diff --git a/docs/cloud-security/code-security/containment-setup.md b/docs/cloud-security/code-security/containment-setup.md new file mode 100644 index 000000000..46f344beb --- /dev/null +++ b/docs/cloud-security/code-security/containment-setup.md @@ -0,0 +1,234 @@ +# Configure evidence, lineage and remediation + +This page covers the settings behind the evidence chain, image lineage, live +pull-request impact, runtime checks and remediation runs. It lists the +permissions each one needs, what to grant on each connection, and which +policies control them. For what these features promise, see +[What Code Security guarantees](guarantees.md). + +!!! note "Availability" + These capabilities are enabled region by region. Until yours is on, the + routes below answer `feature_disabled`, `disabled` or `codesec_disabled`. + Scanning, pull-request checks and the rest of Code Security are unaffected. + +## Permissions + +| Permission | Allows | +|---|---| +| `cloudsec.get` | Reading findings, the evidence chain, coverage, live impact, image lineage and remediation runs. Running a runtime check, which starts collecting package evidence on the finding's sensors and changes nothing else. | +| `cloudsec.set` | Writing policies, pushing Terraform maps and build provenance, rescans. | +| `cloudsec.respond` | Requesting, approving, rejecting and cancelling remediation runs, pressing **Open AutoFix PR**, and writing a `response` policy. | + +`cloudsec.respond` is separate on purpose. `cloudsec.set` does not include it, +because it lets a person or API key change your repositories, detection rules and +endpoints. The **Owner** and **Administrator** roles include it. Operator, +Viewer and Basic do not. A user or API key that got its permissions before +`cloudsec.respond` existed does not receive it automatically: grant it +explicitly, or assign the role again. + +Writing a `cloudsec_policy` record of type `response` needs both +`cloudsec.set` and `cloudsec.respond`, because that record decides what a +remediation may do. + +## Connections and permissions + +All of these are read-only unless the row says otherwise. + +| Connection | Grant | Needed for | +|---|---|---| +| GitHub App | **Contents: Read-only** | Scanning. See [GitHub setup](../provider-setup/github.md). | +| GitHub App | **Attestations: Read-only** | Verifying GitHub Actions artifact attestations, so an image's link to its repository can be `verified`. After adding it to an existing App, an organization owner must approve the new permission on the installation page. Without it, scanning continues and lineage stays inferred or asserted. | +| GitHub App | **Contents** and **Pull requests: Read and write** | AutoFix and fix pull requests. Write access, granted only if you want it. | +| Google Cloud | `roles/artifactregistry.reader` on the project that hosts the image | Pulling private images to scan them. See [Container image scanning](../provider-setup/gcp.md#container-image-scanning-by-code-security). | +| Google Cloud | `roles/containeranalysis.occurrences.viewer` (Artifact Analysis Occurrences Viewer) on the project that stores the image | Reading Cloud Build provenance so the build can be verified. A grant on the project where the workload runs is not enough if the image lives elsewhere. | +| Cloud accounts that run your workloads | The standard Cloud Security collection roles | Seeing which digest each Cloud Run service and GKE workload runs. Only Cloud Run and GKE digests are observed today. | +| LimaCharlie sensors | A sensor on the host or node | Runtime checks. Without one the answer is `unknown` with `no_sensors`. | + +GitLab and Bitbucket connections are scanned with their read tokens. Pull-request +checks, fixes and other writes on GitLab and Bitbucket are not enabled yet. + +### Webhooks + +Push rescans and pull-request checks on GitHub arrive through the App's webhook. +Set it up as described in [Pull-request checks and push +rescans](pull-requests.md#set-up-webhook). LimaCharlie checks the +HMAC-SHA256 signature on every delivery, refuses a delivery without a valid one +with `401`, and refuses a repeated delivery for 24 hours. + +## Image lineage + +Code Security tries to link each running image digest to the source it was built +from, without any change to your build pipeline. Every link carries a `status`: + +| Status | How it is established | +|---|---| +| `inferred` | The image's build steps and file paths match a Dockerfile in a repository you connected. It names the repository and the scanned commit whose Dockerfile matched. It does not identify the build commit. | +| `asserted` | The image carries OCI `org.opencontainers.image.source` and `revision` labels, or an unsigned build record, pointing at a connected repository. Anyone who can build the image can write these. | +| `verified` | A signature checks out for that exact digest and a repository in your connections. It comes from Google Cloud Build, from GitHub Actions artifact attestations, or from a statement you pushed that matches a trusted signing identity in your `provenance_trust` policy. | +| `ambiguous` | More than one source matches, or two sources disagree. Neither is used. | +| `unknown` | No usable evidence, or the evidence is past its window. | + +Among the evidence Code Security collects itself, only Google Cloud Build and +GitHub Actions signatures make a link `verified`. A label or an unsigned build +claim stays `asserted` however it is written. + +The lineage coverage line counts only your own images: images in your own +registries or cloud projects, or matched to your repositories. Public images +from well-known vendor registries that match none of your connections are +counted separately as `third_party`. Images whose ownership cannot be settled +are `ownership_unknown`, and the percentage is withheld until your +connections settle it. + +### Build provenance + +You can also push signed build statements yourself. This is optional. + +- `POST /code/provenance` accepts `lc-build-provenance/v1`, SLSA provenance v1 + and Sigstore bundles, up to 1 MiB, for full sha256 digests and full commits. + It needs `cloudsec.set`. +- A `provenance_trust` policy names the builders you trust, per repository. + A statement verified against a Sigstore identity or GitHub's Sigstore + instance is `verified`. A statement signed with a public key you uploaded is + `asserted`. +- LimaCharlie sets `trust`, `verified`, `signer` and similar fields itself and + refuses a statement that tries to set them. It never fetches a URL found in a + statement. +- Two statements that disagree about the source give `ambiguous` and a coverage + note. The later one does not win. + +## Terraform maps + +Code-to-cloud attribution matches a Terraform declaration to a live resource by +exact identity. When a name is computed at apply time, the scan alone cannot +resolve it, and the finding's attribution reason is `unresolved_name`. A map +fills the gap. + +LimaCharlie never accepts a state or plan file. You extract a map on your own +machine or CI runner and push only the map: + +```bash +terraform show -json > terraform.json + +limacharlie cloudsec code iac-map extract --input terraform.json \ + --source-kind state_identity --repository acme/infra \ + --commit 3f1c2a9e0b7d4c5a6e8f9a0b1c2d3e4f5a6b7c8d --workspace default > map.json + +limacharlie cloudsec code iac-map push --input map.json +``` + +- `extract` runs offline. It does not run Terraform, read your environment, + contact a server or authenticate. It needs the `iac-map-extract` binary on + your `PATH`. +- `--source-kind state_identity` keeps resource identities only. + `plan_desired` adds allowlisted desired settings (yes/no values) so drift can + be checked. +- Every value Terraform marks sensitive or unknown is dropped. On push, a map + with a secret-looking key or a value over 4 KiB is refused. +- `push` needs `cloudsec.set`, is limited to 20 MiB and 30 pushes a minute, and + waits until the map is published. Pushing the same map again writes nothing. + A partial map only adds; it never deletes earlier mappings. + +## Pull-request disclosure + +Pull-request checks can show what a change touches in production. The +`pr_live_context` field of the [code-scanning policy](policy.md) decides how +much: + +| Value | The check shows | +|---|---| +| `off` (default) | Nothing about live resources. The check is exactly what it was without this feature. | +| `risk_summary` | Counts, the environment, yes/no exposure, privilege and sensitivity facts, the highest severity, and a link that needs a LimaCharlie login. No resource names, account IDs, IP addresses or sensor IDs. | +| `resource_details` | The above plus per-declaration impact details. Resource names are still left out of the pull request. The authenticated impact view can show them. | + +When several policies select one repository, the one that discloses least wins. +The impact lookup has a 2-second budget. If it fails or runs out of time, the +check still publishes its normal scan verdict, unchanged, with a note that the +live context is missing. + +## Remediation runs + +Every change Code Security makes to your systems is a remediation run. A run +states its target, which the server derives from the finding. You cannot point +a run at another target. It waits for a person or API key with +`cloudsec.respond` to approve it, and its outcome comes back through an +authenticated callback that the run records. Pre-approval is not available: +every run needs an approval from one of those identities. + +| Action | What it does | Needs | +|---|---|---| +| `open_fix_pr` | Opens a pull request that upgrades a vulnerable dependency or base image. | The GitHub write permissions above. | +| `notify_ticket` | Sends a notification or opens a ticket. | A `code-notify-ticket-v1` playbook. | +| `temporary_detection` | Installs a detection rule that expires on its own. | A `code-temporary-detection-v1` playbook. | +| `isolate_endpoint` | Network-isolates a sensor for a limited time. | A `code-isolate-endpoint-v1` playbook that lists the sensors it may act on. | + +Requesting a run and deciding it are separate calls: + +```bash +limacharlie cloudsec remediation create fnd_0123abcd --action open_fix_pr +limacharlie cloudsec remediation approve rem_0123abcd # prints what you are approving +limacharlie cloudsec remediation approve rem_0123abcd --confirm +``` + +The approval binds to the run's current target and generation. If either +changes after you reviewed it, the approval is refused and you review again. + +Limits: 100 active runs per organization and 10 per finding. A run waits at +most 72 hours for approval and is monitored for up to 7 days after it acts. + +### AutoFix is a remediation run + +Pressing **Open AutoFix PR** creates an `open_fix_pr` run. A person or API key +with `cloudsec.respond` is recorded as both requester and approver. +An API key can approve this run just like an interactive user. The pull request +and its outcome come back through the same callback as any other run. See +[AutoFix pull requests](autofix.md). + +### Response playbooks + +Notify, ticket, temporary detection and isolation run LimaCharlie-managed +templates that you install with a `response` policy. You cannot write a rule +body, command or target through it. + +```yaml +policy_type: response +response: + playbooks: + - template: code-temporary-detection-v1 + version: 1 + mode: live + approval: human + max_duration_seconds: 86400 + - template: code-notify-ticket-v1 + version: 1 + channels: [notify, ticket] +``` + +| Field | Meaning | +|---|---| +| `template`, `version` | The managed template and its version (`1`). A run is bound to the installation it was approved under. If the installation changes before the run acts, the run stops with `installation_changed`. | +| `mode` | `dry_run` (the default) records exactly what would happen and does nothing. `live` acts. | +| `approval` | `human`, the only mode. | +| `max_duration_seconds` | For temporary controls. 5 minutes to 7 days, default 24 hours. Isolation is capped at 4 hours. | +| `channels` | For `code-notify-ticket-v1`: `notify`, `ticket`, or both. | +| `sensors` | For `code-isolate-endpoint-v1`, required: 1 to 10 sensor IDs. The target still comes from the finding. This list only narrows it. | + +At most 8 playbooks, one per template. Isolation also needs an approval less +than 10 minutes old at the moment it acts. + +LimaCharlie never merges, deploys, rolls back or revokes credentials for you. +There is no template for those. + +## Exclusions + +Several settings keep things out of Code Security. They do different things: + +| Setting | Effect | +|---|---| +| `repos.exclude` in the [code-scanning policy](policy.md#fields) | The repository is not scanned. The console's **Exclude from scanning** writes it. Exclude always wins over include. | +| `paths` on a [code rule](code-rules.md) | That rule does not run on those paths. | +| A [`suppression` policy](../configuration.md#suppression-finding-disposition-policy) or a VEX statement | Matching findings get a disposition. They stay recorded. | +| A [`collection` exclusion](../configuration.md#exclusions-the-escape-hatch) | The cloud resources it matches are removed from inventory. | + +A collection exclusion also removes those resources from Code Security's view of +your cloud. Evidence that needs them then reads `resource_not_collected` or +`workload_not_resolved`. It never reads as "not exposed" or "not running". diff --git a/docs/cloud-security/code-security/data-handling.md b/docs/cloud-security/code-security/data-handling.md new file mode 100644 index 000000000..0804b32d2 --- /dev/null +++ b/docs/cloud-security/code-security/data-handling.md @@ -0,0 +1,114 @@ +# Data handling and privacy + +This page lists what Code Security reads, what it keeps, how long it keeps it, +and what happens to it when your organization leaves. It covers the scan lane +described in [Code Security](index.md) and the evidence, lineage and remediation +features described in [What Code Security guarantees](guarantees.md). + +## What it reads + +| Source | What is read | How | +|---|---|---| +| Your repositories | The commit being scanned | A short-lived fetch job downloads the commit and packs it. A separate sandboxed scan job reads the package and is then destroyed. See [How your code is handled](index.md#how-your-code-is-handled). | +| Container images | Image layers and metadata, by digest | The fetch job pulls the image with a registry credential. The scan sandbox never holds that credential. Only images referenced by digest are scanned. | +| Image build records | OCI labels, build attestations, and on Google Cloud the Artifact Analysis occurrences for an image | Read-only, with the permissions listed in [Configure connections and permissions](containment-setup.md#connections-and-permissions). Used to link a running image to its source repository and, where a build record says so, its commit. The lineage status shows how strong that link is. | +| Your cloud accounts | Inventory and configuration | The existing Cloud Security collectors, with read-only credentials. Collection credentials are never used to change anything. | +| LimaCharlie sensors | Process and module activity on hosts that run an affected package | Summarized per sensor into short-lived package evidence (15 minutes). Asking for a runtime check tells the finding's sensors which packages to watch. It never changes the finding. | +| Terraform | A map you produce and push yourself | See [Terraform state and plans](#terraform-state-and-plans). LimaCharlie never receives a state or plan file. | +| Build provenance | Statements you push | Normalized and stored without the raw envelope. LimaCharlie never fetches a URL found inside a statement. | + +## What it keeps + +| Data | Kept | Not kept | +|---|---|---| +| Findings | Rule, severity, file path and line numbers, package and version, a fingerprint, and a hash of the matched text | The matched source lines. Code rules you write keep their own title text, with every captured value replaced by `[code]`. | +| Secret findings | Rule, file, line, a salted hash, and whether the secret is also in history | The secret value. No field on a finding can hold it. Results pushed in SARIF or other foreign formats have their secret findings dropped, because those formats carry the value. | +| Infrastructure-as-code findings | Check, file, lines, the resource address, and the resource name as it is written in the code (for example a bucket name) | The source line. | +| SBOM | One CycloneDX document per repository, as a downloadable file | | +| Terraform map | Resource addresses, identities and allowlisted settings | Any sensitive, unknown or secret-looking value. | +| Build provenance | Repository, commit, builder, workflow, digest and the evidence level | The raw signed envelope. | +| Remediation runs | Who requested and approved, the resolved targets, the state, bounded structured evidence and the callback result | Command output. | + +Source bytes do pass through storage while a scan runs. The fetch job uploads +the commit package for the scan job to read. LimaCharlie deletes it when that +scan ends, and a storage rule deletes anything left after one day. AutoFix +works the same way with the manifest and lockfile it edits. Deleted objects are +not recoverable: the bucket has no soft delete and no versioning. + +## Secrets and credentials + +- **Secrets found in code** are stored only as a salted SHA-256. The salt is + derived for your organization, and the scan sandbox receives only that + derived value. +- **Liveness** comes from the provider's own verdict. LimaCharlie never + presents a found credential to its issuer to test it. +- **Your connection credentials** (GitHub App private keys, GitLab and + Bitbucket tokens, webhook secrets) live in your organization's + [secrets](../../7-administration/config-hive/secrets.md) and are referenced + by name. They never reach the scan sandbox or the browser. +- **The GitHub fetch token** covers one repository and expires within an hour. + +### Terraform state and plans + +State and plan files contain secrets, so LimaCharlie refuses them. The API +recognizes raw Terraform JSON and answers `400 iac_map_raw_terraform` with the +command to run instead. + +You run the extractor yourself, on your own machine or CI runner. It does not +run Terraform, read your environment, contact a server or upload anything. It +reads `terraform show -json` output and writes a map with resource identities +and allowlisted settings. It drops every value Terraform marks sensitive or +unknown. Pushing the map is a separate command. On push, LimaCharlie rejects any +map that carries a key such as `password`, `secret`, `token`, `private_key` or +`connection_string`, or a value over 4 KiB. + +## Pull-request output + +Pull-request checks show the live consequence of a change only if you turn it +on with `pr_live_context`. It is `off` by default. In `risk_summary` mode the +check shows counts, the environment, yes/no exposure facts and the highest +severity, with a link that requires a LimaCharlie login. It shows no resource +names, account IDs, IP addresses or sensor IDs. `resource_details` adds +per-declaration detail, still without resource names. See +[Configure connections and permissions](containment-setup.md#pull-request-disclosure). + +## AI + +AutoFix does not send your code to a language model. It is deterministic. It +raises one dependency to one version and never runs a package manager. No +AI-generated fix feature is available. + +## Retention + +| Data | Retention | +|---|---| +| Source packages and AutoFix working files | Deleted when the job ends. Anything left is deleted after 1 day. | +| Scan reports, SBOMs and other scan result files | 30 days from creation. | +| Findings | While open. A closed finding is removed and a `cloud_finding.closed` event is emitted. | +| Superseded provenance events, and finished remediation runs with their steps and events | 90 days by default. Runs that are still active or `verified` are kept. | +| Runtime package evidence | 15 minutes. | +| Pull-request check bookkeeping | 24 hours. | +| Temporary detections and isolation created by a remediation | Until their expiry: at most 7 days, isolation at most 4 hours. | + +## When your organization leaves + +When an organization is deleted or unsubscribes, Code Security data is removed +from the live databases after a 7-day grace period, and remaining backup copies +expire within 7 days after that. Scan result files expire within 30 days of their +creation. + +Some things are yours and stay with you: tickets, webhook deliveries and pull +requests already created in your systems. An unsubscribed organization also +keeps its own configuration, such as policies and response settings, until the +organization is deleted. + +To ask for your organization's Code Security data to be removed, or to confirm +that removal finished, contact LimaCharlie support. + +## Isolation between organizations + +Every record that holds your data is keyed to your organization. Every API route checks that +the caller belongs to the organization in the path, and the server stamps the +organization itself. Download links for stored files are signed for a single +object and expire, after 15 minutes for SBOMs. Scan jobs run on a network with +no route to the rest of the platform. diff --git a/docs/cloud-security/code-security/guarantees.md b/docs/cloud-security/code-security/guarantees.md new file mode 100644 index 000000000..9389537d6 --- /dev/null +++ b/docs/cloud-security/code-security/guarantees.md @@ -0,0 +1,93 @@ +# What Code Security guarantees + +Code Security connects a code change to what runs in your cloud, lets you act +on it through an approved LimaCharlie workflow, and checks afterwards whether +the risk is gone. When the evidence is incomplete, it says so, as `unknown`, +`partial` or `present`, and gives the reason. It does not fill a gap with a +guess. + +These are the rules it follows. Each one is enforced in the product, and the +reason codes it returns when a rule is not met are listed in +[Unknown, partial and refusal reasons](reasons.md). + +## The rules + +**A fix is `verified` only when it runs everywhere in scope.** Every in-scope +deployment must have fresh, complete evidence. Every digest it runs must carry +an uncontested build record, asserted or verified, that names the fix commit. +Nothing in scope may still run the vulnerable digest. For an image fix, each running fixed digest also needs a completed scan +that does not report the issue. The scan must come from a scanner that reported +the original vulnerability, using vulnerability data at least as fresh as the +finding. An image that was never scanned is never treated as clean. For a +dependency fix in a repository, detection must also close the finding. A merged +repository fix with no deployment in scope ends as unverifiable, not +`verified`. A merged pull +request is progress, not a fix. If the vulnerable digest runs again within 30 +days of verification, the run becomes `regressed`. You do not have to delete the +old image from your registry. While it exists, its own finding stays open. + +**Every image-to-source link says how it was established.** `inferred` means +the image's build steps and files match a repository you connected. It names the +repository, never the exact build commit. `asserted` means a label or build +record names the source without a trusted signature. `verified` means a +signature was checked for that exact digest. The signature can come from a +Google Cloud Build or GitHub Actions artifact attestation, or from a statement +you pushed that matches a signing identity you trust in your `provenance_trust` +policy. Conflicting evidence is `ambiguous`, never resolved by picking one. Public vendor images that match +none of your connections are reported separately as third-party and do not +count against your coverage. Inferred and label-based links need no change to +your build pipeline. A `verified` link needs a signed attestation, which your +build has to produce. + +**"Not observed" needs a complete telemetry window.** A runtime check says +`not_observed` only when every sensor on the resource reported a complete window. +Missing, late or partial telemetry never gives `not_observed`. The answer is +`present` or `unknown`, with a reason. A package already seen loaded or running +stays `loaded` or `executing`. Runtime evidence never changes a finding's risk +score. + +**Every change to your systems has a named approver and a confirmed outcome.** +Fix pull requests (AutoFix included), notifications, tickets, temporary +detections and endpoint isolation all run as remediation runs. A person or API key +with `cloudsec.respond` approves each run, the server derives its target +from the finding, and the outcome comes back through an authenticated callback +that the run records. Temporary controls expire on their own, after at most 7 +days, or 4 hours for isolation. If a control cannot be removed on time, a HIGH +finding opens. LimaCharlie never merges, deploys, rolls back or revokes +credentials for you. + +**Secrets in Terraform state and plans are never stored.** LimaCharlie refuses +state and plan files. An extractor you run yourself turns them into a map of +resource identities and allowlisted settings, and drops sensitive values before +anything is uploaded. The API refuses maps with secret-looking keys. Secrets +found in code are kept as a salted hash, never the value. + +**Code-to-cloud matches are exact or not made.** A finding is attributed to a +cloud resource only on an exact identifier. A declaration that matches several +resources is reported as `ambiguous`. One that matches none is `none`, with a +reason. + +**Pull-request summaries don't expose your estate.** In `risk_summary` mode a +pull-request check shows counts and yes/no facts, with no resource names, +account IDs, IP addresses or sensor IDs. When the live lookup fails, the check +publishes its normal scan verdict unchanged. + +**Your data stays separate, and leaves when you do.** Every record that holds +your data is keyed to your organization, and every API route checks the caller +against it. +When an organization is deleted or unsubscribes, Code Security data is removed +from the live databases after a 7-day grace period, and remaining backup copies +expire within 7 days after that. Scan result files expire within 30 days of +their creation. See [Data handling and privacy](data-handling.md). + +## What it does not claim + +- It does not prevent a vulnerable change from shipping. Pull-request checks + block a merge only if you configure them to. +- A finding without a runtime observation is not "not exploitable", and a + resource without an exposure fact is not "not exposed". +- It does not contain an incident automatically. Every action waits for a + person. +- Coverage is reported per organization with its denominator. Where the + denominator is unknown, no percentage is shown. +- Only Cloud Run and GKE deployments are traced to an image digest today. diff --git a/docs/cloud-security/code-security/incident-response.md b/docs/cloud-security/code-security/incident-response.md new file mode 100644 index 000000000..2006866a1 --- /dev/null +++ b/docs/cloud-security/code-security/incident-response.md @@ -0,0 +1,93 @@ +# Automatic behavior and incident response + +This page describes what Code Security does on its own when something goes +wrong, what it never does on its own, and what you can do during an incident +that involves a remediation. + +## What happens automatically + +| Situation | What Code Security does | +|---|---| +| The security graph is slow or down during a pull-request check | The check publishes its normal scan verdict, unchanged, with a note that the live context is missing. Impact reads give `graph_unavailable` or `deadline`. | +| A sensor stops reporting | A lapse never turns into `not_observed`. The answer becomes `present` or `unknown`, with a reason. A package already seen loaded or running stays so. | +| Deployment or scan evidence goes stale | The affected stage or coverage line reads `stale` or `coverage_stale`. Verification waits. | +| A fix is merged but not deployed everywhere | The run stays in `monitoring` with `old_digest_running` or `deployment_partial`. It is not `verified`. | +| The fixed image has no completed scan | The run stays in `monitoring` with `fix_digest_unscanned`. | +| A fix run's deadline passes without proof | If the old digest is still running, the run ends `persists`. Otherwise it ends `expired` with `deadline`. The finding is not marked fixed. | +| A repository fix is merged but nothing in scope runs it | The run ends `expired` with `pr_merged_unverifiable`. It is never marked `verified`. | +| The vulnerable digest runs again after verification | Within 30 days of verification, the run becomes `regressed` and the evidence chain says so. | +| A temporary control reaches its expiry | It is removed. Detection rules also carry their own expiry, so they are deleted even if Code Security is unavailable. Isolation expires at most 4 hours after it is applied, and cleanup then releases it. | +| A temporary control cannot be removed on time | After three failed removals, or 15 minutes past expiry, a HIGH `code-response-cleanup-failed` finding opens. It closes once the control is confirmed gone. | +| An executor never reports back | The run retries, then fails with `dispatch_exhausted` or `deadline`. A late or repeated callback has no second effect. | +| The playbook installation or the target changes after approval | The run stops before acting, with `installation_changed` or `target_changed`. | +| Someone else isolated the sensor first | The run does not act (`target_already_isolated`), so there is no isolation of its own to release. | + +## What never happens automatically + +- Nothing is merged, deployed or rolled back. +- No credential is revoked. +- No run acts without approval from a person or API key with `cloudsec.respond`. + There is no pre-approval. +- No target is taken from a request. The server derives it from the finding. +- Runtime and exposure evidence never lower a finding's risk score. + +## During an incident + +### A remediation is doing something you did not expect + +1. Find the run: `limacharlie cloudsec remediation list --finding-id `, + or the finding's evidence chain in the console. +2. Cancel it if it has not finished: + `limacharlie cloudsec remediation cancel `, then repeat with the + `--confirm` token it prints. Cancel works even when remediation is disabled + for your organization. +3. A temporary detection it installed expires on its own. To remove it now, + delete its D&R rule, named `cloudsec-rem-*`. +4. An isolation it applied expires after at most 4 hours and cleanup releases + it. If release fails, cleanup keeps retrying and a HIGH finding opens. To + release it earlier, rejoin the sensor to the network as you would any + isolated sensor. +5. A pull request it opened stays open for you to close. The run keeps + monitoring until its deadline and then ends `expired`. + +### Stop all Code Security actions in your organization + +- Set every playbook in your `response` policy to `mode: dry_run`, or delete + the policy. Notify, ticket, detection and isolation runs then record their + plan and act on nothing. A run approved under the old installation stops + with `installation_changed` before it acts. +- Remove `cloudsec.respond` from the users and API keys that hold it. Nobody can + then request or approve a run, or press **Open AutoFix PR**. +- Remove the GitHub App's write permissions to stop fix pull requests at the + source. + +Runs that already acted keep their cleanup: temporary controls are still +removed at expiry. + +### A HIGH `code-response-cleanup-failed` finding opened + +The finding names the control. Check whether its effect is still in place: the +`cloudsec-rem-*` detection rule, or the sensor's isolation state. Remove it by +hand if it is. The finding closes on its own once the control is confirmed gone. +If it does not, contact LimaCharlie support with the finding ID. + +### A fix shows `regressed` + +The vulnerable digest is running again. The run's verification step lists which +workloads run which digests. Find the deployment that reintroduced it, +typically a rollback or a pipeline that still builds from the old commit, and +redeploy the fixed image. + +### Results look wrong + +Check the reason codes first ([Unknown, partial and refusal +reasons](reasons.md)). A stale or partial answer is expected while collection +catches up. If a finding is attributed to the wrong resource, or a stage is +`proven` when you believe it should not be, contact LimaCharlie support with +the finding ID and the evidence chain. + +## Getting help + +Contact LimaCharlie support with the organization ID, the finding or run ID, +and the reason codes shown. Support can check the state of a run, a purge, or a +feature's rollout in your region. diff --git a/docs/cloud-security/code-security/index.md b/docs/cloud-security/code-security/index.md index 6ca9ef692..675f7ca78 100644 --- a/docs/cloud-security/code-security/index.md +++ b/docs/cloud-security/code-security/index.md @@ -50,7 +50,9 @@ keep it. jobs use the connection's token. That is why those connections ask for read-only scopes. - **Only the results leave the sandbox:** findings, the software bill of - materials and hashes. File contents, diffs and secret values are never stored. + materials and hashes. Findings never hold file contents, diffs or secret + values. The commit package the scan reads, and the files AutoFix edits, are + deleted when the job ends. See [Data handling and privacy](data-handling.md). - **Secrets are stored as a salted hash.** No field on a finding can hold the credential itself. - **Nothing is written to your repositories unless you allow it.** Pull-request @@ -85,5 +87,10 @@ source** filters narrow the list. - [Pull-request checks and push rescans](pull-requests.md): scan every push and gate merges on GitHub. - [AutoFix pull requests](autofix.md): let LimaCharlie open dependency upgrade pull requests. - [Bring your own scanner](bring-your-own-scanner.md): scan in your own CI, or push SARIF and CycloneDX results. +- [What Code Security guarantees](guarantees.md): the rules behind `verified`, image lineage, runtime evidence and remediation. +- [Configure evidence, lineage and remediation](containment-setup.md): permissions, connections, pull-request disclosure, remediation runs and playbooks. +- [Automatic behavior and incident response](incident-response.md): what happens on its own, and how to stop or undo a remediation. +- [Data handling and privacy](data-handling.md): what is read, what is kept, retention and purge. - [Reference](reference.md): languages, limits, status codes and API routes. - [Troubleshooting](troubleshooting.md): common problems and how to fix them. +- [Unknown, partial and refusal reasons](reasons.md): every reason code and what to do about it. diff --git a/docs/cloud-security/code-security/reasons.md b/docs/cloud-security/code-security/reasons.md new file mode 100644 index 000000000..6a86ecd15 --- /dev/null +++ b/docs/cloud-security/code-security/reasons.md @@ -0,0 +1,401 @@ +# Unknown, partial and refusal reasons + +Code Security never turns missing evidence into a reassuring answer. When it +cannot prove something, it says `unknown` or `partial` and returns a closed +reason code. Evidence-chain stages and coverage lines also give an `action` code +that names the next step. This page lists the codes for the generally available +workflows and what to do about each. + +For scan-lane problems (repositories not scanned, webhooks, the GitHub App), see +[Troubleshooting](troubleshooting.md). For what the guarantees behind these codes +are, see [What Code Security guarantees](guarantees.md). + +## Reading a reason + +Evidence-chain stages and coverage lines carry these fields. Runtime checks +return `status`, `reason`, `level`, `observed_at` and `stale_at`. Remediation +runs carry `state`, `failure`, `failure_reason` and `playbook_reason` instead. +The action for those codes is listed in each table below. + +| Field | Meaning | +|---|---| +| `status` | `proven`, `partial`, `unknown` or `not_applicable`. `proven` says the stage is evidenced, not that the news is good. Read `outcome` for that. | +| `level` | How strong the evidence is: `verified` (cryptographically checked), `asserted` (a claim from you or a tool), `observed` (LimaCharlie saw it), `derived` (a join of the above), or `unknown`. | +| `reason_domain`, `reason` | The feature the reason belongs to, and the closed reason code. | +| `reason_recognised` | `false` when the code is not in the server's catalog. The code is shown as is, with the action `review_reason`. Report it to LimaCharlie support. | +| `action` | The suggested next step, from the table below. | +| `observed_at`, `stale_at` | When the evidence was observed, and when it stops counting. | + +## Actions + +| `action` | What to do | +|---|---| +| `connect_repository` | Connect the repository in Code Security so its code can be scanned and linked. | +| `rescan_repository` | Rescan the repository so its latest commit is recorded. | +| `push_iac_map` | Push an infrastructure-as-code map so declarations can be matched to live resources. See [Terraform maps](containment-setup.md#terraform-maps). | +| `narrow_iac_scope` | Narrow the infrastructure-as-code scope so each declaration matches one live resource. | +| `push_build_provenance` | Push build provenance so each image digest names the commit it was built from. See [Build provenance](containment-setup.md#build-provenance). | +| `resolve_provenance_conflict` | Two builds claim different sources for this artifact. Check your build records and push the correct one. | +| `deploy_by_digest` | Deploy by immutable image digest instead of a tag. | +| `check_provider_access` | Check that the cloud connection can read this resource's deployments. | +| `connect_cloud_provider` | Connect the cloud account that holds this resource. | +| `wait_for_collection` | Wait for the next collection pass, then reload. | +| `open_impact_view` | Open the live impact view for this commit to see the full answer. | +| `retry_later` | A backend read failed or ran out of time. Try again in a few minutes. | +| `run_runtime_check` | Run a runtime check. | +| `deploy_sensor` | Deploy a LimaCharlie sensor on this resource. | +| `check_sensor_health` | The sensor's telemetry was interrupted or incomplete. Check the sensor. | +| `request_remediation` | Request a remediation run. | +| `approve_remediation` | A run waits for someone with `cloudsec.respond` to approve it. | +| `wait_for_remediation` | A run is in progress. Wait for it to report back. | +| `wait_for_rollout` | Wait for the fix to reach every in-scope deployment. | +| `investigate_regression` | The old artifact came back after the fix was verified. Find the deployment that reintroduced it. | +| `configure_write_access` | Configure write access for the repository. | +| `review_finding` | The evidence does not fit together. A person needs to decide. | +| `review_manually` | LimaCharlie has no evidence for this. Check the cloud provider or the repository directly. | +| `review_reason` | The reason is not one this version knows. Read the raw code. | +| `enable_feature` | The capability is not enabled for your organization. Ask your administrator or LimaCharlie support. | +| `contact_support` | Contact LimaCharlie support with the finding ID. | + +## Evidence chain + +`GET /findings/{finding_id}/evidence-chain` returns eight stages: `declared`, +`committed`, `built`, `running`, `exposed`, `observed`, `responded`, `verified`. +If `chain` is null, the top-level `reason` is `feature_disabled` or +`finding_not_found`. + +| `reason` | Action | Meaning | +|---|---|---| +| `iac_origin_partial` | `push_iac_map` | A declaration inventory was capped or incomplete, so these may not be all the declarations. | +| `no_iac_origin` | `push_iac_map` | No declaration is attributed. That does not prove none exists. | +| `iac_origin_unverified` | `rescan_repository` | The listed declarations carry no fresh, complete evidence. | +| `commit_unknown` | `rescan_repository` | The scan recorded no exact commit. | +| `not_a_workload` | none | The resource is not a workload, so build and deployment stages do not apply. | +| `runtime_not_applicable` | none | The finding names no package, so runtime evidence does not apply. | +| `workload_not_resolved` | `check_provider_access` | No deployment was resolved for this workload. | +| `deployment_not_observed` | `review_manually` | No container deployment was observed. Only Cloud Run and GKE digests are observed today. | +| `deployment_trace_unavailable` | `review_manually` | Which deployments run code from this repository is not traced per finding in this version. | +| `workloads_truncated` | `review_finding` | More workloads run this than one chain lists. | +| `observation_time_unknown` | `wait_for_collection` | The fact has no observation time, so its freshness cannot be stated. | +| `chain_read_failed` | `retry_later` | A read the chain depends on failed. | +| `evidence_contradictory` | `review_finding` | The evidence contradicts itself, so neither side is shown as the answer. | +| `lineage_candidate` | `push_build_provenance` | A source repository is linked, but the exact build commit is not proven. | +| `lineage_ambiguous` | `resolve_provenance_conflict` | Several sources match. | +| `lineage_stale` | `retry_later` | The image lineage expired. It is refreshed on the next image scan. | +| `exposure_not_established` | `review_manually` | Nothing establishes exposure. That does not prove the resource is unexposed. | +| `finding_not_open` | `review_finding` | The finding is not open, so its exposure facts are not current. | +| `runtime_not_checked` | `run_runtime_check` | No runtime check was part of this read. Add `runtime=true` or run a check. | +| `no_remediation_requested` | `request_remediation` | No remediation run exists. | +| `awaiting_approval` | `approve_remediation` | A run waits for approval. | +| `remediation_in_progress` | `wait_for_remediation` | A run is in progress. | +| `remediation_rejected`, `remediation_cancelled`, `remediation_expired` | `request_remediation` | The latest run was rejected, cancelled, or expired without a conclusive result. | +| `remediation_failed` | `contact_support` | The latest run failed. | +| `verification_pending` | `wait_for_rollout` | The rollout is being monitored. No verdict yet. | +| `remediation_persists` | `wait_for_rollout` | The old artifact is still running after the monitoring window. | +| `remediation_regressed` | `investigate_regression` | The old artifact came back after verification. | +| `remediation_state_unrecognised` | `review_reason` | The run's state is not one this server version knows. | + +Stages that are proven carry a positive reason instead: `code_location`, +`iac_origin`, `exposure_established`, `response_executed`, +`remediation_verified`. + +### Code-to-cloud attribution + +A finding's `iac_attribution` is `attributed`, `ambiguous` or `none`. When it is +`none`, the reason says why: + +| `reason` | Action | Meaning | +|---|---|---| +| `unresolved_name` | `push_iac_map` | The declaration's name could not be resolved, so no live resource was matched. A map with the resolved identity fixes this. | +| `unsupported_resource_type` | `review_manually` | This resource type is not supported for attribution. | +| `not_collected` | `connect_cloud_provider` | This resource type is not collected by name, so no live match could be checked. | +| `no_match` | `review_manually` | No matching live resource exists. | + +### Build provenance + +| `reason` | Action | Meaning | +|---|---|---| +| `missing` | `push_build_provenance` | No build provenance is recorded for this artifact. | +| `conflict` | `resolve_provenance_conflict` | Build provenance records disagree about the source. Neither is used. | +| `incomplete` | `push_build_provenance` | The build provenance is incomplete. | +| `declared` | `push_build_provenance` | The source is declared, for example by an image label, but not proven by build provenance. | + +## Deployment coverage and image lineage + +`GET /code/coverage` returns one line per metric, each with a `numerator`, a +`denominator` and a `breakdown`. A percentage is shown only when the line is +complete and fresh. If `coverage` is null, the reason is `feature_disabled`. + +| Line `reason` | Action | Meaning | +|---|---|---| +| `coverage_incomplete` | `check_provider_access` | A collection pass did not read its whole scope, so the denominator is a lower bound. | +| `coverage_stale` | `wait_for_collection` | Part of the count is past its evidence window. | +| `coverage_truncated` | `review_manually` | More rows existed than one report reads. | +| `no_denominator` | `connect_cloud_provider` | There is nothing to count yet. | +| `coverage_read_failed` | `retry_later` | The read for this line failed. | +| `metric_not_materialized` | `review_manually` | This metric is not counted in this version. | + +### Why a workload has no digest + +These codes appear in the workload coverage breakdown and on the `running` +stage of the evidence chain. + +| Code | Action | Meaning | +|---|---|---| +| `tag_only` | `deploy_by_digest` | The deployment names an image tag, not an immutable digest. | +| `revision_unavailable` | `check_provider_access` | The deployment's current revision could not be read. | +| `provider_unreachable` | `check_provider_access` | The cloud provider could not be reached. | +| `not_running` | `wait_for_collection` | Nothing runs for this deployment right now (scaled to zero, or no pods). Reported beside the percentage, not inside it. | +| `malformed_digest` | `contact_support` | The provider reported a malformed digest. | +| `stale` | `wait_for_collection` | The deployment evidence is past its window. | +| `partial` | `check_provider_access` | Only part of the deployment could be resolved. | +| `missing` | `check_provider_access` | No deployment evidence was found. | +| `unattributed` | `push_build_provenance` | The running artifact has no recorded source. | +| `build_unstated` | `push_build_provenance` | The build that produced the artifact is not recorded. | +| `deployment_unknown` | `check_provider_access` | What is deployed could not be determined. | + +`provider_digest`, `resolved` and `rolling` (a rollout in progress, more than +one artifact running) are resolved states. + +### Image lineage breakdown + +The `digests_with_source` line counts your own running image digests that have a +source link. Its breakdown keys: + +| Key | Counted as | +|---|---| +| `inferred` | Covered. Matched from the image's build steps and files. | +| `tool_emitted_asserted`, `signed_push_asserted` | Covered. A label or pushed statement without a trusted signature. | +| `tool_emitted_verified`, `signed_push_verified` | Covered. A trusted signature checked for that digest. | +| `ambiguous` | Not covered. Several sources match. | +| `unknown` | Not covered. No usable evidence, or the evidence is stale. | +| `third_party` | Outside the percentage. A public image that matches none of your registries or repositories. | +| `ownership_unknown` | Withholds the percentage until your registry or source connections show whose images these are. | + +On one image (`GET /code/images/{digest}`), `lineage.reason` explains the +decision. The common ones: + +| `reason` | Meaning | +|---|---| +| `insufficient_evidence`, `lineage_evidence_unavailable` | No candidate matched well enough. | +| `partial_image_evidence` | The image's own metadata is incomplete. | +| `unique_fingerprint_match` | One repository matches (inferred). | +| `competing_candidates` | More than one repository matches (ambiguous). | +| `oci_source_revision_label` | Asserted from the image's OCI source and revision labels. | +| `source_label_unresolved` | The image has no usable source or revision label. | +| `source_label_fingerprint_conflict` | The label and the build-step match name different repositories (ambiguous). | +| `google_cloud_build_signature`, `github_actions_signature` | Verified from a Google Cloud Build or GitHub Actions signature. | +| `native_lineage_conflict`, `signed_claim_conflict`, `producer_claim_conflict`, `lineage_source_conflict` | Two sources of lineage disagree. Neither is used. | +| `signed_push_verified`, `signed_push_asserted` | From a statement you pushed, with or without a trusted signature. | +| `signed_claim_stale`, `signed_claim_incomplete` | A signed claim is out of date or incomplete. | +| `base_candidate_cap`, `build_candidate_cap`, `producer_row_cap`, `signed_claim_bounds` | Too many candidates to decide within limits. | +| `lineage_read_failed` | The read failed. Try again. | + +## Live impact and pull-request consequence + +`GET /code/impact` and the pull-request consequence section return a +`summary.status` of `complete`, `partial` or `unavailable`. A partial answer +never says "no impact". Each impact lists the reasons it is partial: + +| `reason` | Action | Meaning | +|---|---|---| +| `graph_unavailable` | `retry_later` | The security graph did not answer. The pull-request verdict is unaffected. | +| `deadline` | `retry_later` | The impact read ran past its 2-second budget. The pull-request verdict is unaffected. | +| `declarations_truncated` | `open_impact_view` | More than 100 declarations changed, so not all were considered. | +| `graph_truncated` | `open_impact_view` | The graph answered only in part, within its 500-row limit. | +| `mapping_stale` | `push_iac_map` | The code-to-cloud map describes a different revision. | +| `mapping_ambiguous` | `narrow_iac_scope` | A declaration matches more than one live resource. | +| `resource_not_collected` | `connect_cloud_provider` | A matched resource is not collected. | +| `deployment_unknown` | `check_provider_access` | What runs on a workload could not be determined. | +| `runtime_unavailable` | `run_runtime_check` | Runtime information was not available. | +| `disclosure_redacted` | `open_impact_view` | Detail was withheld by the pull-request disclosure setting. | +| `facet_not_collected` | `review_manually` | This fact is not recorded for this kind of resource. | + +If `impact` is null, the reason is `feature_disabled` or `subject_not_found` +(the repository or commit was not found). + +`exposure`, `privilege` and `sensitivity` are `established`, +`not_established` or `unknown`. `not_established` means nothing positive was +found. It never means "not exposed". An impact with status `no_live_match` names +a resource that does not exist yet, typically one the change will create. + +## Runtime checks + +`POST /findings/{finding_id}/runtime-check` returns a `status`: + +| `status` | Meaning | +|---|---| +| `executing` | The package is the running executable. | +| `loaded` | The package is loaded into a running process. | +| `not_observed` | A complete telemetry window never saw the package loaded. This is not "absent" and not "not exploitable". | +| `present` | A sensor is on the resource, but no package-level claim is possible. | +| `unknown` | No usable runtime evidence. | + +When the check could not run at all, `accepted` is `false`: + +| `reason` | Action | Meaning | +|---|---|---| +| `feature_disabled` | `enable_feature` | Runtime evidence is not enabled for your organization. | +| `no_resource` | `review_finding` | The finding names no resource a sensor could run on. | +| `no_packages` | `review_finding` | The finding names no package to look for. | +| `cache_unavailable` | `retry_later` | The runtime evidence store did not answer. | +| `no_sensors` | `deploy_sensor` | No LimaCharlie sensor runs on this resource. | +| `sensors_partial` | `review_manually` | Not every sensor on the resource reported. | + +Reasons on a verdict: + +| `reason` | Action | Meaning | +|---|---|---| +| `no_evidence` | `wait_for_collection` | No runtime evidence has been recorded yet. | +| `expired` | `run_runtime_check` | The evidence is past its 30-minute window. | +| `not_relevant` | `run_runtime_check` | The package was not watched when the evidence was recorded. | +| `window_short` | `wait_for_collection` | The telemetry window is too short to support a claim. | +| `window_interrupted` | `check_sensor_health` | The telemetry window was interrupted. | +| `telemetry_dropped` | `check_sensor_health` | The sensor dropped telemetry during the window. | +| `telemetry_absent` | `check_sensor_health` | The sensor sent no process or module telemetry. | +| `stale_confirmation` | `check_sensor_health` | The sensor's confirmation is out of date. | +| `write_shed` | `retry_later` | Some evidence was dropped under load. | +| `unattributable` | `review_manually` | The activity cannot be attributed to this package. | +| `attribution_incomplete` | `review_manually` | The package's file paths could not all be identified. | +| `relevance_truncated` | `review_manually` | Too many packages were watched to cover them all. | +| `unversioned` | `review_manually` | The package version is not known. | +| `inventory_conflict` | `review_finding` | The sensor's inventory disagrees with the finding. | + +`observed_executing`, `observed_loaded` and `complete_window` are the positive +reasons. + +## Remediation runs + +### Errors from the remediation routes + +The `error` field on `/findings/{id}/remediations`, `/remediations/...` and +`/code/autofix`: + +| HTTP | `error` | Meaning and fix | +|---|---|---| +| 400 | `invalid_request` | A malformed request: bad ID, a body over 4 KiB, an unknown field, or a bad idempotency key or generation. | +| 403 | `missing_permission` | You lack `cloudsec.respond`. `cloudsec.set` does not include it. See [Permissions](containment-setup.md#permissions). | +| 404 | `not_found`, `finding_not_found` | No such run or finding in this organization. | +| 409 | `idempotency_mismatch` | The same idempotency key was used for a different request. | +| 409 | `generation_conflict` | The run changed since you read it. Reload and decide again. | +| 409 | `illegal_transition` | That decision is not possible from the run's current state. | +| 409 | `target_changed` | The target changed after you reviewed it. Reload and review again. | +| 410 | `approval_expired` | The approval window closed. Request a new run. | +| 422 | `action_unavailable` | The action is not enabled, the playbook is not installed, or the finding is not one this action can fix. | +| 429 | `capacity` | Your organization has 100 active runs, or the finding has 10. Wait for runs to finish. | +| 503 | `disabled` | Remediation is not enabled for your organization. `cancel` still works. | +| 502, 503 | `unavailable` | A backend failure. Try again. | + +### Why a run failed + +A run in state `failed` or `expired` carries `failure`: + +| `failure` | Action | Meaning | +|---|---|---| +| `action_unavailable` | `enable_feature` | The action is not available. | +| `policy_refused` | `review_finding` | The response policy refused the run. | +| `executor_error`, `callback_failed` | `contact_support` | The executor reported an error. See `playbook_reason` or `failure_reason`. | +| `dispatch_exhausted` | `retry_later` | The executor could not be reached. | +| `deadline` | `request_remediation` | The run's deadline passed. | +| `window_ended` | `request_remediation` | The run's window ended. The finding is not verified as fixed. | +| `pr_closed` | `request_remediation` | The fix pull request was closed without merging. | +| `pr_merged_unverifiable` | `review_finding` | The pull request merged, but the fix could not be verified. | +| `invalid_run` | `contact_support` | The run was invalid. | + +### Fix pull requests and AutoFix + +`failure_reason` on an `open_fix_pr` run (AutoFix button presses included): + +| `failure_reason` | Action | Meaning | +|---|---|---| +| `repository_not_connected` | `connect_repository` | The repository is not connected, or no enabled policy selects it. | +| `repository_finding_not_found` | `rescan_repository` | The matching finding is not in the repository. | +| `repository_finding_ambiguous` | `review_finding` | More than one repository finding matches. | +| `provenance_unknown` | `push_build_provenance` | The image's source commit is not known, so no fix pull request can be opened. | +| `write_app_not_configured` | `configure_write_access` | No write access is configured for the repository. | +| `write_app_lacks_contents` | `configure_write_access` | Write access lacks permission to change contents. | +| `finding_not_autofixable`, `autofix_not_applicable` | `review_manually` | No automatic fix applies. See [AutoFix](autofix.md#when-no-pull-request-appears). | +| `autofix_pr_already_open` | `review_manually` | An AutoFix pull request for this package is already open. | +| `autofix_budget_exhausted` | `retry_later` | The daily AutoFix limit of 20 per connection is used up. | +| `autofix_budget_unavailable`, `executor_unavailable` | `retry_later` | The service could not be reached. | +| `autofix_job_failed` | `contact_support` | The fix job failed. | +| `autofix_pr_failed` | `configure_write_access` | The pull request could not be opened. | + +### Playbook actions + +`playbook_reason` on a notify, ticket, temporary-detection or isolation run. +The catalog also holds `consent_expired` and `claim_stale`, which apply only to +validation templates that are not generally available. + +| `playbook_reason` | Action | Meaning | +|---|---|---| +| `installation_missing` | `enable_feature` | The response playbook is no longer installed. | +| `installation_changed` | `request_remediation` | The installation changed after the run was approved. | +| `target_changed` | `review_finding` | The finding no longer points at the approved target. | +| `approval_stale` | `request_remediation` | The approval was too old when the playbook was about to act. | +| `target_already_isolated` | `review_manually` | The sensor was already isolated by someone else. LimaCharlie left it alone. | +| `effector_unavailable`, `executor_unavailable` | `contact_support` | A service the playbook needed refused or failed. | +| `effect_unconfirmed` | `review_manually` | The change was applied, but reading it back did not confirm it. | +| `control_ended` | `request_remediation` | The temporary control had already ended, or too little of its window was left. | + +### Verification + +While a run is `monitoring`, each observation step lists why it is not yet +`verified`: + +| `reason` | Action | Meaning | +|---|---|---| +| `old_digest_running` | `wait_for_rollout` | A deployment in scope still runs the old artifact. | +| `deployment_partial` | `wait_for_rollout` | The rollout has reached only part of the scope. | +| `deployment_missing` | `check_provider_access` | A workload in scope has no deployment evidence. | +| `deployment_stale` | `wait_for_collection` | Deployment evidence is past its window. | +| `deployment_scope_empty` | `review_finding` | No deployments are in scope. | +| `scope_truncated` | `review_finding` | The scope is too large to verify. | +| `unexpected_digest` | `review_finding` | A workload runs an artifact that is neither the old nor the fixed one. | +| `finding_open` | `wait_for_collection` | Detection still reports the finding. | +| `finding_operator_closed` | `review_finding` | Someone closed the finding by hand. That does not count as a fix. | +| `finding_unknown` | `wait_for_collection` | The finding's state could not be read. | +| `source_finding_open` | `rescan_repository` | The finding is still open in the source repository. | +| `fix_digest_unproven` | `push_build_provenance` | No build record proves which image contains the fix. | +| `fix_builds_truncated` | `review_manually` | Too many builds to check which contain the fix. | +| `fix_digest_unscanned` | `wait_for_collection` | No completed scan of the fixed image yet. An image without a scan is never treated as clean. | +| `fix_digest_still_vulnerable` | `review_finding` | The image built from the fix is still vulnerable. | +| `fix_scan_source_unavailable` | `review_manually` | None of the scanners that found this vulnerability reports completed scans of the fixed image. Check it by hand. | +| `expectation_unproven` | `review_manually` | What the fix should look like in production cannot be checked automatically. | + +## Pushing maps and provenance + +`POST /code/iac-map`: + +| HTTP | `error` | Meaning and fix | +|---|---|---| +| 400 | `iac_map_raw_terraform` | You sent a raw state or plan. Run the extractor it names and push its output. See [Terraform maps](containment-setup.md#terraform-maps). | +| 400 | `iac_map_invalid` | Not a valid `lc-iac-map/v1` document, or it carries a secret-looking key. | +| 400 | `iac_map_bounds` | Over 20 MiB, 100,000 resources, nesting depth 8 or 4 KiB per string. Split it by workspace. | +| 400 | `iac_map_workspace_limit` | More than 100 workspaces for one repository. | +| 409 | `iac_map_stale` | A newer revision is already published. | +| 415 | | Send uncompressed `application/json`. | +| 429 | | More than 30 pushes a minute. | +| 503 | `codesec_disabled` | This feature is not enabled for your organization. | +| 503 | `iac_map_busy`, `iac_map_storage_unavailable`, `iac_map_ingest_unavailable` | Try again. Resubmitting the same document is safe. | + +`GET /code/iac-map/status` reports `processing`, `published`, `retryable` +(resubmit the same document) or `superseded`, and `iac_map_not_found` for an +unknown receipt. + +`POST /code/provenance` uses `error_code`: + +| HTTP | `error_code` | Meaning and fix | +|---|---|---| +| 400 | `provenance_invalid_document` | Bad JSON, over 1 MiB, or a reserved field such as `trust`, `verified` or `signer`. LimaCharlie sets those itself. | +| 400 | `provenance_identity_required` | The request has no authenticated identity. | +| 400 | `provenance_rejected` | The statement was refused. Details are withheld on purpose. Check it against the format. | +| 429 | | More than 60 pushes a minute. | +| 503 | `provenance_unavailable` | Try again. | + +## Feature switched off + +`feature_disabled` (evidence chain, coverage, impact, runtime check), +`disabled` (remediation) and `codesec_disabled` (map push) all mean the same +thing: the feature is not enabled for your organization. These capabilities are +enabled region by region. Contact LimaCharlie support to find out when yours is. diff --git a/docs/cloud-security/code-security/reference.md b/docs/cloud-security/code-security/reference.md index 69d65c3ae..759cd3355 100644 --- a/docs/cloud-security/code-security/reference.md +++ b/docs/cloud-security/code-security/reference.md @@ -150,8 +150,11 @@ on by default. See [Events](../api-reference.md#events). ## API routes All routes are under `https://api.limacharlie.io/v1/cloudsec/{oid}`. Reads need -`cloudsec.get`, writes need `cloudsec.set`, and the organization must be -subscribed to `ext-cloud-security`. +`cloudsec.get` and writes need `cloudsec.set`. The exceptions: AutoFix and +remediation requests and decisions need `cloudsec.respond`, a runtime check +(`POST`) needs only `cloudsec.get`, and the map status read +(`GET /code/iac-map/status`) needs `cloudsec.set`. See the table below. The organization must be subscribed to +`ext-cloud-security`. | Route | CLI | Purpose | |---|---|---| @@ -163,7 +166,7 @@ subscribed to `ext-cloud-security`. | `GET /code/images`, `GET /code/images/{digest}` | | Container images and one image's detail. | | `GET /code/image-repos`, `GET /code/image-repos/facets` | | Image repositories and their filter counts. | | `POST /code/scan` | `code rescan` | Rescan one repository. Body: `{repo, ref?, provider?}`. | -| `POST /code/autofix` | `code autofix` | Open an AutoFix pull request. Body: `{finding_id, repo?}`. | +| `POST /code/autofix` | `code autofix` | Open an AutoFix pull request as a remediation run. Needs `cloudsec.respond`. Body: `{finding_id, repo?}`. | | `POST /code/ingest` | `code ingest` | Push SARIF, CycloneDX or a scanner report. | | `POST /code/pr_check` | | Check a pull request. Used by the webhook rules. | | `POST /code/webhook` | | Point a GitHub App's webhook at LimaCharlie. See [the webhook API](pull-requests.md#the-webhook-api). | @@ -171,6 +174,23 @@ subscribed to `ext-cloud-security`. Findings are read with the standard [findings routes](../api-reference.md), filtered by `repo`. +Evidence, lineage and remediation routes. See +[Configure evidence, lineage and remediation](containment-setup.md) for the +setup, and [Unknown, partial and refusal reasons](reasons.md) for the codes +they return. + +| Route | Permission | Purpose | +|---|---|---| +| `GET /findings/{finding_id}/evidence-chain` | `cloudsec.get` | The finding's evidence chain. Optional `runtime=true`. | +| `POST /findings/{finding_id}/runtime-check` | `cloudsec.get` | Check whether the finding's package was seen running. Reads only. | +| `GET /code/coverage` | `cloudsec.get` | Coverage lines with their denominators. | +| `GET /code/impact` | `cloudsec.get` | Live impact of a commit (`repo_urn`, `commit`) or a finding (`finding_id`). | +| `GET /code/provenance`, `POST /code/provenance` | `cloudsec.get`, `cloudsec.set` | List or push build provenance. | +| `POST /code/iac-map`, `GET /code/iac-map/status` | `cloudsec.set` | Push a Terraform map, and read its publication status. | +| `POST /findings/{finding_id}/remediations` | `cloudsec.respond` | Request a remediation run. Body: `{action, idempotency_key}`. | +| `GET /remediations`, `GET /remediations/{run_id}` | `cloudsec.get` | List runs, or read one run with its steps. | +| `POST /remediations/{run_id}/approve`, `reject`, `cancel` | `cloudsec.respond` | Decide a run. | + ## Not available yet - **Scanning images from container registries.** `image_sources: ["registries"]` diff --git a/docs/cloud-security/code-security/troubleshooting.md b/docs/cloud-security/code-security/troubleshooting.md index 4fb4036eb..aa9fbe3c1 100644 --- a/docs/cloud-security/code-security/troubleshooting.md +++ b/docs/cloud-security/code-security/troubleshooting.md @@ -1,5 +1,8 @@ # Troubleshooting Code Security +For evidence-chain, lineage, runtime-check and remediation reason codes, see +[Unknown, partial and refusal reasons](reasons.md). + Start with the **Set up code security** checklist on the **Code security** page, if it is shown. It names what is not set up and links to the fix. From the CLI, `limacharlie cloudsec code status` shows whether scans run and what failed. @@ -59,6 +62,8 @@ if it is shown. It names what is not set up and links to the fix. From the CLI, |---|---| | No AutoFix pull request appears | Refusals are reported as `cloudsec.code_autofix_refused` events, once `ops_events` is on. See [When no pull request appears](autofix.md#when-no-pull-request-appears). | | `write_app_lacks_contents` | Grant **Contents: Read and write** to the App and approve on the installation page. | +| `403 missing_permission` when asking for a fix | You need `cloudsec.respond`. `cloudsec.set` does not include it. See [Permissions](containment-setup.md#permissions). | +| `503 disabled` when asking for a fix | Remediation is not enabled for your organization yet. | | The pull request warns that the lockfile is stale | Run the command in the pull request on its branch before merging. See [Lockfiles](autofix.md#lockfiles). | ## Pushed results and local scans diff --git a/docs/cloud-security/configuration.md b/docs/cloud-security/configuration.md index a840f1a65..fad27e1d1 100644 --- a/docs/cloud-security/configuration.md +++ b/docs/cloud-security/configuration.md @@ -16,6 +16,9 @@ tenant onboarding and fleet-wide policy a script, not a UI workflow (see `cloudsec_provider` records are gated by the dedicated `cloudsec_provider.get/set/del` permissions; `cloudsec_policy`, `cloudsec_query` and `cloudsec_code_rule` follow `cloudsec.get`/`cloudsec.set`. + Writing a `cloudsec_policy` record of type `response` also needs + `cloudsec.respond`, because it decides what a + [remediation run](code-security/containment-setup.md#response-playbooks) may do. ## cloudsec_provider @@ -134,10 +137,11 @@ not accept is **rejected when you save** rather than silently ignored: | Dimension | Where it is honored | |---|---| | `tag` | compute resources only. Within `classification` that means the `compute` section; the dimension is also accepted on `coverage` and `exclusions`. | -| `public` | data stores **and** compute. | +| `public` | data stores **and** compute. Not accepted on exclusions. | | `content_class` | data stores only — and not yet populated, see the caveat above. | | `label` / `label_key_present` | resources carrying cloud labels. Not honored on `classification.identities`. | | `services` / `resource_types` | collection exclusions only. | +| `region` | accepted on collection exclusions, where it is judged on each row. | | account / name / provider matchers | every surface, including the `exclusions` emission list — which honors *only* these. | The console's policy editors enforce this per surface and offer live value @@ -241,6 +245,7 @@ has no effect on the others. At least one list must be non-empty: | List | Effect | |---|---| | `collection` | Matching scopes are skipped by the collection sweep. Only this list may add the `services` and `resource_types` narrowers on top of the shared resource matchers — an account-only rule excludes the whole account, while adding `services`/`resource_types` narrows the exclusion to those collector services or resource types inside the matched scope. | +| `scanning` | Accepted for the agentless workload snapshot scanner. That scanner is not available, so this list has no effect today. | | `emission` | Matching events are dropped before delivery to the event stream. Only account/name/provider matchers are honored here — an emission rule constrained on labels or tags can never be satisfied by a lean event and so never drops one. | !!! warning "A collection exclusion deletes the inventory it excludes" @@ -252,6 +257,31 @@ has no effect on the others. At least one list must be non-empty: means excluding a scope you still want recorded loses that inventory until you remove the exclusion and let a sweep repopulate it. +How a `collection` rule is judged: + +- **One account per row.** A row is matched on its own account when it has one, + otherwise on the account the collector ran for. A tenant-wide collector (an + Azure tenant, for example) is excluded wholesale by a rule whose account or + provider matches it. A rule with a negated account pattern is applied row by + row instead. +- **`provider`, `region` and `resource_types` are checked on each row.** A + provider rule never matches another cloud's account. +- **A row that does not state a fact the rule needs is kept.** For example, a + `region` rule keeps rows with no region, and the sweep records a note that + those rows were kept. +- **Relationships go with what they connect.** Removing a resource also removes + its relationships, so a bucket-name rule removes that bucket's access grants. + Containers are the exception: a project-name rule does not remove grants made + at the project level. +- **A fully excluded type is removed even when the pass was partial or failed.** +- **Only deletes the exclusion caused are kept out of the event feed.** A + resource that is really deleted in the cloud still emits + `cloud_resource.deleted`. + +Excluded resources also drop out of [Code Security](code-security/index.md)'s +view of your cloud. Its evidence then reads `resource_not_collected` or +`workload_not_resolved`, never "not exposed". + Exclusions are captured when a sweep starts, so an edit lands on the *next* sweep. Change the provider's `sync_now` nonce to apply it immediately instead of waiting for the refresh cadence. diff --git a/mkdocs.yml b/mkdocs.yml index 32d21fee8..cfdfcb1dc 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -626,8 +626,13 @@ nav: - Pull-Request Checks & Push Rescans: cloud-security/code-security/pull-requests.md - AutoFix Pull Requests: cloud-security/code-security/autofix.md - Bring Your Own Scanner: cloud-security/code-security/bring-your-own-scanner.md + - What Code Security Guarantees: cloud-security/code-security/guarantees.md + - Evidence, Lineage & Remediation Setup: cloud-security/code-security/containment-setup.md + - Automatic Behavior & Incident Response: cloud-security/code-security/incident-response.md + - Data Handling & Privacy: cloud-security/code-security/data-handling.md - Reference: cloud-security/code-security/reference.md - Troubleshooting: cloud-security/code-security/troubleshooting.md + - Unknown, Partial & Refusal Reasons: cloud-security/code-security/reasons.md - Provider Setup: - Overview: cloud-security/provider-setup/index.md - Google Cloud: cloud-security/provider-setup/gcp.md