Skip to content

AppSec: synchronous reverse_dns postoverflow enrichment stalls in-band responses via unbuffered pipeline backpressure (bouncer fail-open bursts) #4600

Description

@roybj

Summary

With the crowdsecurity/rdns postoverflow installed, a slow PTR lookup for an alert's source IP freezes AppSec in-band inspection for the whole engine. Remediation-component requests then exceed APPSEC_PROCESS_TIMEOUT and fail open (APPSEC_FAILURE_ACTION=passthrough), i.e. requests are served uninspected - on our busiest host 76-1,413 timeouts/hour (~0.5% of inspections, bursts to 480/min).

Environment: crowdsec v1.7.8 (rpm), AlmaLinux 8.10, central LAPI, appsec acquisition on 127.0.0.1:7422 with crowdsec-openresty-bouncer v1.2.1 in-band, default routines: 1, ~600 alerts/h (WAF + log scenarios), OS resolver = ISP shared resolver, no local cache, default resolv.conf timeouts (5sx2).

Mechanism (all refs v1.7.8):

  1. crowdsecurity/rdns runs in the postoverflow parse inside runOutput -> parser.reverse_dns (pkg/parser/enrich_dns.go:19) -> net.LookupAddr - synchronous, unbounded by any context/timeout of its own. Hostile alert IPs frequently sit in lame/slow PTR zones -> 0.2-10 s lookups.
  2. While it blocks, the event train jams: overflow chan <- buckets <- parsers <- acquisition output (all unbuffered).
  3. The AppSec runner shares that output channel (runner.outChan = out, pkg/acquisition/modules/appsec/run.go:137) and blocks in handleRequest on r.outChan <- evt (appsec_runner.go:336/341/420/426). For in-band matches this send happens before the response is written to the remediation component (appsec_runner.go:472).
  4. InChan is also unbuffered (appsec/config.go:169), so every handler blocks; all in-flight inspections exceed the client's read budget -> fail-open burst. Burst size = request rate x PTR stall length.

Evidence

  • Goroutine dump captured while a probe to :7422 was hanging (437 ms) - runOutput blocked in the lookup:
goroutine 157 [IO wait]:
internal/poll.runtime_pollWait(...)
...
net.dnsPacketRoundTrip(...)
net.(*Resolver).exchange(...)
net.(*Resolver).tryOneName(...)
net.(*Resolver).goLookupPTR(...)
net.LookupAddr(...)
github.com/crowdsecurity/crowdsec/pkg/parser.reverse_dns(...)
	github.com/crowdsecurity/crowdsec/pkg/parser/enrich_dns.go:19
github.com/crowdsecurity/crowdsec/pkg/parser.(*RuntimeStatic).Apply(...)
	github.com/crowdsecurity/crowdsec/pkg/parser/runtime.go:167
[... postoverflow parse inside main.runOutput ...]
  • The inband parsing histogram is blind to it: 1.04 M inspections, mean 0.47 ms, only 9 ever >250 ms - while the same period logged tens of thousands of >1000 ms client read-timeouts. The stall sits between accept and the timed span.
  • Every clustered burst-second across a 90-min window (22/22) had an alert landing at LAPI +2...+12 s later - the stalled-lookup-then-flush signature.
  • Burst-adjacent alert source IPs re-resolved hours later (warm cache = lower bound): 10/26 still >200 ms, incl. 1,056 ms and a SERVFAIL at 487 ms.
  • Removal experiment: cscli postoverflows remove crowdsecurity/rdns --force + restart dropped the host from 76-1,413/h to 57 in the first full hour (lowest on record) and 15-74/h over the next nine hours, while an untouched sibling host (same stack, same LAPI) ran high (127 that hour). Probe watcher: 0 stalls >=350 ms in 6,038 probes post-change.
  • routines: 4 does not mitigate (measured null): all runners share the same jammed output channel.

Related: cmd/crowdsec/output.go already acknowledges this failure class - "Under high load, LAPI may take between 1 and 2 seconds to process ~100 alerts, which slows down everything including the WAF" - and made PushAlerts async. Parser-stage enrichment has the same coupling but no such fix. Symptom also matches crowdsecurity/cs-openresty-bouncer#75 (closed unresolved: high-traffic timeouts with 0.2 ms engine latency).

Suggested fixes (any one breaks the coupling): bound the enrichment (LookupAddr with a short-timeout context), make enrichment async/cached, buffer the overflow/acquisition channels, or give the AppSec runner a dedicated buffered emission queue so in-band responses never share fate with the alert pipeline.

Side note: in this deployment the reverse_dns value did not appear in any of 1,834 alert payloads at LAPI - if that's expected for appsec/scenario alerts, the enrichment cost bought nothing here.

Workaround for affected users: remove crowdsecurity/rdns (requires --force; owned by crowdsecurity/linux / whitelist-good-actors). seo-bots-whitelist degrades gracefully to its IP-list expression.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions