Add rabbitmq evalya full-coverage fixture - #25224
NouemanKHAL wants to merge 17 commits into
Conversation
Onboards rabbitmq to evalya with a rabbitmq-full task that drives the check to emit the metrics referenced by the OOTB dashboards and the recommended monitors. A single -management broker serves both check backends, since the assets draw metric names from each: the management API and the Prometheus plugin. Continuous AMQP traffic comes from perf-test, because rabbitmqadmin cannot populate channel, connection, or delivery counters over HTTP. Queue churn is separate so the node-wide queue counters keep advancing past perf-test's long-lived queues. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
✅ Dispatcher tests: passed
✅ 12 passed · nothing failed Batches · ✅ batch-01 12/12 Dispatcher finished on |
evalya-impact-summaryevalya impact analysis |
|
✅ All CI checks and tests passed. 🎉 All green!🧪 All tests passed 🎯 Code Coverage (details) 🔗 Commit SHA: a279494 | Docs | View more details | Give us feedback! |
Disk usage changeCommit No integration or dependency changed size. |
evalya starts only a task's target service and its depends_on chain. The task targeted the broker, which nothing depends on for the workload, so seed, load, and activity-gen never ran and every message, consumer, and queue metric read zero under evalya. The task now targets a socat entrypoint gated on the workload that forwards the broker's ports. The broker is renamed rabbitmq-broker because consumers alias this fixture as `rabbitmq`, and evalya gives the entrypoint that hostname, which made socat forward to itself. Host ports are no longer published, so the fixture cannot clash with a local broker or a concurrent run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
perf-test consumers had no prefetch limit, so the whole backlog sat unacked, messages_ready read zero, and unacked grew without bound. --qos 50 keeps the backlog ready and x-max-length=2000 caps it. Nothing published unroutable messages, so the unroutable-dropped counters read zero. A producer-only perf-test now publishes to an unbound routing key; it stays connected because the per-channel counter disappears with the channel. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fixed publish rate leaves every rate metric and the queue depth flat once the backlog hits x-max-length, so OTel vs DD comparisons have no shape to match. Alternating the publish rate every 60s between below and above the consumer rate makes them rise and fall. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The load connections live for the whole run, so the connection and channel opened/closed counters and the consumer count stop moving after startup. A looping short-lived perf-test keeps them advancing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Six dashboard/monitor metrics were emitted but stuck at 0, so they proved the mapping fired but not that the fixture exercised it: - delivered.count (channel, queue) only counts auto-ack deliveries, and every consumer acked manually. An autoack perf-test fixes it. - redeliver/redelivered.count and messages.persistent had no redelivery or persistent backlog. A nack/requeue perf-test on a durable queue with persistent messages drives all three. - messages_unacknowledged.rate stayed 0 because load's unacked count is pinned at its --qos. Capping conn-churn's consumer below its publish rate makes its unacked count climb each cycle. The new services idle with ACTIVITY_GEN=0, like activity-gen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README claimed 56/58 metrics emitted, but coverage should count only non-zero values: 11 of those 56 read 0. It now reports the non-zero figure (51/58), describes the autoack and redeliver services and the ACTIVITY_GEN scope, and names each metric left at 0 with why: sockets_used is zeroed by RabbitMQ 4.0, the alarms are intentionally not raised because an alarm blocks every publisher, and pending_packets needs a client that stops reading its socket. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Capping conn-churn's consumer at 5 msg/s made unacked a 30s sawtooth: it climbed while the 20s consumer lived, then the backlog was requeued at close. That cycle aliases against the scrape intervals. In a 600s comparator run every queue's average unacked matched across scrapers except churn-conn: OTel rabbitmqreceiver 17.8, DD OpenMetrics 35.2, DD management 29.4, so both unacked equivalences failed (6.06% and 6.39%) where they passed before. conn-churn's consumer is uncapped again, so that queue stays near 0. A new unacked-swing service drives messages_unacknowledged.rate from one long-lived connection instead: a constant 4 msg/s publisher and a consumer whose --variable-latency alternates every 120s between 2 and 8 msg/s of capacity, so unacked ramps to about 240 and back on a 240s cycle, slow enough for every scraper's average to agree. It idles with ACTIVITY_GEN=0, like the other added services. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still said conn-churn drives messages_unacknowledged.rate. It now explains why conn-churn's consumer stays uncapped (a backlog requeued every 30s aliases against the scrape intervals), describes unacked-swing and its slow 240s cycle, and adds it to the ACTIVITY_GEN scope and the entrypoint's gate list. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.node.mem_alarm, rabbitmq.node.disk_alarm and rabbitmq.alarms.free_disk_space.watermark are dashboard and monitor targets that read 0 on a healthy broker, so the rabbitmq-full fixture never covered them. The primary broker now runs a background drill that raises both alarms for 25s every 180s via rabbitmqctl, then restores the values it read at startup. Running inside the broker container reuses the node's cookie and name, which a sidecar would need shared. 25s is longer than the 15s DD and OTel scrape intervals plus the ~8s management stats lag, and shorter than perf-test's 30s confirm timeout (evalya ignores restart policies, so an exited perf-test would stay down). Publishers block during a window. ALARM_DRILL=0, or ACTIVITY_GEN=0, turns it off. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.connection.pending_packets counts bytes the broker has queued on a socket but not yet handed to the kernel. It is non-zero only when a client reads slower than the broker delivers, and every perf-test client keeps up, so it read 0 in every scrape. slow-reader is bespoke traffic: a stdlib-only AMQP 0-9-1 client (no pip install at runtime) whose auto-ack consumer reads at ~32 KiB/s through a 4 KiB receive buffer while a second connection publishes ~100 KiB/s. It throttles instead of stopping, because the broker closes a connection whose socket send blocks for 30s. Auto-ack leaves nothing unacked, x-max-length=1000 caps the queue, both rates are fixed, and the consumer reconnects every 600s. Idles with ACTIVITY_GEN=0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.node.sockets_used, rabbitmq.process.max_tcp_sockets and
rabbitmq.process.open_tcp_sockets are unreachable on RabbitMQ 4.x,
which stopped tracking TCP sockets. The primary broker stays on 4.0,
because the 4.x-only queue delivery counters are also targets and
seed.sh needs the v2 rabbitmqadmin.
A second broker on 3.13 (RABBITMQ3_VERSION), with two perf-test
connections, is forwarded by the entrypoint on ports shifted by one and
published as RABBITMQ3_HOST and RABBITMQ3_{AMQP,MANAGEMENT,OPENMETRICS}
_PORT. Every existing provide keeps pointing at the 4.0 broker, so
current consumers see no change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still listed the alarm, pending-packets and socket metrics as left at zero or unreachable. Describe the three new workloads, the ALARM_DRILL and RABBITMQ3_VERSION inputs, the drill's measured side effects, and how a consumer scrapes the 3.13 broker, and record the new measurement: all 58 in-scope metrics emitted and non-zero. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The drill restores the memory watermark and disk free limit it reads at startup. If either read failed or returned something that is not a number, the old code still raised both alarms and then "restored" a bad value, which could leave publishers blocked for the rest of the run. It now logs why and exits before the first raise. A restore that still fails after its retries also stops the loop, so the drill never raises the alarms again on top of alarms it could not clear. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Describe the new startup and restore guards, and how a consumer sets the fixture's knobs. They are Compose interpolation variables read from the environment of the evalya run: a task-level env entry in evalya.yaml, or an alias's override of one, lands only on the rabbitmq-full forwarder container, not on the broker, so declaring ALARM_DRILL there would not let a consumer turn the drill off. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6eb9758e5e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Created DOCS-15842 for the editorial review. |
The README table promised ACTIVITY_GEN=0 idles every workload, and activity-gen.sh calls it a quiescent env, matching the redisdb exemplar. load, unroutable and conn-churn ignored the switch and kept publishing, so the mode never gave a traffic-free broker. They now use the same sleep-infinity guard as autoack. The wrapper runs the image's default entrypoint (java -jar /perf_test/perf-test.jar) with the same arguments, so ACTIVITY_GEN=1 behavior is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AGENTS.md (Python Code Style > Type Hints) asks for type hints on new functions; slow_reader.py is new in this PR. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Validation ReportAll 21 validations passed. Show details
|
What does this PR do?
TL;DR: Adds a
rabbitmq-fullevalya task underrabbitmq/tests/: a full metric-coverage environment where all 58 in-scope dashboard and monitor metrics of the rabbitmq check are emitted live and non-zero. Test fixtures only, no check code touched.Files:
compose/full-coverage.compose: the environment (services below).evalya.yaml: therabbitmq-fulltask, itsprovides.*contract (RABBITMQ_*for the 4.0 broker,RABBITMQ3_*for the 3.13 broker,guest/guest), and a healthcheck through the forwarder.seed.sh,activity-gen.sh,alarm-drill.sh(POSIXsh),slow_reader.py(stdlib Python).README.md: services, the required check instances, knobs, side effects, and how coverage was measured.Services (13):
rabbitmq-broker-managementbroker, 4.0: management API (15672) and Prometheus plugin (15692). Runsalarm-drill.shin the background.seedloadactivity-genrabbitmq.queues.{created,declared,deleted}.count.unroutableconn-churnautoack*.messages.delivered.count.redeliverrabbitmq.queue.messages.persistent.unacked-swingslow-readerrabbitmq.connection.pending_packets.rabbitmq3-brokerload3rabbitmq-fullsocatforwarder (5672/15672/15692 to the 4.0 broker, 5673/15673/15693 to the 3.13 broker), gated on every service above.Knobs (Compose interpolation, read from the environment of the
evalya run, e.g.ALARM_DRILL=0 evalya run ...orevalya run -e ALARM_DRILL=0 ...):RABBITMQ_VERSION4.0seedandactivity-gen.RABBITMQ3_VERSION3.13ACTIVITY_GEN10idles every workload service, including the alarm drill.ALARM_DRILL10turns off only the alarm drill.They are deliberately not declared under the task's
envinevalya.yaml: a task-levelenventry (or an alias's override) lands only on therabbitmq-fullforwarder container, not on the broker. Verified with a two-service compose: taskenvset the variable on the entrypoint only, while host env and-ereached the dependency.Design points for review:
rabbitmq.queue.messages_ready, per-queue*.rate) and OpenMetrics names (rabbitmq.erlang.*,*.count), and neither backend emits the other's.rabbitmq.queue.messages.{acked,delivered,redelivered,delivered.ack}.countare in the target; the 3.13 broker exists only forrabbitmq.node.sockets_usedandrabbitmq.process.{open,max}_tcp_sockets.rabbitmq.node.{mem,disk}_alarmandrabbitmq.alarms.free_disk_space.watermarkread 1. If a baseline read fails or is not a number, it logs why and exits before the first raise; if a restore fails after its retries, it stops instead of raising again.Side effects:
loaddrops to 0 msg/s, then bursts when the alarm clears. A scrape inside a window readrabbitmq.queue.messages.publish.rate0 on every queue; no workload container exited. SetALARM_DRILL=0for uninterrupted publish traffic.slow-reader, and the 3.13 broker are live but not organic: they exist only to move specific metrics.Motivation
The rabbitmq check had no fixture that exercised its dashboard and monitor metrics end to end. The existing
compose/docker-compose.yamlstarts an idle broker, so rate and counter metrics read zero and per-queue metrics have no subjects. A published evalya fixture gives semantic-core's FTF a reusable environment to compare the check with the OTelrabbitmqreceiver.Validation
assets/dashboards/andassets/monitors/: 61 metrics, 58 from this check.data_streams.latency,data_streams.payload_size, andsystem.mem.totalcome from other sources and are excluded.agent check rabbitmq -t 2 --jsonindatadog/agent:7(7.83.1) with this branch's check mounted, four instances (OpenMetrics and management against 4.0.9, the OpenMetrics one also scraping thedetailedendpoint; the same two against 3.13.7), ten-t 2runs 21s apart starting about 3.5 minutes after the fixture turned healthy. 55 are non-zero on the 4.0 broker; the three socket metrics only on 3.13. The alarm metrics read 1 only in runs that land in a drill window.-t 2matters: one scrape emits no OpenMetrics.countmetric.rabbitmq-full: the drill parsed its baseline (vm_memory_high_watermark=0.6 disk_free_limit=50000000, matching the rawrabbitmqctl -q evaloutput on 4.0.9), raised both alarms 60s after start, and cleared them 25s later.rabbitmqctlinside the broker container (dash): an empty, non-numeric, or{absolute,"1GiB"}watermark each exits 1 with a log line and zeroset_*calls.sh, withsleepstubbed too: a non-numeric disk limit and a failing read also exit before any raise; a failing restore stops the loop after one raise; a normal cycle, and an{absolute,N}baseline, restore the values read at startup.evalya manifest validate --path rabbitmq/tests/evalya.yaml: OK, 23 warnings (ignoredrestart:policies, no registry mirror).ddev --no-interactive test --lint rabbitmq: pass.tests/, which is not shipped.RABBITMQ_VERSION.Review checklist (to be filled by reviewers)
qa/requiredif this PR needs QA validation, orqa/skip-qaif it does not. Exactly one of the two is required.backport/<branch-name>label to the PR and it will automatically open a backport PR once this one is merged🤖 Generated with Claude Code