Skip to content

Add rabbitmq evalya full-coverage fixture - #25224

Open
NouemanKHAL wants to merge 17 commits into
masterfrom
noueman/rabbitmq-evalya-fixture
Open

NouemanKHAL wants to merge 17 commits into
masterfrom
noueman/rabbitmq-evalya-fixture

Conversation

@NouemanKHAL

@NouemanKHAL NouemanKHAL commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

TL;DR: Adds a rabbitmq-full evalya task under rabbitmq/tests/: a full metric-coverage environment where all 58 in-scope dashboard and monitor metrics of the rabbitmq check are emitted live and non-zero. Test fixtures only, no check code touched.

Files:

  • compose/full-coverage.compose: the environment (services below).
  • evalya.yaml: the rabbitmq-full task, its provides.* contract (RABBITMQ_* for the 4.0 broker, RABBITMQ3_* for the 3.13 broker, guest/guest), and a healthcheck through the forwarder.
  • seed.sh, activity-gen.sh, alarm-drill.sh (POSIX sh), slow_reader.py (stdlib Python).
  • README.md: services, the required check instances, knobs, side effects, and how coverage was measured.

Services (13):

Service Role
rabbitmq-broker Primary -management broker, 4.0: management API (15672) and Prometheus plugin (15692). Runs alarm-drill.sh in the background.
seed One-shot: vhosts, static queues, exchanges, bindings, cluster name.
load perf-test, 4 producers and 4 consumers; publish rate alternates 30/90 msg/s every 60s against a 45 msg/s consumer, so rates and queue depth rise and fall.
activity-gen Queue declare/delete churn for rabbitmq.queues.{created,declared,deleted}.count.
unroutable Producer-only publisher to an unbound routing key, for the unroutable-dropped counters.
conn-churn Looping 20s perf-test clients, for connection/channel opened/closed counters and a varying consumer count.
autoack Auto-ack consumer, for *.messages.delivered.count.
redeliver nack/requeue consumer on a durable queue, for the redelivery counters and rabbitmq.queue.messages.persistent.
unacked-swing Consumer whose latency alternates every 120s, for a moving unacked count.
slow-reader Bespoke AMQP client that reads its socket slowly, for rabbitmq.connection.pending_packets.
rabbitmq3-broker Second broker on 3.13, for the three socket metrics 4.x dropped.
load3 Light traffic on the 3.13 broker, so its socket gauges count connections.
rabbitmq-full Entrypoint: socat forwarder (5672/15672/15692 to the 4.0 broker, 5673/15673/15693 to the 3.13 broker), gated on every service above.

Knobs (Compose interpolation, read from the environment of the evalya run, e.g. ALARM_DRILL=0 evalya run ... or evalya run -e ALARM_DRILL=0 ...):

Variable Default Effect
RABBITMQ_VERSION 4.0 Primary broker image tag, shared by seed and activity-gen.
RABBITMQ3_VERSION 3.13 Second broker image tag.
ACTIVITY_GEN 1 0 idles every workload service, including the alarm drill.
ALARM_DRILL 1 0 turns off only the alarm drill.

They are deliberately not declared under the task's env in evalya.yaml: a task-level env entry (or an alias's override) lands only on the rabbitmq-full forwarder container, not on the broker. Verified with a two-service compose: task env set the variable on the entrypoint only, while host env and -e reached the dependency.

Design points for review:

  • Two check instances (four with the 3.13 broker). The dashboards and monitors mix management-API names (rabbitmq.queue.messages_ready, per-queue *.rate) and OpenMetrics names (rabbitmq.erlang.*, *.count), and neither backend emits the other's.
  • Primary broker pinned to 4.x. The 4.x-only rabbitmq.queue.messages.{acked,delivered,redelivered,delivered.ack}.count are in the target; the 3.13 broker exists only for rabbitmq.node.sockets_used and rabbitmq.process.{open,max}_tcp_sockets.
  • Alarm drill. Every 180s (first after 60s) it lowers the memory watermark and raises the disk free limit for 25s, then restores the values read at startup, so rabbitmq.node.{mem,disk}_alarm and rabbitmq.alarms.free_disk_space.watermark read 1. If a baseline read fails or is not a number, it logs why and exits before the first raise; if a restore fails after its retries, it stops instead of raising again.

Side effects:

  • During a drill window the broker blocks every publishing connection: about 14% of the time (25s of every 180s). load drops to 0 msg/s, then bursts when the alarm clears. A scrape inside a window read rabbitmq.queue.messages.publish.rate 0 on every queue; no workload container exited. Set ALARM_DRILL=0 for uninterrupted publish traffic.
  • The drill, slow-reader, and the 3.13 broker are live but not organic: they exist only to move specific metrics.

Motivation

The rabbitmq check had no fixture that exercised its dashboard and monitor metrics end to end. The existing compose/docker-compose.yaml starts an idle broker, so rate and counter metrics read zero and per-queue metrics have no subjects. A published evalya fixture gives semantic-core's FTF a reusable environment to compare the check with the OTel rabbitmqreceiver.

Validation

  • Coverage target from assets/dashboards/ and assets/monitors/: 61 metrics, 58 from this check. data_streams.latency, data_streams.payload_size, and system.mem.total come from other sources and are excluded.
  • Coverage 58/58, all live and non-zero. Oracle: agent check rabbitmq -t 2 --json in datadog/agent:7 (7.83.1) with this branch's check mounted, four instances (OpenMetrics and management against 4.0.9, the OpenMetrics one also scraping the detailed endpoint; the same two against 3.13.7), ten -t 2 runs 21s apart starting about 3.5 minutes after the fixture turned healthy. 55 are non-zero on the 4.0 broker; the three socket metrics only on 3.13. The alarm metrics read 1 only in runs that land in a drill window. -t 2 matters: one scrape emits no OpenMetrics .count metric.
  • Alarm drill guard, on this head:
    • Live run of rabbitmq-full: the drill parsed its baseline (vm_memory_high_watermark=0.6 disk_free_limit=50000000, matching the raw rabbitmqctl -q eval output on 4.0.9), raised both alarms 60s after start, and cleared them 25s later.
    • Stubbed rabbitmqctl inside the broker container (dash): an empty, non-numeric, or {absolute,"1GiB"} watermark each exits 1 with a log line and zero set_* calls.
    • Same stubs on macOS sh, with sleep stubbed too: a non-numeric disk limit and a failing read also exit before any raise; a failing restore stops the loop after one raise; a normal cycle, and an {absolute,N} baseline, restore the values read at startup.
  • evalya manifest validate --path rabbitmq/tests/evalya.yaml: OK, 23 warnings (ignored restart: policies, no registry mirror).
  • ddev --no-interactive test --lint rabbitmq: pass.
  • No changelog: every file is under tests/, which is not shipped.
  • Not verified: consumption through an OCI federation context (only local-path runs), non-default RABBITMQ_VERSION.

Review checklist (to be filled by reviewers)

  • Feature or bugfix MUST have appropriate tests (unit, integration, e2e)
  • Add qa/required if this PR needs QA validation, or qa/skip-qa if it does not. Exactly one of the two is required.
  • If you need to backport this PR to another branch, you can add the backport/<branch-name> label to the PR and it will automatically open a backport PR once this one is merged

🤖 Generated with Claude Code

Onboards rabbitmq to evalya with a rabbitmq-full task that drives the
check to emit the metrics referenced by the OOTB dashboards and the
recommended monitors.

A single -management broker serves both check backends, since the assets
draw metric names from each: the management API and the Prometheus
plugin. Continuous AMQP traffic comes from perf-test, because
rabbitmqadmin cannot populate channel, connection, or delivery counters
over HTTP. Queue churn is separate so the node-wide queue counters keep
advancing past perf-test's long-lived queues.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@NouemanKHAL NouemanKHAL added the qa/skip-qa Automatically skip this PR for the next QA label Sep 15, 2026
@dd-octo-sts

dd-octo-sts Bot commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Dispatcher tests: passed

Dispatcher beta: informational only
Existing CI remains the merge signal.

  12/12 jobs

✅ 12 passed · nothing failed

Batches · ✅ batch-01 12/12

Dispatcher finished on 2caa590 — GitHub Run · Dispatcher Logs.

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

evalya-impact-summary

evalya impact analysis
Impact analysis: 0 selected, 0 skipped (of 0 test tasks)
Publish tasks:   3 (always emitted)
Diff (7 files):
  rabbitmq/tests/README.md
  rabbitmq/tests/activity-gen.sh
  rabbitmq/tests/alarm-drill.sh
  rabbitmq/tests/compose/full-coverage.compose
  rabbitmq/tests/evalya.yaml
  rabbitmq/tests/seed.sh
  rabbitmq/tests/slow_reader.py

Debug a specific task: evalya plan impact --path <path> --task <task>

Learn more about CI impact filtering

@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Tests  Code Coverage

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 0.00%
• Overall Coverage: 84.38%

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: a279494 | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Disk usage change

Commit a8787d8 compared against a46ed3e.

No integration or dependency changed size.

NouemanKHAL and others added 4 commits September 24, 2026 16:37
evalya starts only a task's target service and its depends_on chain.
The task targeted the broker, which nothing depends on for the
workload, so seed, load, and activity-gen never ran and every
message, consumer, and queue metric read zero under evalya.

The task now targets a socat entrypoint gated on the workload that
forwards the broker's ports. The broker is renamed rabbitmq-broker
because consumers alias this fixture as `rabbitmq`, and evalya gives
the entrypoint that hostname, which made socat forward to itself. Host
ports are no longer published, so the fixture cannot clash with a
local broker or a concurrent run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
perf-test consumers had no prefetch limit, so the whole backlog sat
unacked, messages_ready read zero, and unacked grew without bound.
--qos 50 keeps the backlog ready and x-max-length=2000 caps it.

Nothing published unroutable messages, so the unroutable-dropped
counters read zero. A producer-only perf-test now publishes to an
unbound routing key; it stays connected because the per-channel
counter disappears with the channel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fixed publish rate leaves every rate metric and the queue depth
flat once the backlog hits x-max-length, so OTel vs DD comparisons
have no shape to match. Alternating the publish rate every 60s
between below and above the consumer rate makes them rise and fall.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The load connections live for the whole run, so the connection and
channel opened/closed counters and the consumer count stop moving
after startup. A looping short-lived perf-test keeps them advancing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NouemanKHAL and others added 10 commits September 29, 2026 21:46
Six dashboard/monitor metrics were emitted but stuck at 0, so they
proved the mapping fired but not that the fixture exercised it:

- delivered.count (channel, queue) only counts auto-ack deliveries,
  and every consumer acked manually. An autoack perf-test fixes it.
- redeliver/redelivered.count and messages.persistent had no
  redelivery or persistent backlog. A nack/requeue perf-test on a
  durable queue with persistent messages drives all three.
- messages_unacknowledged.rate stayed 0 because load's unacked count
  is pinned at its --qos. Capping conn-churn's consumer below its
  publish rate makes its unacked count climb each cycle.

The new services idle with ACTIVITY_GEN=0, like activity-gen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README claimed 56/58 metrics emitted, but coverage should count
only non-zero values: 11 of those 56 read 0. It now reports the
non-zero figure (51/58), describes the autoack and redeliver
services and the ACTIVITY_GEN scope, and names each metric left at
0 with why: sockets_used is zeroed by RabbitMQ 4.0, the alarms are
intentionally not raised because an alarm blocks every publisher,
and pending_packets needs a client that stops reading its socket.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Capping conn-churn's consumer at 5 msg/s made unacked a 30s sawtooth:
it climbed while the 20s consumer lived, then the backlog was requeued
at close. That cycle aliases against the scrape intervals. In a 600s
comparator run every queue's average unacked matched across scrapers
except churn-conn: OTel rabbitmqreceiver 17.8, DD OpenMetrics 35.2,
DD management 29.4, so both unacked equivalences failed (6.06% and
6.39%) where they passed before.

conn-churn's consumer is uncapped again, so that queue stays near 0.
A new unacked-swing service drives messages_unacknowledged.rate from
one long-lived connection instead: a constant 4 msg/s publisher and a
consumer whose --variable-latency alternates every 120s between 2 and
8 msg/s of capacity, so unacked ramps to about 240 and back on a 240s
cycle, slow enough for every scraper's average to agree. It idles
with ACTIVITY_GEN=0, like the other added services.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still said conn-churn drives messages_unacknowledged.rate.
It now explains why conn-churn's consumer stays uncapped (a backlog
requeued every 30s aliases against the scrape intervals), describes
unacked-swing and its slow 240s cycle, and adds it to the
ACTIVITY_GEN scope and the entrypoint's gate list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.node.mem_alarm, rabbitmq.node.disk_alarm and
rabbitmq.alarms.free_disk_space.watermark are dashboard and monitor
targets that read 0 on a healthy broker, so the rabbitmq-full fixture
never covered them.

The primary broker now runs a background drill that raises both alarms
for 25s every 180s via rabbitmqctl, then restores the values it read at
startup. Running inside the broker container reuses the node's cookie
and name, which a sidecar would need shared. 25s is longer than the 15s
DD and OTel scrape intervals plus the ~8s management stats lag, and
shorter than perf-test's 30s confirm timeout (evalya ignores restart
policies, so an exited perf-test would stay down).

Publishers block during a window. ALARM_DRILL=0, or ACTIVITY_GEN=0,
turns it off.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.connection.pending_packets counts bytes the broker has queued
on a socket but not yet handed to the kernel. It is non-zero only when
a client reads slower than the broker delivers, and every perf-test
client keeps up, so it read 0 in every scrape.

slow-reader is bespoke traffic: a stdlib-only AMQP 0-9-1 client (no
pip install at runtime) whose auto-ack consumer reads at ~32 KiB/s
through a 4 KiB receive buffer while a second connection publishes
~100 KiB/s. It throttles instead of stopping, because the broker closes
a connection whose socket send blocks for 30s. Auto-ack leaves nothing
unacked, x-max-length=1000 caps the queue, both rates are fixed, and
the consumer reconnects every 600s. Idles with ACTIVITY_GEN=0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rabbitmq.node.sockets_used, rabbitmq.process.max_tcp_sockets and
rabbitmq.process.open_tcp_sockets are unreachable on RabbitMQ 4.x,
which stopped tracking TCP sockets. The primary broker stays on 4.0,
because the 4.x-only queue delivery counters are also targets and
seed.sh needs the v2 rabbitmqadmin.

A second broker on 3.13 (RABBITMQ3_VERSION), with two perf-test
connections, is forwarded by the entrypoint on ports shifted by one and
published as RABBITMQ3_HOST and RABBITMQ3_{AMQP,MANAGEMENT,OPENMETRICS}
_PORT. Every existing provide keeps pointing at the 4.0 broker, so
current consumers see no change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still listed the alarm, pending-packets and socket metrics
as left at zero or unreachable. Describe the three new workloads, the
ALARM_DRILL and RABBITMQ3_VERSION inputs, the drill's measured side
effects, and how a consumer scrapes the 3.13 broker, and record the
new measurement: all 58 in-scope metrics emitted and non-zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The drill restores the memory watermark and disk free limit it reads
at startup. If either read failed or returned something that is not a
number, the old code still raised both alarms and then "restored" a
bad value, which could leave publishers blocked for the rest of the
run. It now logs why and exits before the first raise.

A restore that still fails after its retries also stops the loop, so
the drill never raises the alarms again on top of alarms it could not
clear.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Describe the new startup and restore guards, and how a consumer sets
the fixture's knobs. They are Compose interpolation variables read
from the environment of the evalya run: a task-level env entry in
evalya.yaml, or an alias's override of one, lands only on the
rabbitmq-full forwarder container, not on the broker, so declaring
ALARM_DRILL there would not let a consumer turn the drill off.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@NouemanKHAL
NouemanKHAL marked this pull request as ready for review September 30, 2026 14:15
@NouemanKHAL
NouemanKHAL requested review from a team as code owners September 30, 2026 14:15
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T14:21:04.380973Z 6eb9758 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6eb9758e5e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread rabbitmq/tests/slow_reader.py Outdated
Comment thread rabbitmq/tests/compose/full-coverage.compose
@buraizu buraizu self-assigned this Sep 30, 2026
@buraizu buraizu added the editorial review Waiting on a more in-depth review from a docs team editor label Sep 30, 2026 — with ddtool CLI

buraizu commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Created DOCS-15842 for the editorial review.

NouemanKHAL and others added 2 commits October 2, 2026 12:19
The README table promised ACTIVITY_GEN=0 idles every workload, and
activity-gen.sh calls it a quiescent env, matching the redisdb
exemplar. load, unroutable and conn-churn ignored the switch and kept
publishing, so the mode never gave a traffic-free broker. They now use
the same sleep-infinity guard as autoack. The wrapper runs the image's
default entrypoint (java -jar /perf_test/perf-test.jar) with the same
arguments, so ACTIVITY_GEN=1 behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AGENTS.md (Python Code Style > Type Hints) asks for type hints on new
functions; slow_reader.py is new in this PR.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chouetz chouetz added the internal Identify a non-fork PR label Oct 2, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Validation Report

All 21 validations passed.

Show details
Validation Description Status
agent-reqs Verify check versions match the Agent requirements file ✅
ci Validate CI configuration and code coverage settings ✅
codeowners Validate every integration has a CODEOWNERS entry ✅
config Validate default configuration files against spec.yaml ✅
dep Verify dependency pins are consistent and Agent-compatible ✅
http Validate integrations use the HTTP wrapper correctly ✅
imports Validate check imports do not use deprecated modules ✅
integration-style Validate check code style conventions ✅
jmx-metrics Validate JMX metrics definition files and config ✅
labeler Validate PR labeler config matches integration directories ✅
legacy-signature Validate no integration uses the legacy Agent check signature ✅
license-headers Validate Python files have proper license headers ✅
licenses Validate third-party license attribution list ✅
metadata Validate metadata.csv metric definitions ✅
models Validate configuration data models match spec.yaml ✅
openmetrics Validate OpenMetrics integrations disable the metric limit ✅
package Validate Python package metadata and naming ✅
qa-label Validate the pull request declares whether it needs QA for the next Agent release ✅
readmes Validate README files have required sections ✅
saved-views Validate saved view JSON file structure and fields ✅
version Validate version consistency between package and changelog ✅

View full run

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

editorial review Waiting on a more in-depth review from a docs team editor integration/rabbitmq internal Identify a non-fork PR qa/skip-qa Automatically skip this PR for the next QA team/agent-integrations team/documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants