Add mongo evalya full-coverage fixture - #25423
NouemanKHAL wants to merge 7 commits into
Conversation
The mongo check had no environment that populates its dashboard and monitor metrics: the pytest compose files start idle nodes, so rate metrics read zero and the mongos-only and replica-set-only metrics never appear together. mongo-full is the smallest sharded cluster that covers both roles, with a delayed secondary so replication lag and repl counters are non-zero, plus a seed and a continuous workload. It is published for downstream consumers comparing the check with the OTel mongodbreceiver. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
evalya-impact-summaryevalya impact analysis |
|
✅ All CI checks and tests passed. 🎉 All green!🧪 All tests passed 🎯 Code Coverage (details) 🔗 Commit SHA: 6715aea | Docs | View more details | Give us feedback! |
serverStatus reports locks.<type>.acquireWaitCount and timeAcquiringMicros only after an acquisition waits on a conflicting mode, and the one-at-a-time workload never conflicts, so the MongoDB FTF comparator saw no data on either side for 27 lock metric pairs. activity-gen now runs bounded contention: a slow $where writer holds intent locks on shard-a while periodic holders request S and X modes on Collection, Database and Global (createIndexes, dropIndexes, dbHash, cross-database rename, fsync lock, setUserWriteBlockMode), and a capped collection write takes the Metadata lock. Every held lock is released in a finally. This makes 21 of the 27 pairs non-zero on the shard primary, from both the mongo check and the mongodbreceiver. The other 6 read lock modes MongoDB 8.0 never acquires in steady state (Metadata S, oplog S, and waits on oplog IX); the README records why. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
setUserWriteBlockMode global:true blocks every user write in the cluster for its window, including writes a consumer sends through the forwarded ports. LOCK_DRILL=0 (default 1, same host-env convention as ACTIVITY_GEN) now skips it and the fsync lock holder, while the rest of the lock contention keeps running. Both locks outlive the mongosh session that took them, so a stop in the middle of a window could leave the cluster blocked. An exit trap now kills the holders, then runs fsyncUnlock and turns the write block off, best effort. fsyncUnlock goes first because setUserWriteBlockMode waits behind a held fsync lock. The main iteration runs in the background and is waited on: sh defers traps until a foreground command returns, and an iteration stuck behind the fsync lock kept docker stop from reaching the trap before SIGKILL. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README described the write block as something only the lock writers see. It blocks every user write in the cluster, so say so, list who is affected, and document LOCK_DRILL, the exit trap, and ACTIVITY_DURATION (read by the script but not passed by the compose file). Also note that the fixture's knobs are read from the evalya run's environment: a task-level env entry reaches only the forwarder. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b2e9c68581
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The compose services read DB_USERNAME/DB_PASSWORD from the host env,
but the provides labels published a literal "datadog", so an override
made consumers authenticate with the wrong user. evalya interpolates
${VAR:-default} in manifest fields from the same env lookup it passes
to compose, so the labels now track the effective values.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The forwarder runs three socat listeners under `wait`, so one can die while the container stays up, and the healthcheck only probed the upstreams. Probe the advertised ports too. The upstream probes stay: a forked socat listener accepts even when its upstream is down. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seed.sh and activity-gen.sh embed DB_USERNAME/DB_PASSWORD unescaped in mongodb:// URIs, so an override with a reserved character misparses. The defaults are safe and the override is a fixture knob, so fail fast in seed with a clear message and document the restriction, rather than rework every mongosh call site in both scripts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Validation ReportAll 21 validations passed. Show details
|
What does this PR do?
TL;DR: Adds a
mongo-fullevalya fixture (mongo/tests/evalya.yaml): the smallest sharded MongoDB 8.0 cluster that exercises both the mongos and the replica-set collectors, a seed, and a continuous workload that includes lock contention. Stacked on #25153 (onboard-evalya skill). Test fixtures only.shard01with a primary (shard-a) and a hidden delayed secondary (shard-b: priority 0, votes 0,secondaryDelaySecs5) so replication lag andrepl.*counters move; mongos; asocatentrypoint forwarding 27017 (mongos), 27018 (shard-a), 27019 (shard-b), same pattern as Add rabbitmq evalya full-coverage fixture #25224.seed.sh: adds the shard, creates thedatadoguser, seeds a pre-split sharded collection with indexes.activity-gen.sh: a mixed workload through mongos (inserts, indexed and scan queries, getMores, updates, aggregates, deletes, periodic DDL, a long-lived session), plus lock contention inlockdrill:$wherewriter holding intent locks, two fast writers anddbStats, and a holder every 3s runningcreateIndexes/dropIndexes,dbHash, a cross-databaserenameCollection, and anfsynclock;setUserWriteBlockModeon then off (Global X).logicalSessionRefreshMillis=10000, sosessions.countappears within a short run.Knobs (Compose interpolation, read from the environment of the
evalya run, e.g.LOCK_DRILL=0 evalya run ...orevalya run -e LOCK_DRILL=0 ...):MONGO_VERSION8.0DB_USERNAME/DB_PASSWORDdatadog/datadogseed.ACTIVITY_GEN10keepsactivity-genup but idle.LOCK_DRILL10skips thefsynclock and thesetUserWriteBlockModeholder; the rest of the lock contention still runs.Side effect of the lock drill:
setUserWriteBlockModewithglobal: trueblocks all user writes cluster-wide for its window, every 3s. That includes the fixture's ownactivity-geniteration, the shard holder's index build (logged, restarted 5s later), and any consumer writing through the forwarded ports. Read-only scrapers are not blocked.LOCK_DRILL=0opts out. An exit trap kills the holders and runsfsyncUnlockthensetUserWriteBlockModeoff, best effort, so stoppingactivity-genmid-window does not leave the cluster blocked.ACTIVITY_DURATION(script-only, not passed by the compose file) is documented in the README.Motivation
semantic-core's FTF federates
*-fullfixtures (redis-full, rabbitmq-full #25224) to compare the Datadogmongocheck with the OTelmongodbreceiver. Mongo had no evalya fixture, and the existingmongo-shard.yamlcompose (11 mongods, no healthchecks,docker execinit) is not usable as one. The lock workload exists because serverStatus reportslocks.<type>.acquireWaitCountandtimeAcquiringMicrosonly after an acquisition waits, which one-at-a-time traffic never causes, so the comparator had no data on either side for 27 lock metric pairs.Validation
mongodbreceiver(MongoDB 8.0.32, three 15s windows,LOCK_DRILL=1). The other 6 read modes 8.0 never acquires in steady state (MetadataR,oplogR, and waits onoplogw); the README records why.agent check mongo(Agent 7.83.1, three instances: mongos, primary, secondary). The 4 zeros:chunks.jumbo,globallock.currentqueue.{readers,writers},extra_info.page_faultsps. The two queue gauges are point-in-time, and the lock workload queues operations only in ~100ms bursts, so they are still expected to read 0 at scrape time; that expectation is reasoned, not measured.metadata.csv: 221/328 emitted, 162 non-zero.mongodbreceiver(otelcol-contrib 0.161.0, all metrics enabled): 50/52 emitted, 43 non-zero.LOCK_DRILL, live on this head:LOCK_DRILL=1: shard-aserverStatus().locksshowed non-zeroacquireWaitCountfor Global, Database, and Collection in all four modes. With anfsynclock held by hand,docker stoponactivity-genreturned in 0.5s, loggedlock drill released, thefsyncLockWorkerlock was gone, and a write through mongos succeeded.LOCK_DRILL=0: the container gotLOCK_DRILL=0, the log showedlock drill disabledand 40+ iterations with noUserWritesBlockedand no failures. Database, Collection, and Metadata lock fields were still populated; the Global wait fields were absent, since the drill's holders are the only Global S/X requests.evalya manifest validate --path mongo/tests/evalya.yaml: OK (warnings: ignoredrestart:policies, no registry mirror).ddev --no-interactive test --lint mongo: pass. Tests-only change, no changelog.MONGO_VERSION,ACTIVITY_GEN=0.Review checklist (to be filled by reviewers)
qa/requiredif this PR needs QA validation, orqa/skip-qaif it does not. Exactly one of the two is required.backport/<branch-name>label to the PR and it will automatically open a backport PR once this one is merged🤖 Generated with Claude Code