Skip to content

Make the warm-up inbox benchmark test deterministic - #1226

Merged
dahlia merged 1 commit into
fedify-dev:2.3-maintenancefrom
dahlia:bugfix/deterministic-inbox-benchmark-test
Oct 4, 2026
Merged

dahlia merged 1 commit into
fedify-dev:2.3-maintenancefrom
dahlia:bugfix/deterministic-inbox-benchmark-test

Conversation

@dahlia

@dahlia dahlia commented Oct 4, 2026

Copy link
Copy Markdown
Member

Fixes #1219. Follow-up to #1220.

Why it could flake

The runner takes its baseline snapshot on the first send scheduled after the warm-up, and reports server metrics only if that snapshot exists. With two closed-loop workers, a 120ms warm-up, and a 400ms real-clock window, warm-up deliveries could keep both workers busy until the window closed. No measured send would start, so server came back null. The success-rate assertion still passed, because an empty measured window reports 1.

How it's fixed

The test now uses constant open-loop arrivals at 10/s over 400ms under a fake clock. Sends at 0 and 100ms fall inside the warm-up; sends at 200 and 300ms are measured. Open-loop arrivals are scheduled up front rather than after earlier sends finish, so the measured sends always happen, at worst a little late, and the baseline with them.

Two new assertions require exactly two requests counted client-side and all four deliveries reaching the inbox listener.

The fake clock from #1220 is now a shared createFakeClock() helper in inbox.test.ts.

Since ServerMetrics exposes only percentiles, the test can't verify that server metrics exclude warm-up verifications. stats-client.test.ts covers the subtraction.

The "reports server metrics scoped past the warm-up" test ran a
closed-loop load for 400ms of real time with a 120ms warm-up.  Server
metrics are only reported once the runner takes its baseline snapshot,
which happens lazily on the first send scheduled past the warm-up.  If
both closed-loop workers were still busy with warm-up deliveries when
the deadline passed (a cold target, CI scheduling), no measured send
was ever dispatched, the baseline was never taken, and server was null.
The success-rate assertion could not catch this, because an empty
measured window reports a success rate of 1.

The test now uses constant open-loop arrivals at 10/s over 400ms under
a fake clock, which schedules exactly four sends: two inside the
warm-up and two measured.  Open-loop scheduling does not depend on
completions, so the measured sends, and with them the baseline, always
happen however slow the deliveries are.  It also asserts that exactly
two measured requests were counted client-side and that all four
deliveries reached the inbox listener, so a vacuous pass is no longer
possible.

The fake clock used by the multi-recipient tests is now a shared
createFakeClock() helper.

Fixes fedify-dev#1219

Assisted-by: Claude Code:claude-opus-5-5
Assisted-by: Codex:gpt-6-astra
@dahlia dahlia self-assigned this Oct 4, 2026
@dahlia dahlia added component/cli CLI tools related component/ci CI/CD workflows and GitHub Actions labels Oct 4, 2026
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (1)
CONTRIBUTING.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 4cb6cd04-7119-47c6-9d2f-234171351a51
📥 Commits

Reviewing files that changed from the base of the PR and between db3f75f and b236d2d.

📒 Files selected for processing (1)
  • packages/cli/src/bench/scenarios/inbox.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.


📝 Walkthrough

Walkthrough

The inbox benchmark tests now use a shared fake clock. The warm-up test uses scheduled arrivals and checks measured deliveries, listener deliveries, and server metrics.

Changes

Inbox benchmark tests

Layer / File(s) Summary
Fake clock and benchmark assertions
packages/cli/src/bench/scenarios/inbox.test.ts
A shared fake clock advances to scheduled arrival times. The warm-up test uses four constant-rate arrivals and checks measured deliveries, listener deliveries, and server signature-verification metrics. The multi-recipient test uses the shared clock.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Other

Merge Risk: ⚪ Minimal · up to b236d

The supplied evidence identifies no issue that needs to be fixed before merging.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly states the main change: making the warm-up inbox benchmark test deterministic.
Description check ✅ Passed The description explains the flaky timing issue and how the fake clock and open-loop arrivals address it.
Linked Issues check ✅ Passed Issue [#1219] requests a deterministic measured request independent of delivery delays. inbox.test.ts now configures constant-rate arrivals at 10/s for 400ms with a fake clock and a 120ms warm-up. T…
Out of Scope Changes check ✅ Passed The shared createFakeClock() helper also replaces the duplicate fake clock in the multi-recipient tests. That refactor supports deterministic arrival scheduling in the same inbox test file. The supp…
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 1 files.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 4, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ All tests successful. No failed tests found.
see 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@dahlia
dahlia merged commit 29f4a6d into fedify-dev:2.3-maintenance Oct 4, 2026
17 checks passed
@dahlia
dahlia deleted the bugfix/deterministic-inbox-benchmark-test branch October 4, 2026 14:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

component/ci CI/CD workflows and GitHub Actions component/cli CLI tools related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant