evmonly/evmonlyapp: size the next block's decode from its cost and the time it has; time PrepareBlock - #4262
Conversation
… PrepareBlock Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
PR SummaryMedium Risk Overview The EVM-only app passes half the smoothed Reviewed by Cursor Bugbot for commit 3ae1d46. Bugbot is set up for automated code reviews on this repo. Configure here. |
|
The latest Buf updates on your PR. Results from workflow Buf / buf (pull_request).
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## giga-1 #4262 +/- ##
===========================================
+ Coverage 65.55% 90.86% +25.31%
===========================================
Files 2081 22 -2059
Lines 157460 3504 -153956
===========================================
- Hits 103222 3184 -100038
+ Misses 54097 320 -53777
+ Partials 141 0 -141
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Caps PrepareBlock's parse pool at max(2, GOMAXPROCS/4) and adds an evmonly_prepare{phase="parse"} timer so the tradeoff against OCC speculation can be measured. The change is correct and the single-timer-per-app assumption holds (PrepareBlock has one caller, the block fetcher goroutine); the only note is that the same knob also throttles the synchronous fallback decode inside FinalizeBlock.
Findings: 0 blocking | 1 non-blocking | 1 posted inline
Blockers
- None at the file/PR level.
Non-blocking
- None at the file/PR level.
- 1 suggestion(s)/nit(s) flagged inline on specific lines.
… pool Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…as, not a fixed share of the processors PrepareBlockWithin picks the fewest parse workers the block's decode is expected to fit in a budget on, from a moving estimate of the per-tx cost measured on earlier decodes; evmonlyapp budgets half the previous block's execution time. No budget (the unprepared FinalizeBlock path, or before the first block) decodes on every worker. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit a5fad63. Configure here.
|
Claude finished @bdchatham's task in 3m 31s —— View job Review
No blockers. The arithmetic is right and the wiring is sound. Four notes on the control loop, all non-blocking for an A/B, and two nits. What I verified
I could not run 1. The budget comes from block n−1 but is spent against block n
The consequence is sharper than "we might miss a prepare", because prepare is synchronously on the critical path: the fetcher does Worth considering: clamp the budget by something that reflects the block actually in flight, or re-check remaining time mid-decode rather than committing the worker count up front. 2. The EWMA measures a quantity that depends on the knob it sets
It converges rather than diverges (fixed point 3. A short block hands the next prepare every worker(Same as Bugbot's, and it's real.) A block that executes in 1 ms gives 4. Sender-cache hit rate moves per-tx cost by ~50×, and the loop recovers slowly
Also note Nits
|
…imate at once and lower it gradually Review follow-ups on the adaptive sizing: the budget is half an EWMA of the prepared blocks' execution time rather than half the last one, so a single short block does not hand the next decode every processor; the per-tx cost jumps up to any decode that overran its estimate and decays down, since an undersized decode delays the block that needs it; unbudgeted decodes (the unprepared FinalizeBlock path) no longer feed the estimate; the parse phase timer ends before the prepared slot is taken. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Re seidroid's review — addressed in 3ae1d46, one item at a time: 1 & 3 (budget from n−1; short block hands out every worker): the budget is now half an EWMA (1/8) of the prepared blocks' execution time ( 2 (the estimate depends on the knob): agreed, that's inherent to measuring wall time under contention; the fixed point sits above the nominal count, which is the safe side. Reworded the 4 (sender-cache hit rate; slow recovery): Nits: the parse phase now ends when |

Describe your changes and provide context
Since #4260,
PrepareBlockdecodes block n+1 (RLP decode + ecrecover for senders CheckTx did not see) onGOMAXPROCSworkers while block n's OCC speculation holds a worker per processor. Theocc_speculateshare rose 0.19 → 0.36 s/s on that roll, so the prepare pool is competing with the block on the critical path for work that only has to finish before block n does.Rather than a fixed share of the processors, the decode is sized per block from the work and the time it has:
So on today's testnet-2 blocks (~1,850 txs, ~50 µs/tx ecrecover, ~20 ms execution) it decodes on ~9–13 workers of 32 (the measured cost includes contention, so the fixed point sits a little above the nominal count — the safe side); on a small block it uses 1–2; on a 5k-tx block it grows; and on a different box or tx mix the per-tx estimate re-fits itself. The asymmetry is deliberate: an undersized decode delays the block that needs it (the fetcher hands the prepared block over synchronously), an oversized one only spends processors, so the estimate jumps up on the first poorly-cached block and comes down slowly. The budget is smoothed for the same reason — a single short block dents it rather than zeroing it, which would hand the next decode every worker. The unprepared
FinalizeBlockpath (executeBlockPipelined) keepsPrepareBlockwith no budget, so a decode that is on the critical path still uses every worker.Config.ParseWorkersremains the ceiling (stillGOMAXPROCSin production wiring).Also adds an
evmonly_preparephase timer (parse) around the decode so prepare wall time vs block time is visible after the roll;evmonly_prepare_phase_duration_seconds_total{phase="parse"}.Measurement after roll:
occ_speculates/s and executed tx/s vs the #4261 baseline (~94k/validator), andevmonly_prepareparse wall staying well underevmonly_finalizeexecute wall. If parse wall approaches execute wall, a mid-decode re-check of the remaining budget is the next step.Testing performed to validate your change
giga/evmonly:parse_sizer_test.go— every worker until the first estimate, fits the estimated cost in the budget (5 → 3 → 1 workers as the budget grows; capped at the ceiling / tx count), no budget uses every worker, estimate rises at once and decays gradually, ignores empty observations.go test -race ./giga/evmonly/green.sei-tendermint/internal/evmonlyapp:TestPrepareBudgetIsAShareOfTheTypicalExecution,TestExecuteEstimateSmoothsOverBlocks;go test -race ./internal/evmonlyapp/green (prepared-hit and unprepared-fallback paths both covered by the existing tests).golangci-lint runon both packages,make fmtcheckclean.Link to Devin session: https://app.devin.ai/sessions/ff612badcded4aa5914ea408dbb41888
Open in Devin Desktop: https://app.devin.ai/desktop/session/ff612badcded4aa5914ea408dbb41888?variant=devin
Requested by: @bdchatham