perf: vectorize small u8 table take with NEON - #9571
1 benchmark regressed
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 2 improved benchmarks
❌ 1 regressed benchmark
✅ 1984 untouched benchmarks
⏩ 54 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | compact[(2048, 90)] |
1.6 µs | 1.8 µs | -10.77% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[16000000] |
8,736.5 µs | 778.6 µs | ×11 |
| ⚡ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
546.5 µs | 50.4 µs | ×11 |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/small-u8-table-take-neon (21fa748) with ji/small-u8-table-take-benchmark (d49393c)
Footnotes
-
54 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩