Skip to content

lowram: Keep w1 in packed form during signature generation - #1345

Merged
mkannwischer merged 5 commits into
mainfrom
lowram-sign-pack-w1
Sep 1, 2026
Merged

lowram: Keep w1 in packed form during signature generation#1345
mkannwischer merged 5 commits into
mainfrom
lowram-sign-pack-w1

Conversation

@mkannwischer

@mkannwischer mkannwischer commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

The signing attempt held w1 as a full mld_polyveck, but w1 is only needed as the w1Encode input to H and, coefficient-wise, as the a1 argument of MakeHint. Both work from the packed encoding, so Decompose and w1Encode now run one polynomial at a time into a MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers each row with the new mld_polyw1_unpack.

With w1 gone, the matrix-vector scratch is the sole remaining user of the buffer the two shared. In REDUCE_RAM mode that scratch is a single polynomial rather than a polyvecl, and it now shares storage with z, so the buffer disappears.

Signing allocation in bytes:

default REDUCE_RAM
ML-DSA-44 44704 -> 44448 13120 -> 9792
ML-DSA-65 69312 -> 68032 17248 -> 11872
ML-DSA-87 108224 -> 107200 21344 -> 14176

mld_polyveck_decompose and mld_polyveck_pack_w1 have no other callers and are replaced by the fused mld_polyveck_decompose_pack_w1.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 46502 cycles 46534 cycles 1.00
ML-DSA-44 sign 131401 cycles 131164 cycles 1.00
ML-DSA-44 verify 47327 cycles 47342 cycles 1.00
ML-DSA-65 keypair 81701 cycles 81716 cycles 1.00
ML-DSA-65 sign 215390 cycles 215441 cycles 1.00
ML-DSA-65 verify 79329 cycles 79334 cycles 1.00
ML-DSA-87 keypair 132436 cycles 132496 cycles 1.00
ML-DSA-87 sign 277552 cycles 277435 cycles 1.00
ML-DSA-87 verify 133580 cycles 133546 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112924 cycles 112777 cycles 1.00
ML-DSA-44 sign 403184 cycles 401308 cycles 1.00
ML-DSA-44 verify 119532 cycles 119390 cycles 1.00
ML-DSA-65 keypair 193857 cycles 192926 cycles 1.00
ML-DSA-65 sign 651728 cycles 649924 cycles 1.00
ML-DSA-65 verify 193025 cycles 193010 cycles 1.00
ML-DSA-87 keypair 318845 cycles 318891 cycles 1.00
ML-DSA-87 sign 831120 cycles 828752 cycles 1.00
ML-DSA-87 verify 321851 cycles 321752 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 52243 cycles 52005 cycles 1.00
ML-DSA-44 sign 156689 cycles 155314 cycles 1.01
ML-DSA-44 verify 54300 cycles 54379 cycles 1.00
ML-DSA-65 keypair 89232 cycles 89891 cycles 0.99
ML-DSA-65 sign 252045 cycles 255341 cycles 0.99
ML-DSA-65 verify 89122 cycles 89591 cycles 0.99
ML-DSA-87 keypair 143951 cycles 142178 cycles 1.01
ML-DSA-87 sign 310297 cycles 310461 cycles 1.00
ML-DSA-87 verify 139225 cycles 139090 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 51564 cycles 51457 cycles 1.00
ML-DSA-44 sign 162252 cycles 161886 cycles 1.00
ML-DSA-44 verify 54672 cycles 54811 cycles 1.00
ML-DSA-65 keypair 90155 cycles 90177 cycles 1.00
ML-DSA-65 sign 266116 cycles 268015 cycles 0.99
ML-DSA-65 verify 89823 cycles 89977 cycles 1.00
ML-DSA-87 keypair 145851 cycles 145957 cycles 1.00
ML-DSA-87 sign 335809 cycles 335685 cycles 1.00
ML-DSA-87 verify 145427 cycles 144930 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 47026 cycles 46740 cycles 1.01
ML-DSA-44 sign 139819 cycles 145382 cycles 0.96
ML-DSA-44 verify 49294 cycles 51476 cycles 0.96
ML-DSA-65 keypair 82252 cycles 82956 cycles 0.99
ML-DSA-65 sign 227024 cycles 229646 cycles 0.99
ML-DSA-65 verify 82074 cycles 82585 cycles 0.99
ML-DSA-87 keypair 130391 cycles 129744 cycles 1.00
ML-DSA-87 sign 280889 cycles 280564 cycles 1.00
ML-DSA-87 verify 129532 cycles 128041 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112309 cycles 112045 cycles 1.00
ML-DSA-44 sign 354052 cycles 353393 cycles 1.00
ML-DSA-44 verify 117053 cycles 117282 cycles 1.00
ML-DSA-65 keypair 194750 cycles 194574 cycles 1.00
ML-DSA-65 sign 583585 cycles 583697 cycles 1.00
ML-DSA-65 verify 192910 cycles 193189 cycles 1.00
ML-DSA-87 keypair 320601 cycles 320560 cycles 1.00
ML-DSA-87 sign 747320 cycles 747028 cycles 1.00
ML-DSA-87 verify 318394 cycles 318140 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 100288 cycles 100193 cycles 1.00
ML-DSA-44 sign 369472 cycles 369630 cycles 1.00
ML-DSA-44 verify 110068 cycles 110334 cycles 1.00
ML-DSA-65 keypair 176009 cycles 175189 cycles 1.00
ML-DSA-65 sign 596526 cycles 597631 cycles 1.00
ML-DSA-65 verify 177530 cycles 177767 cycles 1.00
ML-DSA-87 keypair 286614 cycles 286470 cycles 1.00
ML-DSA-87 sign 758280 cycles 759318 cycles 1.00
ML-DSA-87 verify 295698 cycles 296310 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 133444 cycles 133190 cycles 1.00
ML-DSA-44 sign 520204 cycles 518479 cycles 1.00
ML-DSA-44 verify 146679 cycles 146665 cycles 1.00
ML-DSA-65 keypair 223925 cycles 224540 cycles 1.00
ML-DSA-65 sign 847147 cycles 844646 cycles 1.00
ML-DSA-65 verify 234409 cycles 234394 cycles 1.00
ML-DSA-87 keypair 367457 cycles 367714 cycles 1.00
ML-DSA-87 sign 1062489 cycles 1061302 cycles 1.00
ML-DSA-87 verify 380562 cycles 381497 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 43230 cycles 43292 cycles 1.00
ML-DSA-44 sign 130048 cycles 130069 cycles 1.00
ML-DSA-44 verify 45248 cycles 45143 cycles 1.00
ML-DSA-65 keypair 75909 cycles 75764 cycles 1.00
ML-DSA-65 sign 213758 cycles 213624 cycles 1.00
ML-DSA-65 verify 74418 cycles 74335 cycles 1.00
ML-DSA-87 keypair 123062 cycles 122992 cycles 1.00
ML-DSA-87 sign 271631 cycles 271188 cycles 1.00
ML-DSA-87 verify 120766 cycles 120668 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 61412 cycles 61655 cycles 1.00
ML-DSA-44 sign 189046 cycles 188845 cycles 1.00
ML-DSA-44 verify 66231 cycles 66290 cycles 1.00
ML-DSA-65 keypair 109628 cycles 112698 cycles 0.97
ML-DSA-65 sign 312172 cycles 318478 cycles 0.98
ML-DSA-65 verify 109663 cycles 111904 cycles 0.98
ML-DSA-87 keypair 170175 cycles 171213 cycles 0.99
ML-DSA-87 sign 379730 cycles 379008 cycles 1.00
ML-DSA-87 verify 170594 cycles 170855 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 118667 cycles 118359 cycles 1.00
ML-DSA-44 sign 457628 cycles 458607 cycles 1.00
ML-DSA-44 verify 131501 cycles 129909 cycles 1.01
ML-DSA-65 keypair 201384 cycles 200886 cycles 1.00
ML-DSA-65 sign 740800 cycles 744497 cycles 1.00
ML-DSA-65 verify 208939 cycles 208873 cycles 1.00
ML-DSA-87 keypair 331948 cycles 331103 cycles 1.00
ML-DSA-87 sign 936713 cycles 938146 cycles 1.00
ML-DSA-87 verify 343584 cycles 345239 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 67496 cycles 67139 cycles 1.01
ML-DSA-44 sign 198509 cycles 198230 cycles 1.00
ML-DSA-44 verify 70148 cycles 70284 cycles 1.00
ML-DSA-65 keypair 119095 cycles 119469 cycles 1.00
ML-DSA-65 sign 325504 cycles 326490 cycles 1.00
ML-DSA-65 verify 116584 cycles 116939 cycles 1.00
ML-DSA-87 keypair 196377 cycles 196410 cycles 1.00
ML-DSA-87 sign 421551 cycles 421431 cycles 1.00
ML-DSA-87 verify 193276 cycles 192954 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 92216 cycles 91733 cycles 1.01
ML-DSA-44 sign 349576 cycles 351109 cycles 1.00
ML-DSA-44 verify 99758 cycles 99446 cycles 1.00
ML-DSA-65 keypair 154324 cycles 153920 cycles 1.00
ML-DSA-65 sign 567577 cycles 571127 cycles 0.99
ML-DSA-65 verify 159730 cycles 160330 cycles 1.00
ML-DSA-87 keypair 255364 cycles 255096 cycles 1.00
ML-DSA-87 sign 725239 cycles 721477 cycles 1.01
ML-DSA-87 verify 263819 cycles 264366 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 154136 cycles 154277 cycles 1.00
ML-DSA-44 sign 588309 cycles 588660 cycles 1.00
ML-DSA-44 verify 168888 cycles 169423 cycles 1.00
ML-DSA-65 keypair 262317 cycles 267849 cycles 0.98
ML-DSA-65 sign 958311 cycles 979945 cycles 0.98
ML-DSA-65 verify 271809 cycles 277179 cycles 0.98
ML-DSA-87 keypair 431619 cycles 432934 cycles 1.00
ML-DSA-87 sign 1211290 cycles 1215004 cycles 1.00
ML-DSA-87 verify 446915 cycles 447425 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 211677 cycles 211596 cycles 1.00
ML-DSA-44 sign 761225 cycles 759367 cycles 1.00
ML-DSA-44 verify 228934 cycles 229045 cycles 1.00
ML-DSA-65 keypair 375171 cycles 375197 cycles 1.00
ML-DSA-65 sign 1245358 cycles 1247878 cycles 1.00
ML-DSA-65 verify 371103 cycles 371441 cycles 1.00
ML-DSA-87 keypair 600704 cycles 600145 cycles 1.00
ML-DSA-87 sign 1585805 cycles 1584430 cycles 1.00
ML-DSA-87 verify 616453 cycles 615831 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 128011 cycles 127916 cycles 1.00
ML-DSA-44 sign 441684 cycles 441335 cycles 1.00
ML-DSA-44 verify 136368 cycles 136365 cycles 1.00
ML-DSA-65 keypair 223088 cycles 221710 cycles 1.01
ML-DSA-65 sign 714012 cycles 714085 cycles 1.00
ML-DSA-65 verify 220581 cycles 220606 cycles 1.00
ML-DSA-87 keypair 364536 cycles 365306 cycles 1.00
ML-DSA-87 sign 914776 cycles 916248 cycles 1.00
ML-DSA-87 verify 371022 cycles 370929 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 71530 cycles 71457 cycles 1.00
ML-DSA-44 sign 209907 cycles 208907 cycles 1.00
ML-DSA-44 verify 74817 cycles 75005 cycles 1.00
ML-DSA-65 keypair 125984 cycles 125992 cycles 1.00
ML-DSA-65 sign 344118 cycles 345143 cycles 1.00
ML-DSA-65 verify 124046 cycles 124143 cycles 1.00
ML-DSA-87 keypair 206428 cycles 206682 cycles 1.00
ML-DSA-87 sign 439635 cycles 444029 cycles 0.99
ML-DSA-87 verify 204505 cycles 204086 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 138537 cycles 138455 cycles 1.00
ML-DSA-44 sign 486610 cycles 486114 cycles 1.00
ML-DSA-44 verify 149235 cycles 149248 cycles 1.00
ML-DSA-65 keypair 241159 cycles 242184 cycles 1.00
ML-DSA-65 sign 790421 cycles 791620 cycles 1.00
ML-DSA-65 verify 241516 cycles 241529 cycles 1.00
ML-DSA-87 keypair 395244 cycles 396160 cycles 1.00
ML-DSA-87 sign 1011999 cycles 1013640 cycles 1.00
ML-DSA-87 verify 404107 cycles 403824 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112444 cycles 112659 cycles 1.00
ML-DSA-44 sign 354470 cycles 354364 cycles 1.00
ML-DSA-44 verify 117205 cycles 117531 cycles 1.00
ML-DSA-65 keypair 195775 cycles 195098 cycles 1.00
ML-DSA-65 sign 587160 cycles 584870 cycles 1.00
ML-DSA-65 verify 194231 cycles 193581 cycles 1.00
ML-DSA-87 keypair 320708 cycles 321321 cycles 1.00
ML-DSA-87 sign 747632 cycles 747338 cycles 1.00
ML-DSA-87 verify 318503 cycles 318952 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 212383 cycles 212400 cycles 1.00
ML-DSA-44 sign 761080 cycles 760457 cycles 1.00
ML-DSA-44 verify 229961 cycles 229630 cycles 1.00
ML-DSA-65 keypair 376776 cycles 375851 cycles 1.00
ML-DSA-65 sign 1247361 cycles 1248638 cycles 1.00
ML-DSA-65 verify 372562 cycles 371900 cycles 1.00
ML-DSA-87 keypair 602074 cycles 602038 cycles 1.00
ML-DSA-87 sign 1586617 cycles 1591124 cycles 1.00
ML-DSA-87 verify 618303 cycles 618524 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 218089 cycles 213397 cycles 1.02
ML-DSA-44 sign 599197 cycles 587414 cycles 1.02
ML-DSA-44 verify 217607 cycles 213073 cycles 1.02
ML-DSA-65 keypair 379744 cycles 376687 cycles 1.01
ML-DSA-65 sign 986692 cycles 979172 cycles 1.01
ML-DSA-65 verify 364473 cycles 363486 cycles 1.00
ML-DSA-87 keypair 639246 cycles 647130 cycles 0.99
ML-DSA-87 sign 1311852 cycles 1349859 cycles 0.97
ML-DSA-87 verify 619801 cycles 632424 cycles 0.98

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 300220 cycles 297307 cycles 1.01
ML-DSA-44 sign 1133706 cycles 1124831 cycles 1.01
ML-DSA-44 verify 331808 cycles 327013 cycles 1.01
ML-DSA-65 keypair 549132 cycles 547240 cycles 1.00
ML-DSA-65 sign 1869739 cycles 1861858 cycles 1.00
ML-DSA-65 verify 530042 cycles 527234 cycles 1.01
ML-DSA-87 keypair 849544 cycles 841613 cycles 1.01
ML-DSA-87 sign 2351348 cycles 2350913 cycles 1.00
ML-DSA-87 verify 880873 cycles 873805 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 273189 cycles 271602 cycles 1.01
ML-DSA-44 sign 813853 cycles 811152 cycles 1.00
ML-DSA-44 verify 273807 cycles 273547 cycles 1.00
ML-DSA-65 keypair 465037 cycles 466624 cycles 1.00
ML-DSA-65 sign 1332507 cycles 1341019 cycles 0.99
ML-DSA-65 verify 453524 cycles 452548 cycles 1.00
ML-DSA-87 keypair 799137 cycles 798726 cycles 1.00
ML-DSA-87 sign 1842572 cycles 1835055 cycles 1.00
ML-DSA-87 verify 787003 cycles 773446 cycles 1.02

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 466035 cycles 465745 cycles 1.00
ML-DSA-44 sign 2138837 cycles 2140830 cycles 1.00
ML-DSA-44 verify 556780 cycles 558445 cycles 1.00
ML-DSA-65 keypair 783328 cycles 784167 cycles 1.00
ML-DSA-65 sign 3498055 cycles 3497460 cycles 1.00
ML-DSA-65 verify 868982 cycles 867417 cycles 1.00
ML-DSA-87 keypair 1266853 cycles 1272563 cycles 1.00
ML-DSA-87 sign 4330173 cycles 4331369 cycles 1.00
ML-DSA-87 verify 1392739 cycles 1392417 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SpacemiT K1 8 (Banana Pi F3) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 757718 cycles 760195 cycles 1.00
ML-DSA-44 sign 3137744 cycles 3141418 cycles 1.00
ML-DSA-44 verify 857097 cycles 859027 cycles 1.00
ML-DSA-65 keypair 1288623 cycles 1288559 cycles 1.00
ML-DSA-65 sign 5092434 cycles 5082523 cycles 1.00
ML-DSA-65 verify 1368044 cycles 1368131 cycles 1.00
ML-DSA-87 keypair 2112748 cycles 2109622 cycles 1.00
ML-DSA-87 sign 6377925 cycles 6359264 cycles 1.00
ML-DSA-87 verify 2226700 cycles 2225228 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 31s 12s +158%
mld_attempt_signature_generation ⚠️ 171s 34s +403%
sign_keypair_internal ⚠️ 52s 6s +767%
sign_pk_from_sk ⚠️ 66s 6s +1000%
sign_verify_internal ⚠️ 215s 47s +357%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1414s 1436s -1.5%
sign_verify_internal ⚠️ 215s 47s +357%
mld_attempt_signature_generation ⚠️ 171s 34s +403%
polyvec_matrix_pointwise_montgomery_yvec 128s 196s -35%
poly_pointwise_montgomery_c 72s 125s -42%
sign_pk_from_sk ⚠️ 66s 6s +1000%
mld_invntt_layer 60s 111s -46%
sign_keypair_internal ⚠️ 52s 6s +767%
compute_pack_t0_t1 ⚠️ 31s 12s +158%
fqmul 25s 39s -36%
mld_ntt_layer 24s 44s -45%
sign_signature_internal 14s 3s +367%
keccakf1600x4_permute_native 13s 25s -48%
sig_unpack_hints 13s 4s +225%
mld_ntt_butterfly_block 11s 23s -52%
poly_ntt_c 10s 19s -47%
rej_uniform_c 10s 18s -44%
polyt0_unpack 8s 14s -43%
poly_chknorm_c 7s 13s -46%
poly_invntt_tomont_c 7s 9s -22%
rej_uniform_native_x86_64 7s - new
sign_signature_pre_hash_shake256 7s 3s +133%
intt_native_aarch64 6s 4s +50%
mld_compute_pack_z 6s 6s +0%
pointwise_acc_native_aarch64 6s 5s +20%
polyeta_unpack 6s 13s -54%
decompose 5s 2s +150%
fqscale 5s 1s +400%
keccak_absorb_once_x4 5s 8s -38%
mld_check_pct 5s 15s -67%
mld_sample_s1_s2 5s 8s -38%
mld_sample_s1_s2_serial 5s 6s -17%
poly_uniform_4x 5s 3s +67%
poly_uniform_eta_4x 5s 12s -58%
polyvec_matrix_pointwise_montgomery_row 5s 13s -62%
polyveck_caddq 5s 8s -38%
polyveck_decompose_pack_w1 5s - new
polyw1_unpack_88 5s - new
yvec_init 5s 4s +25%
keccak_squeezeblocks_x4 4s 4s +0%
keccakf1600x4_extract_bytes 4s 1s +300%
mld_sign_resume 4s - new
pointwise_acc_native_x86_64 4s 8s -50%
poly_chknorm_native_x86_64 4s 2s +100%
poly_decompose_c 4s 6s -33%
poly_invntt_tomont_native 4s 2s +100%
poly_power2round 4s 7s -43%
poly_use_hint 4s 4s +0%
polyt1_pack 4s 5s -20%
polyveck_reduce 4s 6s -33%
polyvecl_ntt 4s 8s -50%
rej_uniform 4s 7s -43%
rej_uniform_native 4s 5s -20%
shake128_finalize 4s 2s +100%
sign_keypair 4s 4s +0%
keccak_absorb 3s 4s -25%
keccak_finalize 3s 1s +200%
keccak_init 3s 1s +200%
mld_ct_cmask_nonzero_u32 3s 5s -40%
mld_ct_get_optblocker_u8 3s 3s +0%
mld_ct_memcmp 3s 1s +200%
mld_ct_sel_int32 3s 2s +50%
mld_h 3s 2s +50%
mld_keccakf1600_permute_c 3s 8s -62%
mld_prepare_domain_separation_prefix 3s 3s +0%
mld_sign_finish 3s - new
nttunpack_native_x86_64 3s 3s +0%
pack_sig_z 3s 3s +0%
pack_sk_rho_key_tr_s2 3s 2s +50%
pointwise_native_x86_64 3s 5s -40%
poly_caddq 3s 5s -40%
poly_caddq_native 3s 3s +0%
poly_challenge 3s 4s -25%
poly_chknorm_native 3s 1s +200%
poly_decompose_88_native_aarch64 3s 2s +50%
poly_permute_bitrev_to_custom_optional 3s 2s +50%
poly_permute_bitrev_to_custom_optional_native 3s 3s +0%
poly_shiftl 3s 4s -25%
poly_uniform_eta 3s 5s -40%
poly_uniform_gamma1_4x 3s 3s +0%
poly_use_hint_native_aarch64 3s 3s +0%
poly_use_hint_native_x86_64 3s - new
polyt1_unpack 3s 3s +0%
polyveck_ntt 3s 3s +0%
polyveck_pack_eta 3s 5s -40%
polyvecl_pack_eta 3s 2s +50%
polyvecl_pointwise_acc_montgomery 3s 2s +50%
polyw1_pack_32 3s 3s +0%
polyw1_pack_88 3s 2s +50%
polyw1_unpack_32 3s - new
polyz_unpack_17_native_aarch64 3s 4s -25%
polyz_unpack_c 3s 7s -57%
polyz_unpack_native 3s 1s +200%
polyz_unpack_native_x86_64 3s 3s +0%
rej_uniform_eta_native_aarch64 3s 3s +0%
rej_uniform_eta_native_x86_64 3s - new
sign_signature_pre_hash_internal 3s 2s +50%
sign_verify_pre_hash_shake256 3s 7s -57%
sk_s2hat_get_poly 3s 2s +50%
sk_t0hat_get_poly 3s 3s +0%
unpack_sk 3s 4s -25%
caddq 2s 3s -33%
keccak_squeeze 2s 5s -60%
keccakf1600_permute 2s 2s +0%
keccakf1600_xor_bytes 2s 1s +100%
keccakf1600_xor_bytes (big endian) 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 4s -50%
keccakf1600x4_permute 2s 2s +0%
keccakf1600x4_xor_bytes 2s 1s +100%
mld_ct_abs_i32 2s 1s +100%
mld_ct_cmask_neg_i32 2s 2s +0%
mld_ct_cmask_nonzero_u8 2s 2s +0%
mld_ct_get_optblocker_i64 2s 2s +0%
mld_keccakf1600_extract_bytes 2s 1s +100%
mld_keccakf1600x4_extract_bytes_c 2s 3s -33%
mld_polymat_expand_entry 2s 4s -50%
mld_value_barrier_u8 2s 2s +0%
montgomery_reduce 2s 2s +0%
ntt_native_aarch64 2s 4s -50%
ntt_native_x86_64 2s 2s +0%
pack_sk_s1 2s 2s +0%
poly_add 2s 8s -75%
poly_caddq_c 2s 3s -33%
poly_caddq_native_aarch64 2s 3s -33%
poly_decompose 2s 2s +0%
poly_decompose_native 2s 2s +0%
poly_decompose_native_x86_64 2s 3s -33%
poly_ntt 2s 2s +0%
poly_reduce 2s 4s -50%
poly_sub 2s 5s -60%
poly_uniform 2s 4s -50%
poly_uniform_gamma1 2s 3s -33%
polyvec_matrix_expand 2s 3s -33%
polyveck_chknorm 2s 9s -78%
polyveck_invntt_tomont 2s 4s -50%
polyvecl_pointwise_acc_montgomery_native 2s 2s +0%
polyvecl_uniform_gamma1 2s 3s -33%
polyvecl_unpack_eta 2s 3s -33%
polyw1_pack 2s 2s +0%
polyz_pack 2s 4s -50%
polyz_unpack 2s 3s -33%
power2round 2s 3s -33%
rej_eta_c 2s 4s -50%
rej_eta_native 2s 4s -50%
shake128_absorb 2s 2s +0%
shake128_init 2s 2s +0%
shake128x4_absorb_once 2s 4s -50%
shake256 2s 3s -33%
shake256_release 2s 5s -60%
shake256x4_absorb_once 2s 5s -60%
shake256x4_squeezeblocks 2s 4s -50%
sign_signature 2s 4s -50%
sign_signature_extmu 2s 4s -50%
sign_verify 2s 2s +0%
sign_verify_pre_hash_internal 2s 3s -33%
sk_s1hat_get_poly 2s 3s -33%
unpack_pk_t1 2s 2s +0%
unpack_sk_s1hat 2s 3s -33%
unpack_sk_t0hat 2s 4s -50%
use_hint 2s 3s -33%
yvec_get_poly 2s 2s +0%
intt_native_x86_64 1s 4s -75%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_f1600_x1_native_aarch64_v84a 1s 2s -50%
keccak_f1600_x4_native_aarch64_v84a 1s 1s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 1s 2s -50%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 2s -50%
keccak_f1600_x4_native_avx2 1s 3s -67%
keccakf1600_extract_bytes (big endian) 1s 3s -67%
keccakf1600_permute_native 1s 3s -67%
keccakf1600x4_xor_bytes_native 1s 2s -50%
make_hint 1s 2s -50%
mld_ct_get_optblocker_u32 1s 2s -50%
mld_keccakf1600x4_xor_bytes_c 1s 1s +0%
mld_sign_attempt 1s - new
mld_value_barrier_i64 1s 2s -50%
mld_value_barrier_u32 1s 3s -67%
pack_sig_c 1s 2s -50%
pack_sig_h 1s 3s -67%
pointwise_native_aarch64 1s 4s -75%
poly_caddq_native_x86_64 1s 3s -67%
poly_chknorm 1s 4s -75%
poly_chknorm_native_aarch64 1s 5s -80%
poly_decompose_32_native_aarch64 1s 3s -67%
poly_invntt_tomont 1s 4s -75%
poly_ntt_native 1s 4s -75%
poly_pointwise_montgomery 1s 3s -67%
poly_pointwise_montgomery_native 1s 3s -67%
poly_use_hint_c 1s 4s -75%
poly_use_hint_native 1s 1s +0%
polyeta_pack 1s 3s -67%
polyt0_pack 1s 3s -67%
polyvec_matrix_expand_serial 1s 3s -67%
polyveck_unpack_eta 1s 3s -67%
polyvecl_chknorm 1s 38s -97%
polyvecl_pointwise_acc_montgomery_c 1s 2s -50%
polyvecl_uniform_gamma1_serial 1s 2s -50%
polyvecl_unpack_z 1s 3s -67%
polyw1_unpack 1s - new
polyz_unpack_19_native_aarch64 1s 5s -80%
reduce32 1s 3s -67%
rej_eta 1s 2s -50%
rej_uniform_native_aarch64 1s 3s -67%
shake128_release 1s 3s -67%
shake128_squeeze 1s 1s +0%
shake128x4_squeezeblocks 1s 1s +0%
shake256_absorb 1s 3s -67%
shake256_finalize 1s 3s -67%
shake256_init 1s 3s -67%
shake256_squeeze 1s 2s -50%
sign_verify_extmu 1s 4s -75%
sys_check_capability 1s 2s -50%
unpack_sk_s2hat 1s 3s -67%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-65, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 32s 7s +357%
mld_attempt_signature_generation ⚠️ 117s 31s +277%
sign_keypair_internal ⚠️ 25s 3s +733%
sign_pk_from_sk ⚠️ 49s 5s +880%
sign_verify_internal ⚠️ 155s 71s +118%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1174s 1482s -20.8%
sign_verify_internal ⚠️ 155s 71s +118%
mld_attempt_signature_generation ⚠️ 117s 31s +277%
polyvec_matrix_pointwise_montgomery_yvec 79s 201s -61%
poly_pointwise_montgomery_c 64s 116s -45%
mld_invntt_layer 58s 108s -46%
sign_pk_from_sk ⚠️ 49s 5s +880%
compute_pack_t0_t1 ⚠️ 32s 7s +357%
sign_keypair_internal ⚠️ 25s 3s +733%
fqmul 23s 40s -43%
mld_ntt_layer 22s 43s -49%
keccakf1600x4_permute_native 12s 22s -45%
sign_signature_internal 12s 6s +100%
mld_ntt_butterfly_block 11s 25s -56%
sig_unpack_hints 10s 1s +900%
poly_ntt_c 9s 22s -59%
polyveck_decompose_pack_w1 8s - new
rej_uniform_c 8s 16s -50%
poly_uniform_eta_4x 7s 13s -46%
polyveck_chknorm 7s 36s -81%
mld_check_pct 6s 12s -50%
poly_chknorm_c 6s 13s -54%
polyt0_unpack 6s 13s -54%
rej_uniform_native_x86_64 6s - new
decompose 5s 4s +25%
poly_chknorm_native_aarch64 5s 2s +150%
poly_decompose_c 5s 8s -38%
poly_invntt_tomont_c 5s 10s -50%
polyveck_reduce 5s 6s -17%
polyz_unpack_c 5s 9s -44%
sign_signature_pre_hash_internal 5s 3s +67%
unpack_sk 5s 3s +67%
caddq 4s 2s +100%
keccak_absorb_once_x4 4s 10s -60%
mld_compute_pack_z 4s 5s -20%
mld_sample_s1_s2_serial 4s 3s +33%
mld_sign_attempt 4s - new
poly_add 4s 7s -43%
poly_challenge 4s 5s -20%
poly_decompose_88_native_aarch64 4s 2s +100%
poly_power2round 4s 7s -43%
poly_uniform_eta 4s 5s -20%
polyvec_matrix_expand 4s 6s -33%
polyvecl_chknorm 4s 43s -91%
polyvecl_ntt 4s 7s -43%
reduce32 4s 2s +100%
rej_uniform 4s 7s -43%
shake256_init 4s 4s +0%
shake256x4_absorb_once 4s 2s +100%
sign_keypair 4s 4s +0%
sign_signature_pre_hash_shake256 4s 7s -43%
sign_verify_extmu 4s 2s +100%
sign_verify_pre_hash_shake256 4s 6s -33%
sys_check_capability 4s 1s +300%
keccak_f1600_x1_native_aarch64 3s 2s +50%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 3s 3s +0%
keccakf1600_extract_bytes (big endian) 3s 3s +0%
keccakf1600_permute 3s 4s -25%
keccakf1600x4_extract_bytes 3s 4s -25%
mld_ct_cmask_neg_i32 3s 3s +0%
mld_ct_cmask_nonzero_u8 3s 2s +50%
mld_ct_get_optblocker_u32 3s 3s +0%
mld_ct_get_optblocker_u8 3s 4s -25%
mld_keccakf1600_permute_c 3s 7s -57%
mld_sign_finish 3s - new
montgomery_reduce 3s 1s +200%
nttunpack_native_x86_64 3s 3s +0%
pack_sig_z 3s 1s +200%
pointwise_acc_native_x86_64 3s 5s -40%
pointwise_native_x86_64 3s 2s +50%
poly_caddq_c 3s 5s -40%
poly_caddq_native 3s 4s -25%
poly_caddq_native_aarch64 3s 2s +50%
poly_chknorm_native 3s 4s -25%
poly_ntt 3s 2s +50%
poly_permute_bitrev_to_custom_optional 3s 3s +0%
poly_use_hint_c 3s 2s +50%
polyeta_pack 3s 1s +200%
polyvec_matrix_pointwise_montgomery_row 3s 8s -62%
polyveck_caddq 3s 7s -57%
polyvecl_pointwise_acc_montgomery_native 3s 3s +0%
polyvecl_uniform_gamma1 3s 4s -25%
polyvecl_unpack_eta 3s 2s +50%
polyw1_pack 3s 3s +0%
polyw1_pack_32 3s 2s +50%
polyw1_unpack_32 3s - new
polyw1_unpack_88 3s - new
polyz_unpack_native 3s 3s +0%
rej_uniform_native 3s 4s -25%
shake128_init 3s 6s -50%
shake128x4_squeezeblocks 3s 2s +50%
shake256 3s 2s +50%
shake256_absorb 3s 4s -25%
sign_signature 3s 5s -40%
sign_verify_pre_hash_internal 3s 5s -40%
unpack_pk_t1 3s 2s +50%
unpack_sk_t0hat 3s 4s -25%
fqscale 2s 3s -33%
keccak_absorb 2s 3s -33%
keccak_f1600_x1_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v84a 2s 2s +0%
keccak_squeeze 2s 2s +0%
keccakf1600_permute_native 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 1s +100%
keccakf1600x4_permute 2s 1s +100%
keccakf1600x4_xor_bytes 2s 2s +0%
make_hint 2s 3s -33%
mld_ct_cmask_nonzero_u32 2s 3s -33%
mld_ct_sel_int32 2s 2s +0%
mld_h 2s 3s -33%
mld_keccakf1600_extract_bytes 2s 2s +0%
mld_keccakf1600x4_extract_bytes_c 2s 2s +0%
mld_keccakf1600x4_xor_bytes_c 2s 2s +0%
mld_sample_s1_s2 2s 6s -67%
mld_value_barrier_i64 2s 2s +0%
mld_value_barrier_u32 2s 4s -50%
pack_sk_rho_key_tr_s2 2s 2s +0%
pack_sk_s1 2s 3s -33%
pointwise_acc_native_aarch64 2s 6s -67%
pointwise_native_aarch64 2s 3s -33%
poly_caddq 2s 3s -33%
poly_caddq_native_x86_64 2s 1s +100%
poly_chknorm 2s 3s -33%
poly_chknorm_native_x86_64 2s 4s -50%
poly_decompose 2s 2s +0%
poly_decompose_native 2s 5s -60%
poly_decompose_native_x86_64 2s 2s +0%
poly_ntt_native 2s 2s +0%
poly_pointwise_montgomery 2s 2s +0%
poly_pointwise_montgomery_native 2s 4s -50%
poly_reduce 2s 4s -50%
poly_shiftl 2s 3s -33%
poly_sub 2s 4s -50%
poly_uniform_4x 2s 4s -50%
poly_uniform_gamma1_4x 2s 3s -33%
poly_use_hint_native 2s 7s -71%
poly_use_hint_native_aarch64 2s 4s -50%
poly_use_hint_native_x86_64 2s - new
polyeta_unpack 2s 4s -50%
polyt0_pack 2s 5s -60%
polyt1_pack 2s 3s -33%
polyt1_unpack 2s 3s -33%
polyveck_invntt_tomont 2s 7s -71%
polyveck_pack_eta 2s 3s -33%
polyvecl_pack_eta 2s 2s +0%
polyvecl_pointwise_acc_montgomery_c 2s 3s -33%
polyvecl_uniform_gamma1_serial 2s 1s +100%
polyvecl_unpack_z 2s 1s +100%
polyw1_unpack 2s - new
polyz_pack 2s 3s -33%
polyz_unpack_native_x86_64 2s 3s -33%
power2round 2s 2s +0%
rej_eta_native 2s 3s -33%
rej_uniform_eta_native_x86_64 2s - new
rej_uniform_native_aarch64 2s 5s -60%
shake128_absorb 2s 2s +0%
shake128_squeeze 2s 3s -33%
shake256_finalize 2s 1s +100%
shake256x4_squeezeblocks 2s 2s +0%
sign_signature_extmu 2s 4s -50%
sign_verify 2s 6s -67%
sk_s2hat_get_poly 2s 1s +100%
sk_t0hat_get_poly 2s 1s +100%
unpack_sk_s2hat 2s 3s -33%
use_hint 2s 2s +0%
yvec_init 2s 5s -60%
intt_native_aarch64 1s 2s -50%
intt_native_x86_64 1s 4s -75%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 1s 3s -67%
keccak_f1600_x4_native_avx2 1s 2s -50%
keccak_finalize 1s 2s -50%
keccak_init 1s 4s -75%
keccak_squeezeblocks_x4 1s 4s -75%
keccakf1600_xor_bytes 1s 4s -75%
keccakf1600_xor_bytes (big endian) 1s 2s -50%
keccakf1600x4_xor_bytes_native 1s 2s -50%
mld_ct_abs_i32 1s 6s -83%
mld_ct_get_optblocker_i64 1s 1s +0%
mld_ct_memcmp 1s 4s -75%
mld_polymat_expand_entry 1s 3s -67%
mld_prepare_domain_separation_prefix 1s 4s -75%
mld_sign_resume 1s - new
mld_value_barrier_u8 1s 1s +0%
ntt_native_aarch64 1s 3s -67%
ntt_native_x86_64 1s 5s -80%
pack_sig_c 1s 3s -67%
pack_sig_h 1s 2s -50%
poly_decompose_32_native_aarch64 1s 4s -75%
poly_invntt_tomont 1s 5s -80%
poly_invntt_tomont_native 1s 5s -80%
poly_permute_bitrev_to_custom_optional_native 1s 3s -67%
poly_uniform 1s 2s -50%
poly_uniform_gamma1 1s 3s -67%
poly_use_hint 1s 3s -67%
polyvec_matrix_expand_serial 1s 3s -67%
polyveck_ntt 1s 4s -75%
polyveck_unpack_eta 1s 3s -67%
polyvecl_pointwise_acc_montgomery 1s 3s -67%
polyw1_pack_88 1s 3s -67%
polyz_unpack 1s 3s -67%
polyz_unpack_17_native_aarch64 1s 2s -50%
polyz_unpack_19_native_aarch64 1s 5s -80%
rej_eta 1s 5s -80%
rej_eta_c 1s 4s -75%
rej_uniform_eta_native_aarch64 1s 5s -80%
shake128_finalize 1s 2s -50%
shake128_release 1s 2s -50%
shake128x4_absorb_once 1s 2s -50%
shake256_release 1s 3s -67%
shake256_squeeze 1s 4s -75%
sk_s1hat_get_poly 1s 3s -67%
unpack_sk_s1hat 1s 1s +0%
yvec_get_poly 1s 3s -67%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 188s 19s +889%
mld_attempt_signature_generation ⚠️ 255s 55s +364%
sig_unpack_hints ⚠️ 26s 3s +767%
sign_keypair_internal ⚠️ 39s 6s +550%
sign_pk_from_sk ⚠️ 65s 6s +983%
sign_signature_internal ⚠️ 321s 42s +664%
sign_verify_internal ⚠️ 521s 97s +437%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 2448s 2151s +13.8%
sign_verify_internal ⚠️ 521s 97s +437%
sign_signature_internal ⚠️ 321s 42s +664%
mld_attempt_signature_generation ⚠️ 255s 55s +364%
polyvecl_pointwise_acc_montgomery_c 213s 353s -40%
compute_pack_t0_t1 ⚠️ 188s 19s +889%
polyvec_matrix_expand 86s 324s -73%
sign_pk_from_sk ⚠️ 65s 6s +983%
mld_invntt_layer 58s 115s -50%
poly_pointwise_montgomery_c 44s 144s -69%
sign_keypair_internal ⚠️ 39s 6s +550%
sig_unpack_hints ⚠️ 26s 3s +767%
mld_ntt_layer 24s 45s -47%
fqmul 22s 43s -49%
polyvec_matrix_expand_serial 22s 38s -42%
keccakf1600x4_permute_native 14s 23s -39%
rej_uniform 12s 15s -20%
mld_ntt_butterfly_block 11s 24s -54%
polyvec_matrix_pointwise_montgomery_yvec 11s 18s -39%
polyt0_unpack 10s 14s -29%
rej_uniform_c 10s 19s -47%
poly_ntt_c 9s 21s -57%
poly_uniform_eta_4x 9s 11s -18%
poly_chknorm_c 8s 15s -47%
polyveck_decompose_pack_w1 8s - new
polyveck_caddq 7s 8s -12%
rej_uniform_native_x86_64 7s - new
poly_invntt_tomont_c 6s 12s -50%
poly_uniform_4x 6s 13s -54%
shake256_squeeze 6s 2s +200%
sign_keypair 6s 4s +50%
sign_signature_pre_hash_shake256 6s 6s +0%
mld_check_pct 5s 16s -69%
mld_compute_pack_z 5s 8s -38%
mld_keccakf1600_permute_c 5s 7s -29%
pointwise_acc_native_aarch64 5s 6s -17%
poly_decompose_native 5s 2s +150%
poly_reduce 5s 2s +150%
polyvec_matrix_pointwise_montgomery_row 5s 3s +67%
polyveck_invntt_tomont 5s 9s -44%
polyvecl_ntt 5s 7s -29%
shake256x4_absorb_once 5s 2s +150%
mld_ct_get_optblocker_u32 4s 2s +100%
mld_sign_resume 4s - new
ntt_native_aarch64 4s 6s -33%
pointwise_acc_native_x86_64 4s 6s -33%
poly_invntt_tomont_native 4s 5s -20%
poly_use_hint 4s 1s +300%
polyeta_unpack 4s 17s -76%
polyveck_ntt 4s 11s -64%
polyw1_unpack_88 4s - new
sk_s1hat_get_poly 4s 3s +33%
unpack_sk_t0hat 4s 7s -43%
decompose 3s 3s +0%
intt_native_aarch64 3s 3s +0%
keccak_absorb 3s 3s +0%
keccak_absorb_once_x4 3s 9s -67%
keccak_f1600_x4_native_aarch64_v84a 3s 4s -25%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccak_init 3s 3s +0%
keccak_squeeze 3s 2s +50%
keccakf1600_extract_bytes (big endian) 3s 3s +0%
keccakf1600_xor_bytes (big endian) 3s 2s +50%
keccakf1600x4_xor_bytes 3s 3s +0%
keccakf1600x4_xor_bytes_native 3s 3s +0%
mld_ct_abs_i32 3s 3s +0%
mld_ct_get_optblocker_i64 3s 3s +0%
mld_h 3s 3s +0%
mld_sample_s1_s2 3s 6s -50%
montgomery_reduce 3s 3s +0%
nttunpack_native_x86_64 3s 3s +0%
pack_sig_h 3s 4s -25%
pack_sig_z 3s 5s -40%
pointwise_native_aarch64 3s 5s -40%
pointwise_native_x86_64 3s 2s +50%
poly_add 3s 6s -50%
poly_caddq 3s 2s +50%
poly_chknorm 3s 3s +0%
poly_chknorm_native 3s 4s -25%
poly_decompose_c 3s 5s -40%
poly_ntt 3s 3s +0%
poly_ntt_native 3s 3s +0%
poly_pointwise_montgomery 3s 4s -25%
poly_power2round 3s 3s +0%
poly_sub 3s 3s +0%
poly_use_hint_c 3s 2s +50%
polyveck_pack_eta 3s 6s -50%
polyvecl_pointwise_acc_montgomery 3s 5s -40%
polyvecl_uniform_gamma1_serial 3s 4s -25%
polyz_unpack_c 3s 3s +0%
polyz_unpack_native_x86_64 3s 3s +0%
reduce32 3s 3s +0%
rej_eta_native 3s 6s -50%
rej_uniform_eta_native_x86_64 3s - new
rej_uniform_native 3s 8s -62%
sign_signature_extmu 3s 4s -25%
sk_t0hat_get_poly 3s 4s -25%
unpack_pk_t1 3s 1s +200%
use_hint 3s 3s +0%
yvec_init 3s 4s -25%
fqscale 2s 3s -33%
intt_native_x86_64 2s 3s -33%
keccak_finalize 2s 1s +100%
keccakf1600x4_extract_bytes 2s 2s +0%
make_hint 2s 3s -33%
mld_ct_cmask_neg_i32 2s 4s -50%
mld_ct_cmask_nonzero_u32 2s 2s +0%
mld_ct_cmask_nonzero_u8 2s 3s -33%
mld_ct_get_optblocker_u8 2s 3s -33%
mld_ct_memcmp 2s 3s -33%
mld_keccakf1600_extract_bytes 2s 4s -50%
mld_polymat_expand_entry 2s 3s -33%
mld_prepare_domain_separation_prefix 2s 4s -50%
mld_sample_s1_s2_serial 2s 8s -75%
mld_sign_attempt 2s - new
mld_sign_finish 2s - new
mld_value_barrier_i64 2s 1s +100%
mld_value_barrier_u32 2s 2s +0%
mld_value_barrier_u8 2s 3s -33%
ntt_native_x86_64 2s 2s +0%
pack_sk_s1 2s 1s +100%
poly_caddq_native_aarch64 2s 3s -33%
poly_caddq_native_x86_64 2s 2s +0%
poly_chknorm_native_aarch64 2s 4s -50%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose 2s 3s -33%
poly_decompose_88_native_aarch64 2s 2s +0%
poly_invntt_tomont 2s 1s +100%
poly_permute_bitrev_to_custom_optional_native 2s 5s -60%
poly_pointwise_montgomery_native 2s 5s -60%
poly_uniform 2s 4s -50%
poly_uniform_eta 2s 4s -50%
poly_uniform_gamma1 2s 4s -50%
poly_uniform_gamma1_4x 2s 3s -33%
poly_use_hint_native 2s 3s -33%
poly_use_hint_native_x86_64 2s - new
polyeta_pack 2s 4s -50%
polyt1_pack 2s 3s -33%
polyt1_unpack 2s 4s -50%
polyveck_unpack_eta 2s 4s -50%
polyvecl_chknorm 2s 7s -71%
polyvecl_pack_eta 2s 3s -33%
polyvecl_pointwise_acc_montgomery_native 2s 3s -33%
polyvecl_uniform_gamma1 2s 2s +0%
polyvecl_unpack_eta 2s 5s -60%
polyvecl_unpack_z 2s 4s -50%
polyw1_pack 2s 4s -50%
polyw1_pack_32 2s 4s -50%
polyw1_pack_88 2s 3s -33%
polyw1_unpack_32 2s - new
polyz_unpack 2s 2s +0%
polyz_unpack_17_native_aarch64 2s 3s -33%
polyz_unpack_native 2s 4s -50%
rej_eta_c 2s 5s -60%
rej_uniform_eta_native_aarch64 2s 5s -60%
shake128_absorb 2s 4s -50%
shake128_finalize 2s 2s +0%
shake128_release 2s 4s -50%
shake128_squeeze 2s 2s +0%
shake128x4_absorb_once 2s 5s -60%
shake128x4_squeezeblocks 2s 3s -33%
shake256 2s 2s +0%
shake256_absorb 2s 2s +0%
shake256_finalize 2s 2s +0%
shake256x4_squeezeblocks 2s 1s +100%
sign_signature_pre_hash_internal 2s 5s -60%
sign_verify_pre_hash_internal 2s 3s -33%
sign_verify_pre_hash_shake256 2s 3s -33%
sys_check_capability 2s 3s -33%
unpack_sk_s1hat 2s 2s +0%
caddq 1s 4s -75%
keccak_f1600_x1_native_aarch64 1s 1s +0%
keccak_f1600_x1_native_aarch64_v84a 1s 1s +0%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 4s -75%
keccak_f1600_x4_native_avx2 1s 2s -50%
keccak_squeezeblocks_x4 1s 4s -75%
keccakf1600_permute 1s 3s -67%
keccakf1600_permute_native 1s 4s -75%
keccakf1600_xor_bytes 1s 3s -67%
keccakf1600x4_extract_bytes_native 1s 2s -50%
keccakf1600x4_permute 1s 3s -67%
mld_ct_sel_int32 1s 3s -67%
mld_keccakf1600x4_extract_bytes_c 1s 2s -50%
mld_keccakf1600x4_xor_bytes_c 1s 4s -75%
pack_sig_c 1s 4s -75%
pack_sk_rho_key_tr_s2 1s 3s -67%
poly_caddq_c 1s 3s -67%
poly_caddq_native 1s 4s -75%
poly_challenge 1s 5s -80%
poly_decompose_32_native_aarch64 1s 5s -80%
poly_decompose_native_x86_64 1s 2s -50%
poly_permute_bitrev_to_custom_optional 1s 3s -67%
poly_shiftl 1s 4s -75%
poly_use_hint_native_aarch64 1s 2s -50%
polyt0_pack 1s 4s -75%
polyveck_chknorm 1s 3s -67%
polyveck_reduce 1s 3s -67%
polyw1_unpack 1s - new
polyz_pack 1s 5s -80%
polyz_unpack_19_native_aarch64 1s 4s -75%
power2round 1s 4s -75%
rej_eta 1s 4s -75%
rej_uniform_native_aarch64 1s 2s -50%
shake128_init 1s 1s +0%
shake256_init 1s 2s -50%
shake256_release 1s 3s -67%
sign_signature 1s 5s -80%
sign_verify 1s 5s -80%
sign_verify_extmu 1s 3s -67%
sk_s2hat_get_poly 1s 3s -67%
unpack_sk 1s 5s -80%
unpack_sk_s2hat 1s 3s -67%
yvec_get_poly 1s 4s -75%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-44)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 97s 50s +94%
mld_attempt_signature_generation ⚠️ 164s 58s +183%
polyveck_chknorm ⚠️ 56s 5s +1020%
sign_keypair_internal ⚠️ 21s 4s +425%
sign_pk_from_sk ⚠️ 30s 5s +500%
sign_signature_internal ⚠️ 91s 26s +250%
sign_verify_internal ⚠️ 263s 114s +131%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1605s 1514s +6.0%
sign_verify_internal ⚠️ 263s 114s +131%
mld_attempt_signature_generation ⚠️ 164s 58s +183%
polyvecl_pointwise_acc_montgomery_c 150s 131s +15%
compute_pack_t0_t1 ⚠️ 97s 50s +94%
sign_signature_internal ⚠️ 91s 26s +250%
mld_invntt_layer 58s 107s -46%
polyveck_chknorm ⚠️ 56s 5s +1020%
poly_pointwise_montgomery_c 44s 114s -61%
sign_pk_from_sk ⚠️ 30s 5s +500%
fqmul 23s 39s -41%
mld_ntt_layer 22s 42s -48%
sign_keypair_internal ⚠️ 21s 4s +425%
polyvec_matrix_expand 19s 28s -32%
sig_unpack_hints 19s 2s +850%
polyvec_matrix_pointwise_montgomery_yvec 13s 15s -13%
keccakf1600x4_permute_native 12s 22s -45%
mld_ntt_butterfly_block 11s 23s -52%
rej_uniform 11s 18s -39%
poly_chknorm_c 9s 17s -47%
poly_ntt_c 9s 19s -53%
polyt0_unpack 8s 15s -47%
rej_uniform_c 8s 15s -47%
rej_uniform_native_x86_64 8s - new
mld_compute_pack_z 7s 7s +0%
poly_invntt_tomont_c 7s 11s -36%
poly_uniform_eta_4x 7s 11s -36%
sign_keypair 7s 4s +75%
poly_uniform_4x 6s 14s -57%
polyeta_unpack 6s 14s -57%
polyvecl_ntt 6s 2s +200%
keccak_absorb_once_x4 5s 9s -44%
keccak_finalize 5s 3s +67%
polyw1_unpack_88 5s - new
polyz_unpack_c 5s 11s -55%
polyz_unpack_native_x86_64 5s 3s +67%
mld_check_pct 4s 14s -71%
mld_ct_cmask_nonzero_u32 4s 2s +100%
mld_h 4s 4s +0%
mld_keccakf1600_permute_c 4s 8s -50%
pointwise_acc_native_x86_64 4s 7s -43%
poly_challenge 4s 3s +33%
poly_chknorm_native_x86_64 4s 2s +100%
poly_decompose_native 4s 3s +33%
poly_decompose_native_x86_64 4s 3s +33%
poly_pointwise_montgomery_native 4s 2s +100%
polyvec_matrix_expand_serial 4s 8s -50%
polyz_unpack 4s 3s +33%
rej_uniform_native 4s 4s +0%
sign_signature_pre_hash_shake256 4s 3s +33%
sign_verify_extmu 4s 3s +33%
unpack_pk_t1 4s 5s -20%
yvec_init 4s 3s +33%
intt_native_aarch64 3s 9s -67%
intt_native_x86_64 3s 4s -25%
keccak_absorb 3s 5s -40%
keccak_f1600_x1_native_aarch64_v84a 3s 3s +0%
keccakf1600_xor_bytes 3s 1s +200%
keccakf1600_xor_bytes (big endian) 3s 4s -25%
keccakf1600x4_extract_bytes 3s 2s +50%
make_hint 3s 3s +0%
mld_prepare_domain_separation_prefix 3s 4s -25%
mld_sign_attempt 3s - new
mld_value_barrier_i64 3s 2s +50%
ntt_native_aarch64 3s 3s +0%
ntt_native_x86_64 3s 3s +0%
pack_sig_z 3s 4s -25%
pointwise_acc_native_aarch64 3s 4s -25%
poly_add 3s 7s -57%
poly_caddq_c 3s 2s +50%
poly_caddq_native 3s 5s -40%
poly_caddq_native_x86_64 3s 4s -25%
poly_chknorm_native 3s 4s -25%
poly_decompose_c 3s 4s -25%
poly_invntt_tomont_native 3s 4s -25%
poly_ntt_native 3s 5s -40%
poly_permute_bitrev_to_custom_optional 3s 2s +50%
poly_reduce 3s 3s +0%
poly_sub 3s 3s +0%
poly_uniform_eta 3s 5s -40%
poly_use_hint 3s 4s -25%
poly_use_hint_native 3s 4s -25%
polyt1_unpack 3s 5s -40%
polyveck_invntt_tomont 3s 5s -40%
polyveck_pack_eta 3s 2s +50%
polyveck_reduce 3s 5s -40%
polyvecl_pointwise_acc_montgomery_native 3s 2s +50%
polyw1_pack 3s 3s +0%
power2round 3s 3s +0%
rej_eta 3s 3s +0%
rej_eta_c 3s 4s -25%
rej_eta_native 3s 6s -50%
rej_uniform_eta_native_x86_64 3s - new
rej_uniform_native_aarch64 3s 5s -40%
shake256_init 3s 3s +0%
sign_signature 3s 3s +0%
sign_verify 3s 4s -25%
sign_verify_pre_hash_internal 3s 3s +0%
sign_verify_pre_hash_shake256 3s 5s -40%
sk_s1hat_get_poly 3s 3s +0%
unpack_sk_s2hat 3s 4s -25%
unpack_sk_t0hat 3s 3s +0%
decompose 2s 2s +0%
keccak_f1600_x1_native_aarch64 2s 2s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 2s 3s -33%
keccak_squeeze 2s 1s +100%
keccak_squeezeblocks_x4 2s 4s -50%
keccakf1600_permute 2s 2s +0%
keccakf1600x4_xor_bytes 2s 2s +0%
mld_ct_abs_i32 2s 2s +0%
mld_ct_cmask_nonzero_u8 2s 2s +0%
mld_ct_get_optblocker_u8 2s 1s +100%
mld_ct_memcmp 2s 3s -33%
mld_keccakf1600_extract_bytes 2s 2s +0%
mld_sample_s1_s2 2s 4s -50%
mld_sign_finish 2s - new
mld_sign_resume 2s - new
mld_value_barrier_u32 2s 2s +0%
mld_value_barrier_u8 2s 3s -33%
montgomery_reduce 2s 3s -33%
nttunpack_native_x86_64 2s 1s +100%
pack_sig_h 2s 3s -33%
pack_sk_rho_key_tr_s2 2s 3s -33%
pack_sk_s1 2s 2s +0%
pointwise_native_aarch64 2s 5s -60%
pointwise_native_x86_64 2s 4s -50%
poly_caddq_native_aarch64 2s 4s -50%
poly_chknorm 2s 4s -50%
poly_chknorm_native_aarch64 2s 4s -50%
poly_decompose_32_native_aarch64 2s 2s +0%
poly_decompose_88_native_aarch64 2s 4s -50%
poly_invntt_tomont 2s 4s -50%
poly_permute_bitrev_to_custom_optional_native 2s 4s -50%
poly_power2round 2s 4s -50%
poly_uniform 2s 6s -67%
poly_uniform_gamma1 2s 4s -50%
poly_use_hint_c 2s 3s -33%
poly_use_hint_native_aarch64 2s 2s +0%
poly_use_hint_native_x86_64 2s - new
polyt0_pack 2s 2s +0%
polyvec_matrix_pointwise_montgomery_row 2s 2s +0%
polyveck_caddq 2s 3s -33%
polyveck_decompose_pack_w1 2s - new
polyveck_ntt 2s 5s -60%
polyvecl_chknorm 2s 10s -80%
polyvecl_pack_eta 2s 3s -33%
polyvecl_unpack_eta 2s 2s +0%
polyw1_pack_88 2s 1s +100%
polyw1_unpack 2s - new
polyw1_unpack_32 2s - new
polyz_pack 2s 2s +0%
polyz_unpack_19_native_aarch64 2s 5s -60%
polyz_unpack_native 2s 2s +0%
rej_uniform_eta_native_aarch64 2s 2s +0%
shake128_finalize 2s 2s +0%
shake128_init 2s 3s -33%
shake256_finalize 2s 2s +0%
sign_signature_extmu 2s 4s -50%
sk_s2hat_get_poly 2s 2s +0%
sk_t0hat_get_poly 2s 2s +0%
unpack_sk 2s 3s -33%
unpack_sk_s1hat 2s 1s +100%
use_hint 2s 2s +0%
yvec_get_poly 2s 3s -33%
caddq 1s 2s -50%
fqscale 1s 3s -67%
keccak_f1600_x4_native_aarch64_v84a 1s 2s -50%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 1s +0%
keccak_f1600_x4_native_avx2 1s 3s -67%
keccak_init 1s 3s -67%
keccakf1600_extract_bytes (big endian) 1s 2s -50%
keccakf1600_permute_native 1s 3s -67%
keccakf1600x4_extract_bytes_native 1s 3s -67%
keccakf1600x4_permute 1s 4s -75%
keccakf1600x4_xor_bytes_native 1s 3s -67%
mld_ct_cmask_neg_i32 1s 3s -67%
mld_ct_get_optblocker_i64 1s 4s -75%
mld_ct_get_optblocker_u32 1s 1s +0%
mld_ct_sel_int32 1s 2s -50%
mld_keccakf1600x4_extract_bytes_c 1s 2s -50%
mld_keccakf1600x4_xor_bytes_c 1s 2s -50%
mld_polymat_expand_entry 1s 3s -67%
mld_sample_s1_s2_serial 1s 3s -67%
pack_sig_c 1s 2s -50%
poly_caddq 1s 2s -50%
poly_decompose 1s 3s -67%
poly_ntt 1s 3s -67%
poly_pointwise_montgomery 1s 3s -67%
poly_shiftl 1s 3s -67%
poly_uniform_gamma1_4x 1s 5s -80%
polyeta_pack 1s 3s -67%
polyt1_pack 1s 4s -75%
polyveck_unpack_eta 1s 2s -50%
polyvecl_pointwise_acc_montgomery 1s 4s -75%
polyvecl_uniform_gamma1 1s 3s -67%
polyvecl_uniform_gamma1_serial 1s 2s -50%
polyvecl_unpack_z 1s 2s -50%
polyw1_pack_32 1s 2s -50%
polyz_unpack_17_native_aarch64 1s 5s -80%
reduce32 1s 2s -50%
shake128_absorb 1s 2s -50%
shake128_release 1s 1s +0%
shake128_squeeze 1s 1s +0%
shake128x4_absorb_once 1s 3s -67%
shake128x4_squeezeblocks 1s 1s +0%
shake256 1s 1s +0%
shake256_absorb 1s 3s -67%
shake256_release 1s 2s -50%
shake256_squeeze 1s 2s -50%
shake256x4_absorb_once 1s 3s -67%
shake256x4_squeezeblocks 1s 4s -75%
sign_signature_pre_hash_internal 1s 6s -83%
sys_check_capability 1s 4s -75%

Comment thread mldsa/src/polyvec_lazy.h
Comment thread mldsa/src/poly.h Outdated
Comment thread mldsa/src/poly.h Outdated
Comment thread mldsa/src/polyvec_lazy.c
Comment thread mldsa/src/poly.c Outdated
Comment thread mldsa/src/poly.c Outdated
Comment thread mldsa/src/poly.c Outdated
Comment thread mldsa/src/poly.c Outdated
Comment thread mldsa/src/params.h Outdated
Comment thread mldsa/src/poly.c

@hanno-becker hanno-becker left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am unhappy with the complexity creep, but have so far not found any approach to avoiding it.

In the meantime, some notes:

  • Can you split the work into multiple commits? As indicated in the comments, you can at least hoist the yvec scratch refactor and the w1_unpack changes into standalone commits.
  • The use of MLDSA_POLYW1_PACKED_BITS_XX is inconsistent, and its presence inconsistent with other pack bit width, e.g. for t1.

@mkannwischer
mkannwischer force-pushed the lowram-sign-pack-w1 branch 2 times, most recently from 6ba0a18 to dad850e Compare September 1, 2026 07:33
Settle on the postfix form everywhere in mldsa/src.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
The scratch argument of mld_polyvec_matrix_pointwise_montgomery_yvec was
typed mld_polyvecl in both variants, even though the lazy variant only
ever uses its first polynomial.

This will enable memory savings in a following commit.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
Comment thread mldsa/src/params.h Outdated
Comment thread mldsa/src/polyvec.c
Comment thread mldsa/src/params.h Outdated

@hanno-becker hanno-becker left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @mkannwischer. This is a significant memory reduction. The added complexity is tolerable; as mentioned, I would prefer to hide the scratch object better, but I cannot see a way to do it at the moment.

Add mld_polyw1_unpack_{88,32} and the mld_polyw1_unpack dispatcher, the
inverses of mld_polyw1_pack_{88,32} and mld_polyw1_pack, together with
their CBMC proofs.

The routines have no caller yet - they will be used in the next commit to
enable some memory optimizations.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
The signing attempt held w1 as a full mld_polyveck, but w1 is only
needed as the w1Encode input to H and, coefficient-wise, as the a1
argument of MakeHint. Both work from the packed encoding, so Decompose
and w1Encode now run one polynomial at a time into a
MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers
each row with mld_polyw1_unpack.

Signing allocation in bytes:

              default           REDUCE_RAM
  ML-DSA-44   44704 -> 44448    13120 ->  9792
  ML-DSA-65   69312 -> 68032    17248 -> 11872
  ML-DSA-87  108224 -> 107200   21344 -> 14176

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
Follow MLDSA_POLYW1_PACKED_BITS for the remaining packing routines and
derive MLDSA_POLY*_PACKEDBYTES instead of hardcoding 320, 416, 576/640
and 96/128.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
@mkannwischer
mkannwischer added this pull request to the merge queue Sep 1, 2026
Merged via the queue into main with commit b05bb09 Sep 1, 2026
380 of 390 checks passed
@mkannwischer
mkannwischer deleted the lowram-sign-pack-w1 branch September 1, 2026 14:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants