Skip to content

Add Armv8.1-M Keccak x1 backend - #1277

Open
bremoran wants to merge 4 commits into
mainfrom
armv81m-keccak-x1
Open

Add Armv8.1-M Keccak x1 backend#1277
bremoran wants to merge 4 commits into
mainfrom
armv81m-keccak-x1

Conversation

@bremoran

@bremoran bremoran commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

Add the unmodified upstream Adomnicai Armv7-M Keccak source under armv81m_clean and the Slothy-generated M7 optimized x1 permutation under armv81m_opt, with a Makefile regeneration path.

Move the Armv8.1-M FIPS202 development sources to armv81m_opt and synchronize the production backend. Route x1 state XOR/extract through the clean-source assembly helpers so the optimized permutation can keep the state in the bit-interleaved native representation.

Keep the test changes focused on representation-aware Keccak x1/x4 unit coverage, including extracted-byte state dumps on x1 failures, and the static ML-DSA-87 unit-test workspace needed for Zephyr stack pressure.

Fixes #1322

@bremoran
bremoran requested a review from a team as a code owner July 10, 2026 08:11
@mkannwischer
mkannwischer marked this pull request as draft July 10, 2026 08:18
@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-44)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 98s 50s +96%
mld_attempt_signature_generation ⚠️ 311s 58s +436%
sig_unpack_hints ⚠️ 20s 2s +900%
sign_keypair_internal ⚠️ 20s 4s +400%
sign_pk_from_sk ⚠️ 37s 5s +640%
sign_signature_internal ⚠️ 99s 26s +281%
sign_verify_internal ⚠️ 294s 114s +158%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** 1839s 1514s +21.5%
mld_attempt_signature_generation ⚠️ 311s 58s +436%
sign_verify_internal ⚠️ 294s 114s +158%
polyvecl_pointwise_acc_montgomery_c 151s 131s +15%
poly_pointwise_montgomery_c 145s 114s +27%
sign_signature_internal ⚠️ 99s 26s +281%
compute_pack_t0_t1 ⚠️ 98s 50s +96%
mld_invntt_layer 58s 107s -46%
sign_pk_from_sk ⚠️ 37s 5s +640%
mld_ntt_layer 24s 42s -43%
fqmul 23s 39s -41%
sig_unpack_hints ⚠️ 20s 2s +900%
sign_keypair_internal ⚠️ 20s 4s +400%
polyvec_matrix_expand 18s 28s -36%
mld_ntt_butterfly_block 14s 23s -39%
poly_ntt_c 12s 19s -37%
keccakf1600x4_permute_native 11s 22s -50%
poly_uniform_eta_4x 9s 11s -18%
polyeta_unpack 9s 14s -36%
polyt0_unpack 8s 15s -47%
polyvec_matrix_pointwise_montgomery_yvec 8s 15s -47%
rej_uniform_native_x86_64 8s - new
poly_invntt_tomont_c 7s 11s -36%
rej_uniform 7s 18s -61%
mld_check_pct 6s 14s -57%
poly_chknorm_c 6s 17s -65%
polyz_unpack_c 6s 11s -45%
rej_uniform_c 6s 15s -60%
sign_verify_extmu 6s 3s +100%
keccak_absorb_once_x4 5s 9s -44%
keccak_squeezeblocks_x4 5s 4s +25%
poly_add 5s 7s -29%
poly_uniform_4x 5s 14s -64%
poly_use_hint_c 5s 3s +67%
polyvecl_ntt 5s 2s +150%
sign_keypair 5s 4s +25%
sign_signature_extmu 5s 4s +25%
keccak_absorb 4s 5s -20%
keccakf1600_xor_bytes 4s 1s +300%
mld_compute_pack_z 4s 7s -43%
mld_keccakf1600_permute_c 4s 8s -50%
mld_sample_s1_s2_serial 4s 3s +33%
ntt_native_x86_64 4s 3s +33%
poly_sub 4s 3s +33%
polyveck_pack_w1 4s 1s +300%
polyvecl_unpack_eta 4s 2s +100%
polyvecl_unpack_z 4s 2s +100%
rej_eta_c 4s 4s +0%
rej_uniform_eta_native_aarch64 4s 2s +100%
sign_signature_pre_hash_shake256 4s 3s +33%
sign_verify_pre_hash_shake256 4s 5s -20%
unpack_sk 4s 3s +33%
make_hint 3s 3s +0%
mld_ct_cmask_neg_i32 3s 3s +0%
mld_h 3s 4s -25%
mld_sign_finish 3s - new
ntt_native_aarch64 3s 3s +0%
pack_sig_h 3s 3s +0%
pack_sig_z 3s 4s -25%
pointwise_acc_native_aarch64 3s 4s -25%
pointwise_acc_native_x86_64 3s 7s -57%
pointwise_native_aarch64 3s 5s -40%
poly_caddq_native_x86_64 3s 4s -25%
poly_chknorm_native_x86_64 3s 2s +50%
poly_decompose_32_native_aarch64 3s 2s +50%
poly_decompose_88_native_aarch64 3s 4s -25%
poly_ntt 3s 3s +0%
poly_ntt_native 3s 5s -40%
poly_pointwise_montgomery 3s 3s +0%
poly_power2round 3s 4s -25%
poly_uniform 3s 6s -50%
poly_use_hint_native 3s 4s -25%
poly_use_hint_native_aarch64 3s 2s +50%
polyt0_pack 3s 2s +50%
polyt1_unpack 3s 5s -40%
polyvec_matrix_expand_serial 3s 8s -62%
polyveck_caddq 3s 3s +0%
polyveck_chknorm 3s 5s -40%
polyveck_ntt 3s 5s -40%
polyveck_pack_eta 3s 2s +50%
polyveck_unpack_eta 3s 2s +50%
polyvecl_pointwise_acc_montgomery 3s 4s -25%
polyvecl_uniform_gamma1_serial 3s 2s +50%
polyw1_pack_88 3s 1s +200%
polyz_unpack_17_native_aarch64 3s 5s -40%
polyz_unpack_19_native_aarch64 3s 5s -40%
polyz_unpack_native_x86_64 3s 3s +0%
reduce32 3s 2s +50%
rej_uniform_native 3s 4s -25%
shake128_absorb 3s 2s +50%
sign_signature 3s 3s +0%
sign_signature_pre_hash_internal 3s 6s -50%
sk_t0hat_get_poly 3s 2s +50%
yvec_get_poly 3s 3s +0%
yvec_init 3s 3s +0%
decompose 2s 2s +0%
fqscale 2s 3s -33%
intt_native_aarch64 2s 9s -78%
intt_native_x86_64 2s 4s -50%
keccak_f1600_x1_native_aarch64_v84a 2s 3s -33%
keccak_f1600_x4_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 2s 3s -33%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 1s +100%
keccak_finalize 2s 3s -33%
keccak_init 2s 3s -33%
keccak_squeeze 2s 1s +100%
keccakf1600_extract_bytes (big endian) 2s 2s +0%
keccakf1600_permute_native 2s 3s -33%
keccakf1600x4_extract_bytes 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 3s -33%
keccakf1600x4_xor_bytes 2s 2s +0%
keccakf1600x4_xor_bytes_native 2s 3s -33%
mld_ct_abs_i32 2s 2s +0%
mld_ct_sel_int32 2s 2s +0%
mld_keccakf1600_extract_bytes 2s 2s +0%
mld_prepare_domain_separation_prefix 2s 4s -50%
mld_sign_attempt 2s - new
mld_sign_resume 2s - new
mld_value_barrier_i64 2s 2s +0%
nttunpack_native_x86_64 2s 1s +100%
pack_sk_rho_key_tr_s2 2s 3s -33%
pack_sk_s1 2s 2s +0%
pointwise_native_x86_64 2s 4s -50%
poly_caddq 2s 2s +0%
poly_caddq_native 2s 5s -60%
poly_caddq_native_aarch64 2s 4s -50%
poly_chknorm_native 2s 4s -50%
poly_chknorm_native_aarch64 2s 4s -50%
poly_decompose 2s 3s -33%
poly_decompose_native 2s 3s -33%
poly_decompose_native_x86_64 2s 3s -33%
poly_invntt_tomont 2s 4s -50%
poly_invntt_tomont_native 2s 4s -50%
poly_permute_bitrev_to_custom_optional 2s 2s +0%
poly_shiftl 2s 3s -33%
poly_uniform_eta 2s 5s -60%
poly_uniform_gamma1_4x 2s 5s -60%
poly_use_hint_native_x86_64 2s - new
polyt1_pack 2s 4s -50%
polyvec_matrix_pointwise_montgomery_row 2s 2s +0%
polyveck_decompose 2s 4s -50%
polyveck_invntt_tomont 2s 5s -60%
polyvecl_chknorm 2s 10s -80%
polyvecl_pack_eta 2s 3s -33%
polyvecl_pointwise_acc_montgomery_native 2s 2s +0%
polyz_pack 2s 2s +0%
polyz_unpack 2s 3s -33%
polyz_unpack_native 2s 2s +0%
power2round 2s 3s -33%
rej_eta 2s 3s -33%
rej_eta_native 2s 6s -67%
rej_uniform_eta_native_x86_64 2s - new
shake128_finalize 2s 2s +0%
shake128_init 2s 3s -33%
shake128x4_absorb_once 2s 3s -33%
shake128x4_squeezeblocks 2s 1s +100%
shake256 2s 1s +100%
shake256_absorb 2s 3s -33%
shake256_finalize 2s 2s +0%
shake256_init 2s 3s -33%
shake256_release 2s 2s +0%
shake256_squeeze 2s 2s +0%
shake256x4_squeezeblocks 2s 4s -50%
sign_verify 2s 4s -50%
sign_verify_pre_hash_internal 2s 3s -33%
sys_check_capability 2s 4s -50%
unpack_sk_s1hat 2s 1s +100%
unpack_sk_s2hat 2s 4s -50%
caddq 1s 2s -50%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_f1600_x4_native_avx2 1s 3s -67%
keccakf1600_permute 1s 2s -50%
keccakf1600_xor_bytes (big endian) 1s 4s -75%
keccakf1600x4_permute 1s 4s -75%
mld_ct_cmask_nonzero_u32 1s 2s -50%
mld_ct_cmask_nonzero_u8 1s 2s -50%
mld_ct_get_optblocker_i64 1s 4s -75%
mld_ct_get_optblocker_u32 1s 1s +0%
mld_ct_get_optblocker_u8 1s 1s +0%
mld_ct_memcmp 1s 3s -67%
mld_keccakf1600x4_extract_bytes_c 1s 2s -50%
mld_keccakf1600x4_xor_bytes_c 1s 2s -50%
mld_polymat_expand_entry 1s 3s -67%
mld_sample_s1_s2 1s 4s -75%
mld_value_barrier_u32 1s 2s -50%
mld_value_barrier_u8 1s 3s -67%
montgomery_reduce 1s 3s -67%
pack_sig_c 1s 2s -50%
poly_caddq_c 1s 2s -50%
poly_challenge 1s 3s -67%
poly_chknorm 1s 4s -75%
poly_decompose_c 1s 4s -75%
poly_permute_bitrev_to_custom_optional_native 1s 4s -75%
poly_pointwise_montgomery_native 1s 2s -50%
poly_reduce 1s 3s -67%
poly_uniform_gamma1 1s 4s -75%
poly_use_hint 1s 4s -75%
polyeta_pack 1s 3s -67%
polyveck_reduce 1s 5s -80%
polyvecl_uniform_gamma1 1s 3s -67%
polyw1_pack 1s 3s -67%
polyw1_pack_32 1s 2s -50%
rej_uniform_native_aarch64 1s 5s -80%
shake128_release 1s 1s +0%
shake128_squeeze 1s 1s +0%
shake256x4_absorb_once 1s 3s -67%
sk_s1hat_get_poly 1s 3s -67%
sk_s2hat_get_poly 1s 2s -50%
unpack_pk_t1 1s 5s -80%
unpack_sk_t0hat 1s 3s -67%
use_hint 1s 2s -50%

@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-44, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 32s 12s +167%
mld_attempt_signature_generation ⚠️ 140s 19s +637%
sign_keypair_internal ⚠️ 21s 4s +425%
sign_pk_from_sk ⚠️ 40s 5s +700%
sign_verify_internal ⚠️ 158s 49s +222%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** 1218s 1364s -10.7%
sign_verify_internal ⚠️ 158s 49s +222%
mld_attempt_signature_generation ⚠️ 140s 19s +637%
polyvec_matrix_pointwise_montgomery_yvec 98s 149s -34%
poly_pointwise_montgomery_c 64s 112s -43%
mld_invntt_layer 53s 105s -50%
sign_pk_from_sk ⚠️ 40s 5s +700%
compute_pack_t0_t1 ⚠️ 32s 12s +167%
fqmul 22s 39s -44%
sign_keypair_internal ⚠️ 21s 4s +425%
mld_ntt_layer 20s 41s -51%
keccakf1600x4_permute_native 12s 22s -45%
mld_ntt_butterfly_block 12s 23s -48%
sig_unpack_hints 12s 2s +500%
sign_signature_internal 11s 4s +175%
rej_uniform_c 10s 16s -38%
poly_invntt_tomont_c 9s 11s -18%
polyt0_unpack 9s 12s -25%
poly_ntt_c 8s 21s -62%
rej_uniform_native_x86_64 8s - new
sign_signature_pre_hash_shake256 8s 4s +100%
poly_uniform_eta_4x 7s 12s -42%
polyeta_unpack 7s 14s -50%
polyz_unpack_c 7s 8s -12%
rej_uniform 7s 8s -12%
keccak_absorb_once_x4 6s 8s -25%
poly_chknorm_c 6s 11s -45%
mld_check_pct 5s 15s -67%
mld_h 5s 4s +25%
mld_sample_s1_s2_serial 5s 4s +25%
poly_use_hint_native_x86_64 5s - new
polyveck_chknorm 5s 67s -93%
polyw1_pack 5s 4s +25%
sign_signature_pre_hash_internal 5s 4s +25%
sign_verify_pre_hash_internal 5s 5s +0%
keccak_absorb 4s 3s +33%
keccak_squeezeblocks_x4 4s 3s +33%
mld_keccakf1600_permute_c 4s 7s -43%
pointwise_native_x86_64 4s 4s +0%
poly_add 4s 7s -43%
poly_decompose 4s 2s +100%
poly_permute_bitrev_to_custom_optional_native 4s 3s +33%
poly_pointwise_montgomery 4s 3s +33%
poly_uniform 4s 3s +33%
polyt0_pack 4s 4s +0%
polyvec_matrix_pointwise_montgomery_row 4s 6s -33%
polyveck_decompose 4s 7s -43%
polyvecl_chknorm 4s 11s -64%
polyvecl_uniform_gamma1_serial 4s 2s +100%
polyz_unpack_19_native_aarch64 4s 4s +0%
sign_signature 4s 6s -33%
sign_verify_pre_hash_shake256 4s 5s -20%
sk_s2hat_get_poly 4s 5s -20%
unpack_pk_t1 4s 3s +33%
caddq 3s 6s -50%
fqscale 3s 2s +50%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 2s +50%
keccak_init 3s 3s +0%
keccakf1600_permute_native 3s 3s +0%
keccakf1600x4_xor_bytes 3s 3s +0%
mld_ct_abs_i32 3s 3s +0%
mld_ct_cmask_neg_i32 3s 2s +50%
mld_sign_attempt 3s - new
mld_sign_finish 3s - new
ntt_native_aarch64 3s 4s -25%
pointwise_acc_native_aarch64 3s 6s -50%
pointwise_acc_native_x86_64 3s 4s -25%
poly_caddq_native_x86_64 3s 4s -25%
poly_chknorm_native 3s 4s -25%
poly_decompose_c 3s 5s -40%
poly_uniform_eta 3s 2s +50%
polyt1_pack 3s 3s +0%
polyt1_unpack 3s 2s +50%
polyvec_matrix_expand 3s 3s +0%
polyvec_matrix_expand_serial 3s 4s -25%
polyveck_caddq 3s 3s +0%
polyveck_pack_w1 3s 3s +0%
polyvecl_ntt 3s 3s +0%
polyvecl_pointwise_acc_montgomery_native 3s 3s +0%
polyvecl_unpack_eta 3s 4s -25%
polyz_unpack_17_native_aarch64 3s 2s +50%
power2round 3s 2s +50%
rej_eta 3s 5s -40%
rej_eta_c 3s 4s -25%
rej_uniform_eta_native_x86_64 3s - new
rej_uniform_native 3s 5s -40%
rej_uniform_native_aarch64 3s 3s +0%
shake128x4_squeezeblocks 3s 2s +50%
sign_keypair 3s 5s -40%
sign_signature_extmu 3s 4s -25%
sign_verify_extmu 3s 5s -40%
sk_t0hat_get_poly 3s 1s +200%
unpack_sk_s1hat 3s 3s +0%
unpack_sk_s2hat 3s 4s -25%
unpack_sk_t0hat 3s 2s +50%
intt_native_aarch64 2s 3s -33%
keccak_f1600_x1_native_aarch64 2s 3s -33%
keccak_f1600_x1_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v84a 2s 3s -33%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccak_finalize 2s 2s +0%
keccak_squeeze 2s 2s +0%
keccakf1600_extract_bytes (big endian) 2s 4s -50%
keccakf1600_permute 2s 2s +0%
keccakf1600_xor_bytes 2s 3s -33%
keccakf1600_xor_bytes (big endian) 2s 6s -67%
keccakf1600x4_permute 2s 4s -50%
keccakf1600x4_xor_bytes_native 2s 5s -60%
mld_compute_pack_z 2s 6s -67%
mld_ct_cmask_nonzero_u32 2s 3s -33%
mld_ct_cmask_nonzero_u8 2s 2s +0%
mld_ct_get_optblocker_i64 2s 4s -50%
mld_ct_memcmp 2s 1s +100%
mld_polymat_expand_entry 2s 2s +0%
mld_prepare_domain_separation_prefix 2s 3s -33%
mld_sample_s1_s2 2s 1s +100%
mld_sign_resume 2s - new
mld_value_barrier_u32 2s 1s +100%
mld_value_barrier_u8 2s 4s -50%
montgomery_reduce 2s 2s +0%
pack_sig_c 2s 5s -60%
pack_sk_s1 2s 3s -33%
poly_caddq 2s 4s -50%
poly_caddq_c 2s 3s -33%
poly_caddq_native 2s 3s -33%
poly_caddq_native_aarch64 2s 3s -33%
poly_chknorm 2s 3s -33%
poly_chknorm_native_aarch64 2s 2s +0%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose_88_native_aarch64 2s 3s -33%
poly_invntt_tomont 2s 2s +0%
poly_invntt_tomont_native 2s 3s -33%
poly_ntt_native 2s 3s -33%
poly_permute_bitrev_to_custom_optional 2s 5s -60%
poly_pointwise_montgomery_native 2s 2s +0%
poly_power2round 2s 4s -50%
poly_reduce 2s 4s -50%
poly_sub 2s 3s -33%
poly_uniform_4x 2s 3s -33%
poly_uniform_gamma1 2s 3s -33%
poly_uniform_gamma1_4x 2s 5s -60%
poly_use_hint_c 2s 5s -60%
poly_use_hint_native 2s 3s -33%
poly_use_hint_native_aarch64 2s 3s -33%
polyeta_pack 2s 3s -33%
polyveck_ntt 2s 2s +0%
polyveck_pack_eta 2s 4s -50%
polyveck_reduce 2s 4s -50%
polyvecl_pointwise_acc_montgomery_c 2s 4s -50%
polyvecl_unpack_z 2s 3s -33%
polyw1_pack_32 2s 3s -33%
polyw1_pack_88 2s 2s +0%
polyz_pack 2s 2s +0%
polyz_unpack 2s 3s -33%
polyz_unpack_native 2s 3s -33%
rej_eta_native 2s 3s -33%
rej_uniform_eta_native_aarch64 2s 3s -33%
shake128_absorb 2s 3s -33%
shake128_finalize 2s 3s -33%
shake128_release 2s 2s +0%
shake128_squeeze 2s 2s +0%
shake128x4_absorb_once 2s 2s +0%
shake256 2s 1s +100%
shake256_absorb 2s 2s +0%
shake256_finalize 2s 2s +0%
shake256x4_absorb_once 2s 1s +100%
sk_s1hat_get_poly 2s 2s +0%
sys_check_capability 2s 2s +0%
unpack_sk 2s 3s -33%
use_hint 2s 2s +0%
yvec_get_poly 2s 3s -33%
decompose 1s 3s -67%
intt_native_x86_64 1s 2s -50%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 3s -67%
keccakf1600x4_extract_bytes 1s 2s -50%
keccakf1600x4_extract_bytes_native 1s 4s -75%
make_hint 1s 3s -67%
mld_ct_get_optblocker_u32 1s 2s -50%
mld_ct_get_optblocker_u8 1s 2s -50%
mld_ct_sel_int32 1s 3s -67%
mld_keccakf1600_extract_bytes 1s 2s -50%
mld_keccakf1600x4_extract_bytes_c 1s 4s -75%
mld_keccakf1600x4_xor_bytes_c 1s 4s -75%
mld_value_barrier_i64 1s 3s -67%
ntt_native_x86_64 1s 3s -67%
nttunpack_native_x86_64 1s 4s -75%
pack_sig_h 1s 4s -75%
pack_sig_z 1s 2s -50%
pack_sk_rho_key_tr_s2 1s 5s -80%
pointwise_native_aarch64 1s 2s -50%
poly_challenge 1s 6s -83%
poly_decompose_32_native_aarch64 1s 4s -75%
poly_decompose_native 1s 2s -50%
poly_decompose_native_x86_64 1s 2s -50%
poly_ntt 1s 3s -67%
poly_shiftl 1s 4s -75%
poly_use_hint 1s 2s -50%
polyveck_invntt_tomont 1s 5s -80%
polyveck_unpack_eta 1s 2s -50%
polyvecl_pack_eta 1s 2s -50%
polyvecl_pointwise_acc_montgomery 1s 3s -67%
polyvecl_uniform_gamma1 1s 2s -50%
polyz_unpack_native_x86_64 1s 3s -67%
reduce32 1s 3s -67%
shake128_init 1s 2s -50%
shake256_init 1s 4s -75%
shake256_release 1s 1s +0%
shake256_squeeze 1s 2s -50%
shake256x4_squeezeblocks 1s 2s -50%
sign_verify 1s 4s -75%
yvec_init 1s 1s +0%

@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
**TOTAL** ⚠️ 1952s 1436s +35.9%
compute_pack_t0_t1 ⚠️ 34s 12s +183%
mld_attempt_signature_generation ⚠️ 294s 34s +765%
sign_keypair_internal ⚠️ 60s 6s +900%
sign_pk_from_sk ⚠️ 75s 6s +1150%
sign_verify_internal ⚠️ 384s 47s +717%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** ⚠️ 1952s 1436s +35.9%
sign_verify_internal ⚠️ 384s 47s +717%
mld_attempt_signature_generation ⚠️ 294s 34s +765%
polyvec_matrix_pointwise_montgomery_yvec 233s 196s +19%
poly_pointwise_montgomery_c 85s 125s -32%
sign_pk_from_sk ⚠️ 75s 6s +1150%
mld_invntt_layer 73s 111s -34%
sign_keypair_internal ⚠️ 60s 6s +900%
compute_pack_t0_t1 ⚠️ 34s 12s +183%
mld_ntt_layer 28s 44s -36%
fqmul 27s 39s -31%
sig_unpack_hints 15s 4s +275%
sign_signature_internal 15s 3s +400%
mld_ntt_butterfly_block 13s 23s -43%
rej_uniform_c 13s 18s -28%
keccakf1600x4_permute_native 11s 25s -56%
poly_ntt_c 11s 19s -42%
rej_uniform 11s 7s +57%
polyt0_unpack 10s 14s -29%
polyveck_decompose 10s 12s -17%
poly_uniform_eta_4x 9s 12s -25%
rej_uniform_native_x86_64 9s - new
keccak_absorb_once_x4 7s 8s -12%
poly_invntt_tomont_c 7s 9s -22%
polyeta_unpack 7s 13s -46%
sign_keypair 7s 4s +75%
sign_verify_extmu 7s 4s +75%
pointwise_acc_native_x86_64 6s 8s -25%
poly_invntt_tomont 6s 4s +50%
poly_use_hint_native 6s 1s +500%
polyvec_matrix_pointwise_montgomery_row 6s 13s -54%
polyvecl_ntt 6s 8s -25%
sign_verify_pre_hash_shake256 6s 7s -14%
unpack_sk_s1hat 6s 3s +100%
fqscale 5s 1s +400%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 5s 2s +150%
mld_check_pct 5s 15s -67%
mld_compute_pack_z 5s 6s -17%
mld_sample_s1_s2_serial 5s 6s -17%
poly_chknorm_c 5s 13s -62%
polyveck_caddq 5s 8s -38%
polyvecl_chknorm 5s 38s -87%
polyvecl_pointwise_acc_montgomery_c 5s 2s +150%
polyvecl_unpack_z 5s 3s +67%
sign_signature 5s 4s +25%
sign_signature_pre_hash_shake256 5s 3s +67%
keccak_absorb 4s 4s +0%
keccakf1600_permute 4s 2s +100%
keccakf1600_xor_bytes (big endian) 4s 2s +100%
mld_keccakf1600_permute_c 4s 8s -50%
mld_sign_resume 4s - new
poly_add 4s 8s -50%
poly_caddq_native 4s 3s +33%
poly_chknorm_native_x86_64 4s 2s +100%
poly_decompose_32_native_aarch64 4s 3s +33%
poly_decompose_c 4s 6s -33%
poly_decompose_native 4s 2s +100%
poly_permute_bitrev_to_custom_optional_native 4s 3s +33%
polyt1_unpack 4s 3s +33%
polyveck_invntt_tomont 4s 4s +0%
polyveck_reduce 4s 6s -33%
polyz_unpack_19_native_aarch64 4s 5s -20%
rej_uniform_native 4s 5s -20%
shake256 4s 3s +33%
shake256_release 4s 5s -20%
sign_signature_pre_hash_internal 4s 2s +100%
unpack_sk_s2hat 4s 3s +33%
intt_native_aarch64 3s 4s -25%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 3s 2s +50%
keccak_squeeze 3s 5s -40%
keccak_squeezeblocks_x4 3s 4s -25%
keccakf1600_extract_bytes (big endian) 3s 3s +0%
keccakf1600x4_extract_bytes 3s 1s +200%
keccakf1600x4_xor_bytes_native 3s 2s +50%
make_hint 3s 2s +50%
mld_ct_cmask_nonzero_u8 3s 2s +50%
mld_ct_get_optblocker_u32 3s 2s +50%
mld_keccakf1600x4_extract_bytes_c 3s 3s +0%
mld_prepare_domain_separation_prefix 3s 3s +0%
mld_sample_s1_s2 3s 8s -62%
mld_sign_finish 3s - new
mld_value_barrier_i64 3s 2s +50%
mld_value_barrier_u32 3s 3s +0%
ntt_native_x86_64 3s 2s +50%
nttunpack_native_x86_64 3s 3s +0%
pack_sig_c 3s 2s +50%
pointwise_acc_native_aarch64 3s 5s -40%
poly_caddq 3s 5s -40%
poly_caddq_native_aarch64 3s 3s +0%
poly_challenge 3s 4s -25%
poly_chknorm 3s 4s -25%
poly_chknorm_native 3s 1s +200%
poly_invntt_tomont_native 3s 2s +50%
poly_pointwise_montgomery_native 3s 3s +0%
poly_power2round 3s 7s -57%
poly_uniform 3s 4s -25%
poly_uniform_4x 3s 3s +0%
poly_uniform_gamma1_4x 3s 3s +0%
poly_use_hint_c 3s 4s -25%
poly_use_hint_native_x86_64 3s - new
polyt1_pack 3s 5s -40%
polyvec_matrix_expand 3s 3s +0%
polyvec_matrix_expand_serial 3s 3s +0%
polyveck_chknorm 3s 9s -67%
polyvecl_pack_eta 3s 2s +50%
polyvecl_pointwise_acc_montgomery_native 3s 2s +50%
polyvecl_unpack_eta 3s 3s +0%
polyw1_pack_88 3s 2s +50%
polyz_unpack_17_native_aarch64 3s 4s -25%
polyz_unpack_native 3s 1s +200%
rej_eta_c 3s 4s -25%
rej_eta_native 3s 4s -25%
rej_uniform_eta_native_x86_64 3s - new
shake128x4_absorb_once 3s 4s -25%
sign_verify 3s 2s +50%
sign_verify_pre_hash_internal 3s 3s +0%
sk_s1hat_get_poly 3s 3s +0%
sys_check_capability 3s 2s +50%
unpack_pk_t1 3s 2s +50%
unpack_sk 3s 4s -25%
use_hint 3s 3s +0%
yvec_init 3s 4s -25%
caddq 2s 3s -33%
intt_native_x86_64 2s 4s -50%
keccak_f1600_x1_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v84a 2s 1s +100%
keccak_f1600_x4_native_avx2 2s 3s -33%
keccakf1600_permute_native 2s 3s -33%
keccakf1600_xor_bytes 2s 1s +100%
keccakf1600x4_permute 2s 2s +0%
keccakf1600x4_xor_bytes 2s 1s +100%
mld_ct_abs_i32 2s 1s +100%
mld_ct_get_optblocker_i64 2s 2s +0%
mld_ct_get_optblocker_u8 2s 3s -33%
mld_h 2s 2s +0%
mld_polymat_expand_entry 2s 4s -50%
mld_sign_attempt 2s - new
montgomery_reduce 2s 2s +0%
ntt_native_aarch64 2s 4s -50%
pack_sig_h 2s 3s -33%
pack_sk_rho_key_tr_s2 2s 2s +0%
pack_sk_s1 2s 2s +0%
pointwise_native_aarch64 2s 4s -50%
pointwise_native_x86_64 2s 5s -60%
poly_caddq_c 2s 3s -33%
poly_caddq_native_x86_64 2s 3s -33%
poly_decompose 2s 2s +0%
poly_decompose_native_x86_64 2s 3s -33%
poly_ntt 2s 2s +0%
poly_ntt_native 2s 4s -50%
poly_uniform_eta 2s 5s -60%
poly_uniform_gamma1 2s 3s -33%
poly_use_hint 2s 4s -50%
poly_use_hint_native_aarch64 2s 3s -33%
polyeta_pack 2s 3s -33%
polyveck_ntt 2s 3s -33%
polyveck_pack_eta 2s 5s -60%
polyveck_unpack_eta 2s 3s -33%
polyvecl_uniform_gamma1 2s 3s -33%
polyvecl_uniform_gamma1_serial 2s 2s +0%
polyw1_pack_32 2s 3s -33%
polyz_unpack_c 2s 7s -71%
power2round 2s 3s -33%
reduce32 2s 3s -33%
rej_eta 2s 2s +0%
rej_uniform_eta_native_aarch64 2s 3s -33%
rej_uniform_native_aarch64 2s 3s -33%
shake128_absorb 2s 2s +0%
shake128_finalize 2s 2s +0%
shake128_init 2s 2s +0%
shake128_release 2s 3s -33%
shake128x4_squeezeblocks 2s 1s +100%
shake256_finalize 2s 3s -33%
shake256_squeeze 2s 2s +0%
shake256x4_absorb_once 2s 5s -60%
sign_signature_extmu 2s 4s -50%
sk_t0hat_get_poly 2s 3s -33%
unpack_sk_t0hat 2s 4s -50%
yvec_get_poly 2s 2s +0%
decompose 1s 2s -50%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_finalize 1s 1s +0%
keccak_init 1s 1s +0%
keccakf1600x4_extract_bytes_native 1s 4s -75%
mld_ct_cmask_neg_i32 1s 2s -50%
mld_ct_cmask_nonzero_u32 1s 5s -80%
mld_ct_memcmp 1s 1s +0%
mld_ct_sel_int32 1s 2s -50%
mld_keccakf1600_extract_bytes 1s 1s +0%
mld_keccakf1600x4_xor_bytes_c 1s 1s +0%
mld_value_barrier_u8 1s 2s -50%
pack_sig_z 1s 3s -67%
poly_chknorm_native_aarch64 1s 5s -80%
poly_decompose_88_native_aarch64 1s 2s -50%
poly_permute_bitrev_to_custom_optional 1s 2s -50%
poly_pointwise_montgomery 1s 3s -67%
poly_reduce 1s 4s -75%
poly_shiftl 1s 4s -75%
poly_sub 1s 5s -80%
polyt0_pack 1s 3s -67%
polyveck_pack_w1 1s 3s -67%
polyvecl_pointwise_acc_montgomery 1s 2s -50%
polyw1_pack 1s 2s -50%
polyz_pack 1s 4s -75%
polyz_unpack 1s 3s -67%
polyz_unpack_native_x86_64 1s 3s -67%
shake128_squeeze 1s 1s +0%
shake256_absorb 1s 3s -67%
shake256_init 1s 3s -67%
shake256x4_squeezeblocks 1s 4s -75%
sk_s2hat_get_poly 1s 2s -50%

@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-65)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 135s 13s +938%
mld_attempt_signature_generation ⚠️ 263s 63s +317%
sig_unpack_hints ⚠️ 28s 2s +1300%
sign_keypair_internal ⚠️ 23s 5s +360%
sign_pk_from_sk ⚠️ 45s 6s +650%
sign_signature_internal ⚠️ 195s 55s +255%
sign_verify_internal ⚠️ 410s 177s +132%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** 2217s 1784s +24.3%
sign_verify_internal ⚠️ 410s 177s +132%
mld_attempt_signature_generation ⚠️ 263s 63s +317%
sign_signature_internal ⚠️ 195s 55s +255%
polyvecl_pointwise_acc_montgomery_c 178s 210s -15%
poly_pointwise_montgomery_c 146s 125s +17%
compute_pack_t0_t1 ⚠️ 135s 13s +938%
polyvec_matrix_expand 76s 110s -31%
mld_invntt_layer 60s 109s -45%
sign_pk_from_sk ⚠️ 45s 6s +650%
sig_unpack_hints ⚠️ 28s 2s +1300%
mld_ntt_layer 24s 41s -41%
sign_keypair_internal ⚠️ 23s 5s +360%
fqmul 22s 40s -45%
polyvec_matrix_expand_serial 18s 27s -33%
keccakf1600x4_permute_native 15s 23s -35%
polyveck_ntt 14s 7s +100%
mld_ntt_butterfly_block 13s 24s -46%
poly_ntt_c 11s 19s -42%
polyveck_decompose 11s 12s -8%
poly_decompose_c 9s 5s +80%
rej_uniform_c 9s 14s -36%
polyt0_unpack 8s 16s -50%
polyvec_matrix_pointwise_montgomery_yvec 8s 18s -56%
rej_uniform 8s 18s -56%
rej_uniform_native_x86_64 8s - new
poly_chknorm_c 7s 15s -53%
sign_signature_pre_hash_shake256 7s 3s +133%
mld_sample_s1_s2_serial 6s 4s +50%
poly_uniform_4x 6s 13s -54%
rej_uniform_native 6s 4s +50%
sign_signature_extmu 6s 2s +200%
decompose 5s 1s +400%
keccak_absorb 5s 3s +67%
mld_compute_pack_z 5s 7s -29%
poly_decompose_native_x86_64 5s 3s +67%
poly_invntt_tomont_c 5s 11s -55%
poly_uniform_eta_4x 5s 14s -64%
poly_uniform_gamma1 5s 3s +67%
sign_signature 5s 6s -17%
sign_signature_pre_hash_internal 5s 3s +67%
intt_native_aarch64 4s 3s +33%
keccak_absorb_once_x4 4s 8s -50%
keccak_squeezeblocks_x4 4s 5s -20%
mld_check_pct 4s 14s -71%
mld_keccakf1600_permute_c 4s 6s -33%
pack_sig_h 4s 5s -20%
pointwise_acc_native_x86_64 4s 8s -50%
poly_caddq 4s 2s +100%
poly_chknorm_native_aarch64 4s 3s +33%
poly_uniform_eta 4s 4s +0%
poly_use_hint_native_aarch64 4s 3s +33%
polyveck_caddq 4s 6s -33%
polyveck_pack_w1 4s 3s +33%
polyw1_pack_88 4s 2s +100%
polyz_pack 4s 3s +33%
rej_eta_c 4s 5s -20%
rej_eta_native 4s 3s +33%
rej_uniform_native_aarch64 4s 3s +33%
sign_keypair 4s 3s +33%
sk_t0hat_get_poly 4s 2s +100%
unpack_sk_s1hat 4s 2s +100%
use_hint 4s 2s +100%
keccakf1600_xor_bytes 3s 2s +50%
keccakf1600x4_extract_bytes 3s 4s -25%
keccakf1600x4_permute 3s 4s -25%
make_hint 3s 2s +50%
mld_ct_cmask_neg_i32 3s 2s +50%
mld_ct_get_optblocker_u32 3s 3s +0%
mld_ct_sel_int32 3s 3s +0%
mld_h 3s 2s +50%
mld_keccakf1600x4_xor_bytes_c 3s 2s +50%
mld_polymat_expand_entry 3s 4s -25%
mld_prepare_domain_separation_prefix 3s 2s +50%
mld_sample_s1_s2 3s 6s -50%
mld_sign_finish 3s - new
mld_value_barrier_i64 3s 2s +50%
ntt_native_aarch64 3s 2s +50%
pack_sig_z 3s 2s +50%
pack_sk_s1 3s 5s -40%
pointwise_acc_native_aarch64 3s 6s -50%
poly_add 3s 9s -67%
poly_caddq_c 3s 2s +50%
poly_caddq_native_aarch64 3s 2s +50%
poly_decompose_32_native_aarch64 3s 3s +0%
poly_invntt_tomont_native 3s 2s +50%
poly_permute_bitrev_to_custom_optional_native 3s 4s -25%
poly_power2round 3s 4s -25%
poly_reduce 3s 3s +0%
poly_shiftl 3s 3s +0%
poly_sub 3s 3s +0%
poly_use_hint_c 3s 5s -40%
polyt1_unpack 3s 5s -40%
polyveck_unpack_eta 3s 6s -50%
polyvecl_chknorm 3s 7s -57%
polyvecl_ntt 3s 4s -25%
polyvecl_pointwise_acc_montgomery 3s 3s +0%
polyvecl_uniform_gamma1 3s 2s +50%
polyz_unpack_17_native_aarch64 3s 3s +0%
polyz_unpack_19_native_aarch64 3s 3s +0%
polyz_unpack_c 3s 13s -77%
polyz_unpack_native 3s 4s -25%
rej_uniform_eta_native_x86_64 3s - new
shake128_absorb 3s 2s +50%
shake128x4_absorb_once 3s 3s +0%
shake256_init 3s 2s +50%
shake256x4_absorb_once 3s 3s +0%
sign_verify_extmu 3s 3s +0%
sign_verify_pre_hash_internal 3s 5s -40%
sk_s1hat_get_poly 3s 6s -50%
unpack_sk_s2hat 3s 4s -25%
caddq 2s 3s -33%
fqscale 2s 3s -33%
intt_native_x86_64 2s 3s -33%
keccak_f1600_x1_native_aarch64 2s 2s +0%
keccak_f1600_x1_native_aarch64_v84a 2s 3s -33%
keccak_f1600_x4_native_aarch64_v84a 2s 4s -50%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 2s 2s +0%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 2s +0%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccak_init 2s 1s +100%
keccak_squeeze 2s 4s -50%
keccakf1600_extract_bytes (big endian) 2s 2s +0%
keccakf1600_permute 2s 2s +0%
keccakf1600_xor_bytes (big endian) 2s 2s +0%
mld_ct_abs_i32 2s 2s +0%
mld_ct_get_optblocker_i64 2s 2s +0%
mld_ct_memcmp 2s 2s +0%
mld_keccakf1600_extract_bytes 2s 1s +100%
mld_sign_attempt 2s - new
mld_value_barrier_u8 2s 3s -33%
ntt_native_x86_64 2s 2s +0%
nttunpack_native_x86_64 2s 3s -33%
pack_sig_c 2s 5s -60%
pack_sk_rho_key_tr_s2 2s 2s +0%
pointwise_native_aarch64 2s 5s -60%
pointwise_native_x86_64 2s 3s -33%
poly_caddq_native 2s 2s +0%
poly_challenge 2s 6s -67%
poly_chknorm 2s 4s -50%
poly_chknorm_native 2s 2s +0%
poly_decompose 2s 2s +0%
poly_decompose_88_native_aarch64 2s 1s +100%
poly_invntt_tomont 2s 3s -33%
poly_ntt 2s 2s +0%
poly_permute_bitrev_to_custom_optional 2s 2s +0%
poly_pointwise_montgomery 2s 3s -33%
poly_uniform 2s 4s -50%
poly_uniform_gamma1_4x 2s 4s -50%
poly_use_hint 2s 3s -33%
poly_use_hint_native_x86_64 2s - new
polyeta_pack 2s 3s -33%
polyeta_unpack 2s 8s -75%
polyt0_pack 2s 3s -33%
polyt1_pack 2s 3s -33%
polyveck_chknorm 2s 6s -67%
polyveck_invntt_tomont 2s 8s -75%
polyveck_reduce 2s 4s -50%
polyvecl_pack_eta 2s 3s -33%
polyvecl_uniform_gamma1_serial 2s 3s -33%
polyvecl_unpack_eta 2s 3s -33%
polyw1_pack 2s 4s -50%
polyz_unpack_native_x86_64 2s 3s -33%
power2round 2s 1s +100%
rej_eta 2s 2s +0%
shake128_init 2s 2s +0%
shake256 2s 2s +0%
shake256_release 2s 2s +0%
shake256x4_squeezeblocks 2s 2s +0%
sign_verify 2s 3s -33%
sign_verify_pre_hash_shake256 2s 5s -60%
sk_s2hat_get_poly 2s 2s +0%
unpack_pk_t1 2s 5s -60%
unpack_sk 2s 3s -33%
yvec_get_poly 2s 3s -33%
yvec_init 2s 3s -33%
keccak_finalize 1s 2s -50%
keccakf1600_permute_native 1s 3s -67%
keccakf1600x4_extract_bytes_native 1s 4s -75%
keccakf1600x4_xor_bytes 1s 1s +0%
keccakf1600x4_xor_bytes_native 1s 3s -67%
mld_ct_cmask_nonzero_u32 1s 4s -75%
mld_ct_cmask_nonzero_u8 1s 3s -67%
mld_ct_get_optblocker_u8 1s 2s -50%
mld_keccakf1600x4_extract_bytes_c 1s 3s -67%
mld_sign_resume 1s - new
mld_value_barrier_u32 1s 3s -67%
montgomery_reduce 1s 4s -75%
poly_caddq_native_x86_64 1s 3s -67%
poly_chknorm_native_x86_64 1s 3s -67%
poly_decompose_native 1s 3s -67%
poly_ntt_native 1s 4s -75%
poly_pointwise_montgomery_native 1s 3s -67%
poly_use_hint_native 1s 2s -50%
polyvec_matrix_pointwise_montgomery_row 1s 1s +0%
polyveck_pack_eta 1s 4s -75%
polyvecl_pointwise_acc_montgomery_native 1s 2s -50%
polyvecl_unpack_z 1s 3s -67%
polyw1_pack_32 1s 4s -75%
polyz_unpack 1s 3s -67%
reduce32 1s 3s -67%
rej_uniform_eta_native_aarch64 1s 4s -75%
shake128_finalize 1s 1s +0%
shake128_release 1s 3s -67%
shake128_squeeze 1s 2s -50%
shake128x4_squeezeblocks 1s 3s -67%
shake256_absorb 1s 2s -50%
shake256_finalize 1s 1s +0%
shake256_squeeze 1s 3s -67%
sys_check_capability 1s 4s -75%
unpack_sk_t0hat 1s 5s -80%

@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-65, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 33s 7s +371%
mld_attempt_signature_generation ⚠️ 153s 31s +394%
sign_keypair_internal ⚠️ 27s 3s +800%
sign_pk_from_sk ⚠️ 55s 5s +1000%
sign_verify_internal ⚠️ 176s 71s +148%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** 1394s 1482s -5.9%
sign_verify_internal ⚠️ 176s 71s +148%
mld_attempt_signature_generation ⚠️ 153s 31s +394%
polyvec_matrix_pointwise_montgomery_yvec 138s 201s -31%
poly_pointwise_montgomery_c 69s 116s -41%
mld_invntt_layer 56s 108s -48%
sign_pk_from_sk ⚠️ 55s 5s +1000%
polyvecl_chknorm 51s 43s +19%
compute_pack_t0_t1 ⚠️ 33s 7s +371%
sign_keypair_internal ⚠️ 27s 3s +800%
fqmul 23s 40s -43%
mld_ntt_layer 21s 43s -51%
keccakf1600x4_permute_native 13s 22s -41%
rej_uniform_c 13s 16s -19%
sign_signature_internal 13s 6s +117%
mld_ntt_butterfly_block 11s 25s -56%
poly_ntt_c 10s 22s -55%
polyt0_unpack 10s 13s -23%
sig_unpack_hints 10s 1s +900%
rej_uniform 9s 7s +29%
poly_uniform_eta_4x 8s 13s -38%
rej_uniform_native_x86_64 8s - new
polyveck_decompose 7s 14s -50%
mld_check_pct 6s 12s -50%
mld_keccakf1600_permute_c 6s 7s -14%
polyeta_unpack 6s 4s +50%
sign_keypair 6s 4s +50%
sign_signature_pre_hash_shake256 6s 7s -14%
caddq 5s 2s +150%
keccak_absorb 5s 3s +67%
pointwise_acc_native_x86_64 5s 5s +0%
poly_chknorm_c 5s 13s -62%
poly_invntt_tomont_c 5s 10s -50%
poly_permute_bitrev_to_custom_optional 5s 3s +67%
polyvec_matrix_pointwise_montgomery_row 5s 8s -38%
polyvecl_ntt 5s 7s -29%
rej_uniform_native 5s 4s +25%
shake256x4_squeezeblocks 5s 2s +150%
intt_native_x86_64 4s 4s +0%
keccak_absorb_once_x4 4s 10s -60%
keccak_finalize 4s 2s +100%
pointwise_acc_native_aarch64 4s 6s -33%
poly_challenge 4s 5s -20%
poly_chknorm_native_aarch64 4s 2s +100%
poly_decompose_c 4s 8s -50%
poly_ntt 4s 2s +100%
poly_pointwise_montgomery 4s 2s +100%
poly_pointwise_montgomery_native 4s 4s +0%
polyeta_pack 4s 1s +300%
polyveck_caddq 4s 7s -43%
polyvecl_pack_eta 4s 2s +100%
polyw1_pack_32 4s 2s +100%
polyz_unpack_native 4s 3s +33%
rej_eta_c 4s 4s +0%
rej_uniform_native_aarch64 4s 5s -20%
sign_signature_extmu 4s 4s +0%
sk_t0hat_get_poly 4s 1s +300%
unpack_sk 4s 3s +33%
unpack_sk_t0hat 4s 4s +0%
intt_native_aarch64 3s 2s +50%
keccak_f1600_x1_native_aarch64 3s 2s +50%
keccak_f1600_x1_native_aarch64_v84a 3s 2s +50%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccakf1600x4_extract_bytes_native 3s 1s +200%
mld_ct_abs_i32 3s 6s -50%
mld_h 3s 3s +0%
mld_keccakf1600x4_xor_bytes_c 3s 2s +50%
mld_polymat_expand_entry 3s 3s +0%
mld_prepare_domain_separation_prefix 3s 4s -25%
mld_sign_finish 3s - new
mld_sign_resume 3s - new
montgomery_reduce 3s 1s +200%
pack_sig_h 3s 2s +50%
pack_sk_rho_key_tr_s2 3s 2s +50%
pointwise_native_x86_64 3s 2s +50%
poly_add 3s 7s -57%
poly_caddq 3s 3s +0%
poly_caddq_c 3s 5s -40%
poly_chknorm 3s 3s +0%
poly_decompose 3s 2s +50%
poly_decompose_native 3s 5s -40%
poly_invntt_tomont 3s 5s -40%
poly_permute_bitrev_to_custom_optional_native 3s 3s +0%
poly_power2round 3s 7s -57%
poly_uniform_eta 3s 5s -40%
poly_use_hint_native 3s 7s -57%
poly_use_hint_native_aarch64 3s 4s -25%
polyveck_invntt_tomont 3s 7s -57%
polyvecl_pointwise_acc_montgomery_native 3s 3s +0%
polyvecl_unpack_eta 3s 2s +50%
polyw1_pack_88 3s 3s +0%
polyz_unpack_17_native_aarch64 3s 2s +50%
polyz_unpack_native_x86_64 3s 3s +0%
power2round 3s 2s +50%
rej_uniform_eta_native_x86_64 3s - new
shake128_init 3s 6s -50%
shake128_release 3s 2s +50%
shake128_squeeze 3s 3s +0%
shake128x4_squeezeblocks 3s 2s +50%
shake256_absorb 3s 4s -25%
sign_signature 3s 5s -40%
sign_signature_pre_hash_internal 3s 3s +0%
sk_s1hat_get_poly 3s 3s +0%
unpack_sk_s1hat 3s 1s +200%
yvec_init 3s 5s -40%
fqscale 2s 3s -33%
keccak_f1600_x4_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 3s -33%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccak_squeeze 2s 2s +0%
keccak_squeezeblocks_x4 2s 4s -50%
keccakf1600_extract_bytes (big endian) 2s 3s -33%
keccakf1600_permute 2s 4s -50%
keccakf1600_permute_native 2s 2s +0%
keccakf1600_xor_bytes 2s 4s -50%
keccakf1600_xor_bytes (big endian) 2s 2s +0%
keccakf1600x4_extract_bytes 2s 4s -50%
keccakf1600x4_permute 2s 1s +100%
keccakf1600x4_xor_bytes 2s 2s +0%
keccakf1600x4_xor_bytes_native 2s 2s +0%
mld_compute_pack_z 2s 5s -60%
mld_ct_cmask_neg_i32 2s 3s -33%
mld_ct_cmask_nonzero_u32 2s 3s -33%
mld_ct_cmask_nonzero_u8 2s 2s +0%
mld_ct_get_optblocker_i64 2s 1s +100%
mld_ct_get_optblocker_u32 2s 3s -33%
mld_ct_sel_int32 2s 2s +0%
mld_keccakf1600_extract_bytes 2s 2s +0%
mld_sample_s1_s2 2s 6s -67%
mld_sample_s1_s2_serial 2s 3s -33%
mld_sign_attempt 2s - new
mld_value_barrier_i64 2s 2s +0%
mld_value_barrier_u32 2s 4s -50%
mld_value_barrier_u8 2s 1s +100%
ntt_native_aarch64 2s 3s -33%
pack_sig_c 2s 3s -33%
pack_sig_z 2s 1s +100%
pack_sk_s1 2s 3s -33%
pointwise_native_aarch64 2s 3s -33%
poly_caddq_native_aarch64 2s 2s +0%
poly_chknorm_native_x86_64 2s 4s -50%
poly_decompose_32_native_aarch64 2s 4s -50%
poly_decompose_native_x86_64 2s 2s +0%
poly_invntt_tomont_native 2s 5s -60%
poly_ntt_native 2s 2s +0%
poly_shiftl 2s 3s -33%
poly_sub 2s 4s -50%
poly_uniform_4x 2s 4s -50%
poly_uniform_gamma1 2s 3s -33%
poly_use_hint 2s 3s -33%
poly_use_hint_native_x86_64 2s - new
polyt0_pack 2s 5s -60%
polyveck_chknorm 2s 36s -94%
polyveck_ntt 2s 4s -50%
polyveck_pack_w1 2s 4s -50%
polyveck_reduce 2s 6s -67%
polyvecl_pointwise_acc_montgomery 2s 3s -33%
polyvecl_uniform_gamma1 2s 4s -50%
polyvecl_uniform_gamma1_serial 2s 1s +100%
polyvecl_unpack_z 2s 1s +100%
polyw1_pack 2s 3s -33%
polyz_unpack_c 2s 9s -78%
rej_eta 2s 5s -60%
rej_eta_native 2s 3s -33%
rej_uniform_eta_native_aarch64 2s 5s -60%
shake128_finalize 2s 2s +0%
shake256_finalize 2s 1s +100%
shake256_release 2s 3s -33%
shake256_squeeze 2s 4s -50%
shake256x4_absorb_once 2s 2s +0%
sign_verify_extmu 2s 2s +0%
sign_verify_pre_hash_internal 2s 5s -60%
sign_verify_pre_hash_shake256 2s 6s -67%
sk_s2hat_get_poly 2s 1s +100%
sys_check_capability 2s 1s +100%
unpack_sk_s2hat 2s 3s -33%
yvec_get_poly 2s 3s -33%
decompose 1s 4s -75%
keccak_init 1s 4s -75%
make_hint 1s 3s -67%
mld_ct_get_optblocker_u8 1s 4s -75%
mld_ct_memcmp 1s 4s -75%
mld_keccakf1600x4_extract_bytes_c 1s 2s -50%
ntt_native_x86_64 1s 5s -80%
nttunpack_native_x86_64 1s 3s -67%
poly_caddq_native 1s 4s -75%
poly_caddq_native_x86_64 1s 1s +0%
poly_chknorm_native 1s 4s -75%
poly_decompose_88_native_aarch64 1s 2s -50%
poly_reduce 1s 4s -75%
poly_uniform 1s 2s -50%
poly_uniform_gamma1_4x 1s 3s -67%
poly_use_hint_c 1s 2s -50%
polyt1_pack 1s 3s -67%
polyt1_unpack 1s 3s -67%
polyvec_matrix_expand 1s 6s -83%
polyvec_matrix_expand_serial 1s 3s -67%
polyveck_pack_eta 1s 3s -67%
polyveck_unpack_eta 1s 3s -67%
polyvecl_pointwise_acc_montgomery_c 1s 3s -67%
polyz_pack 1s 3s -67%
polyz_unpack 1s 3s -67%
polyz_unpack_19_native_aarch64 1s 5s -80%
reduce32 1s 2s -50%
shake128_absorb 1s 2s -50%
shake128x4_absorb_once 1s 2s -50%
shake256 1s 2s -50%
shake256_init 1s 4s -75%
sign_verify 1s 6s -83%
unpack_pk_t1 1s 2s -50%
use_hint 1s 2s -50%

@oqs-bot

oqs-bot commented Jul 10, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87)

⚠️ Attention Required

Proof Status Current Previous Change
**TOTAL** ⚠️ 2815s 2151s +30.9%
compute_pack_t0_t1 ⚠️ 189s 19s +895%
mld_attempt_signature_generation ⚠️ 288s 55s +424%
sig_unpack_hints ⚠️ 28s 3s +833%
sign_keypair_internal ⚠️ 51s 6s +750%
sign_pk_from_sk ⚠️ 47s 6s +683%
sign_signature_internal ⚠️ 314s 42s +648%
sign_verify_internal ⚠️ 751s 97s +674%
Full Results (210 proofs)
Proof Status Current Previous Change
**TOTAL** ⚠️ 2815s 2151s +30.9%
sign_verify_internal ⚠️ 751s 97s +674%
sign_signature_internal ⚠️ 314s 42s +648%
mld_attempt_signature_generation ⚠️ 288s 55s +424%
polyvecl_pointwise_acc_montgomery_c 211s 353s -40%
compute_pack_t0_t1 ⚠️ 189s 19s +895%
poly_pointwise_montgomery_c 138s 144s -4%
polyvec_matrix_expand 82s 324s -75%
mld_invntt_layer 58s 115s -50%
sign_keypair_internal ⚠️ 51s 6s +750%
sign_pk_from_sk ⚠️ 47s 6s +683%
sig_unpack_hints ⚠️ 28s 3s +833%
fqmul 25s 43s -42%
polyvec_matrix_expand_serial 22s 38s -42%
mld_ntt_layer 20s 45s -56%
mld_ntt_butterfly_block 14s 24s -42%
keccakf1600x4_permute_native 12s 23s -48%
polyvec_matrix_pointwise_montgomery_yvec 11s 18s -39%
poly_ntt_c 10s 21s -52%
polyt0_unpack 9s 14s -36%
poly_chknorm_c 8s 15s -47%
rej_uniform 8s 15s -47%
poly_uniform_eta_4x 7s 11s -36%
polyeta_unpack 7s 17s -59%
polyveck_decompose 7s 13s -46%
rej_uniform_c 7s 19s -63%
rej_uniform_native_x86_64 7s - new
sign_keypair 7s 4s +75%
keccak_absorb 6s 3s +100%
mld_compute_pack_z 6s 8s -25%
pack_sk_rho_key_tr_s2 6s 3s +100%
poly_chknorm_native 6s 4s +50%
poly_uniform_4x 6s 13s -54%
keccak_absorb_once_x4 5s 9s -44%
mld_check_pct 5s 16s -69%
pointwise_acc_native_x86_64 5s 6s -17%
poly_add 5s 6s -17%
poly_invntt_tomont_c 5s 12s -58%
poly_power2round 5s 3s +67%
polyveck_invntt_tomont 5s 9s -44%
polyvecl_ntt 5s 7s -29%
polyz_unpack_c 5s 3s +67%
polyz_unpack_native_x86_64 5s 3s +67%
rej_uniform_eta_native_aarch64 5s 5s +0%
sign_signature_pre_hash_shake256 5s 6s -17%
intt_native_aarch64 4s 3s +33%
keccakf1600x4_extract_bytes 4s 2s +100%
mld_keccakf1600_permute_c 4s 7s -43%
mld_sample_s1_s2 4s 6s -33%
mld_sample_s1_s2_serial 4s 8s -50%
pack_sig_h 4s 4s +0%
poly_chknorm_native_aarch64 4s 4s +0%
poly_decompose_32_native_aarch64 4s 5s -20%
poly_invntt_tomont_native 4s 5s -20%
poly_permute_bitrev_to_custom_optional_native 4s 5s -20%
poly_uniform_gamma1 4s 4s +0%
poly_use_hint_c 4s 2s +100%
polyt1_pack 4s 3s +33%
polyt1_unpack 4s 4s +0%
polyveck_unpack_eta 4s 4s +0%
polyvecl_chknorm 4s 7s -43%
polyvecl_pointwise_acc_montgomery 4s 5s -20%
polyvecl_uniform_gamma1 4s 2s +100%
polyw1_pack_88 4s 3s +33%
shake256x4_squeezeblocks 4s 1s +300%
sign_signature_extmu 4s 4s +0%
sign_signature_pre_hash_internal 4s 5s -20%
sign_verify 4s 5s -20%
sign_verify_pre_hash_internal 4s 3s +33%
yvec_get_poly 4s 4s +0%
intt_native_x86_64 3s 3s +0%
keccak_f1600_x4_native_aarch64_v84a 3s 4s -25%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccakf1600_permute_native 3s 4s -25%
keccakf1600_xor_bytes (big endian) 3s 2s +50%
keccakf1600x4_xor_bytes_native 3s 3s +0%
mld_sign_resume 3s - new
pointwise_acc_native_aarch64 3s 6s -50%
pointwise_native_aarch64 3s 5s -40%
poly_caddq 3s 2s +50%
poly_caddq_native_aarch64 3s 3s +0%
poly_challenge 3s 5s -40%
poly_decompose 3s 3s +0%
poly_decompose_c 3s 5s -40%
poly_decompose_native_x86_64 3s 2s +50%
poly_invntt_tomont 3s 1s +200%
poly_permute_bitrev_to_custom_optional 3s 3s +0%
poly_pointwise_montgomery_native 3s 5s -40%
poly_reduce 3s 2s +50%
poly_sub 3s 3s +0%
poly_uniform 3s 4s -25%
poly_uniform_eta 3s 4s -25%
poly_uniform_gamma1_4x 3s 3s +0%
polyveck_caddq 3s 8s -62%
polyveck_ntt 3s 11s -73%
polyveck_pack_eta 3s 6s -50%
polyveck_pack_w1 3s 3s +0%
polyw1_pack 3s 4s -25%
polyw1_pack_32 3s 4s -25%
polyz_unpack_native 3s 4s -25%
reduce32 3s 3s +0%
rej_eta_c 3s 5s -40%
rej_uniform_eta_native_x86_64 3s - new
rej_uniform_native 3s 8s -62%
sign_signature 3s 5s -40%
sign_verify_extmu 3s 3s +0%
sys_check_capability 3s 3s +0%
use_hint 3s 3s +0%
caddq 2s 4s -50%
decompose 2s 3s -33%
fqscale 2s 3s -33%
keccak_f1600_x1_native_aarch64 2s 1s +100%
keccak_f1600_x1_native_aarch64_v84a 2s 1s +100%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccak_squeeze 2s 2s +0%
keccak_squeezeblocks_x4 2s 4s -50%
keccakf1600_permute 2s 3s -33%
keccakf1600_xor_bytes 2s 3s -33%
keccakf1600x4_extract_bytes_native 2s 2s +0%
keccakf1600x4_permute 2s 3s -33%
keccakf1600x4_xor_bytes 2s 3s -33%
mld_ct_cmask_neg_i32 2s 4s -50%
mld_ct_get_optblocker_i64 2s 3s -33%
mld_ct_get_optblocker_u8 2s 3s -33%
mld_ct_sel_int32 2s 3s -33%
mld_h 2s 3s -33%
mld_keccakf1600x4_extract_bytes_c 2s 2s +0%
mld_polymat_expand_entry 2s 3s -33%
mld_prepare_domain_separation_prefix 2s 4s -50%
mld_sign_attempt 2s - new
mld_sign_finish 2s - new
mld_value_barrier_i64 2s 1s +100%
mld_value_barrier_u32 2s 2s +0%
ntt_native_aarch64 2s 6s -67%
nttunpack_native_x86_64 2s 3s -33%
pack_sig_c 2s 4s -50%
pack_sig_z 2s 5s -60%
pack_sk_s1 2s 1s +100%
pointwise_native_x86_64 2s 2s +0%
poly_caddq_c 2s 3s -33%
poly_caddq_native 2s 4s -50%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose_native 2s 2s +0%
poly_ntt 2s 3s -33%
poly_ntt_native 2s 3s -33%
poly_pointwise_montgomery 2s 4s -50%
poly_shiftl 2s 4s -50%
poly_use_hint 2s 1s +100%
poly_use_hint_native 2s 3s -33%
poly_use_hint_native_aarch64 2s 2s +0%
poly_use_hint_native_x86_64 2s - new
polyeta_pack 2s 4s -50%
polyt0_pack 2s 4s -50%
polyvec_matrix_pointwise_montgomery_row 2s 3s -33%
polyveck_chknorm 2s 3s -33%
polyveck_reduce 2s 3s -33%
polyvecl_uniform_gamma1_serial 2s 4s -50%
polyvecl_unpack_z 2s 4s -50%
polyz_pack 2s 5s -60%
polyz_unpack 2s 2s +0%
polyz_unpack_17_native_aarch64 2s 3s -33%
power2round 2s 4s -50%
rej_eta 2s 4s -50%
rej_uniform_native_aarch64 2s 2s +0%
shake128_finalize 2s 2s +0%
shake128x4_absorb_once 2s 5s -60%
shake128x4_squeezeblocks 2s 3s -33%
shake256_absorb 2s 2s +0%
shake256_init 2s 2s +0%
shake256_release 2s 3s -33%
shake256_squeeze 2s 2s +0%
sign_verify_pre_hash_shake256 2s 3s -33%
sk_s1hat_get_poly 2s 3s -33%
sk_t0hat_get_poly 2s 4s -50%
unpack_pk_t1 2s 1s +100%
unpack_sk 2s 5s -60%
unpack_sk_s1hat 2s 2s +0%
unpack_sk_t0hat 2s 7s -71%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 4s -75%
keccak_finalize 1s 1s +0%
keccak_init 1s 3s -67%
keccakf1600_extract_bytes (big endian) 1s 3s -67%
make_hint 1s 3s -67%
mld_ct_abs_i32 1s 3s -67%
mld_ct_cmask_nonzero_u32 1s 2s -50%
mld_ct_cmask_nonzero_u8 1s 3s -67%
mld_ct_get_optblocker_u32 1s 2s -50%
mld_ct_memcmp 1s 3s -67%
mld_keccakf1600_extract_bytes 1s 4s -75%
mld_keccakf1600x4_xor_bytes_c 1s 4s -75%
mld_value_barrier_u8 1s 3s -67%
montgomery_reduce 1s 3s -67%
ntt_native_x86_64 1s 2s -50%
poly_caddq_native_x86_64 1s 2s -50%
poly_chknorm 1s 3s -67%
poly_decompose_88_native_aarch64 1s 2s -50%
polyvecl_pack_eta 1s 3s -67%
polyvecl_pointwise_acc_montgomery_native 1s 3s -67%
polyvecl_unpack_eta 1s 5s -80%
polyz_unpack_19_native_aarch64 1s 4s -75%
rej_eta_native 1s 6s -83%
shake128_absorb 1s 4s -75%
shake128_init 1s 1s +0%
shake128_release 1s 4s -75%
shake128_squeeze 1s 2s -50%
shake256 1s 2s -50%
shake256_finalize 1s 2s -50%
shake256x4_absorb_once 1s 2s -50%
sk_s2hat_get_poly 1s 3s -67%
unpack_sk_s2hat 1s 3s -67%
yvec_init 1s 4s -75%

@bremoran
bremoran force-pushed the armv81m-keccak-x1 branch 3 times, most recently from a1dde44 to e8a49a1 Compare July 14, 2026 10:42
@bremoran
bremoran force-pushed the armv81m-keccak-x1 branch from d0543bf to c45a108 Compare July 23, 2026 11:03
@bremoran

Copy link
Copy Markdown
Contributor Author

Depends on either #1312 or #1318

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 46545 cycles 46515 cycles 1.00
ML-DSA-44 sign 131210 cycles 131104 cycles 1.00
ML-DSA-44 verify 47369 cycles 47340 cycles 1.00
ML-DSA-65 keypair 81706 cycles 81714 cycles 1.00
ML-DSA-65 sign 215415 cycles 215483 cycles 1.00
ML-DSA-65 verify 79313 cycles 79331 cycles 1.00
ML-DSA-87 keypair 132481 cycles 132496 cycles 1.00
ML-DSA-87 sign 277601 cycles 277751 cycles 1.00
ML-DSA-87 verify 133534 cycles 133550 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 112752 cycles 112732 cycles 1.00
ML-DSA-44 sign 401039 cycles 401165 cycles 1.00
ML-DSA-44 verify 119360 cycles 119350 cycles 1.00
ML-DSA-65 keypair 192953 cycles 192974 cycles 1.00
ML-DSA-65 sign 650062 cycles 649955 cycles 1.00
ML-DSA-65 verify 193069 cycles 193006 cycles 1.00
ML-DSA-87 keypair 318951 cycles 318940 cycles 1.00
ML-DSA-87 sign 828946 cycles 828798 cycles 1.00
ML-DSA-87 verify 321793 cycles 321792 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 51479 cycles 51560 cycles 1.00
ML-DSA-44 sign 162053 cycles 162124 cycles 1.00
ML-DSA-44 verify 54702 cycles 54700 cycles 1.00
ML-DSA-65 keypair 90292 cycles 90132 cycles 1.00
ML-DSA-65 sign 267890 cycles 267374 cycles 1.00
ML-DSA-65 verify 90005 cycles 89790 cycles 1.00
ML-DSA-87 keypair 146069 cycles 145982 cycles 1.00
ML-DSA-87 sign 335529 cycles 336138 cycles 1.00
ML-DSA-87 verify 145250 cycles 145433 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5 (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 100181 cycles 99983 cycles 1.00
ML-DSA-44 sign 369540 cycles 369043 cycles 1.00
ML-DSA-44 verify 110294 cycles 110043 cycles 1.00
ML-DSA-65 keypair 175195 cycles 175209 cycles 1.00
ML-DSA-65 sign 597735 cycles 596868 cycles 1.00
ML-DSA-65 verify 177754 cycles 177498 cycles 1.00
ML-DSA-87 keypair 286479 cycles 286715 cycles 1.00
ML-DSA-87 sign 759319 cycles 757791 cycles 1.00
ML-DSA-87 verify 296527 cycles 295756 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks (opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 112292 cycles 112421 cycles 1.00
ML-DSA-44 sign 353509 cycles 353943 cycles 1.00
ML-DSA-44 verify 117045 cycles 117107 cycles 1.00
ML-DSA-65 keypair 194768 cycles 194811 cycles 1.00
ML-DSA-65 sign 583684 cycles 583921 cycles 1.00
ML-DSA-65 verify 192841 cycles 192856 cycles 1.00
ML-DSA-87 keypair 320755 cycles 321141 cycles 1.00
ML-DSA-87 sign 746990 cycles 747840 cycles 1.00
ML-DSA-87 verify 318459 cycles 318856 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3 (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 138441 cycles 138205 cycles 1.00
ML-DSA-44 sign 486106 cycles 485764 cycles 1.00
ML-DSA-44 verify 149269 cycles 149259 cycles 1.00
ML-DSA-65 keypair 242153 cycles 241613 cycles 1.00
ML-DSA-65 sign 791700 cycles 791618 cycles 1.00
ML-DSA-65 verify 241534 cycles 241503 cycles 1.00
ML-DSA-87 keypair 396162 cycles 395195 cycles 1.00
ML-DSA-87 sign 1013608 cycles 1014135 cycles 1.00
ML-DSA-87 verify 403785 cycles 404031 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 43290 cycles 43407 cycles 1.00
ML-DSA-44 sign 130107 cycles 130440 cycles 1.00
ML-DSA-44 verify 45233 cycles 45231 cycles 1.00
ML-DSA-65 keypair 75800 cycles 75611 cycles 1.00
ML-DSA-65 sign 213743 cycles 213750 cycles 1.00
ML-DSA-65 verify 74411 cycles 74462 cycles 1.00
ML-DSA-87 keypair 122926 cycles 122917 cycles 1.00
ML-DSA-87 sign 270981 cycles 271210 cycles 1.00
ML-DSA-87 verify 120778 cycles 120613 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 46735 cycles 46709 cycles 1.00
ML-DSA-44 sign 139947 cycles 139435 cycles 1.00
ML-DSA-44 verify 49469 cycles 49461 cycles 1.00
ML-DSA-65 keypair 82621 cycles 81967 cycles 1.01
ML-DSA-65 sign 227154 cycles 226735 cycles 1.00
ML-DSA-65 verify 81978 cycles 82665 cycles 0.99
ML-DSA-87 keypair 130551 cycles 129492 cycles 1.01
ML-DSA-87 sign 279924 cycles 280340 cycles 1.00
ML-DSA-87 verify 128506 cycles 128405 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i) (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 91704 cycles 91774 cycles 1.00
ML-DSA-44 sign 351502 cycles 351709 cycles 1.00
ML-DSA-44 verify 99387 cycles 99709 cycles 1.00
ML-DSA-65 keypair 154237 cycles 154289 cycles 1.00
ML-DSA-65 sign 571927 cycles 570738 cycles 1.00
ML-DSA-65 verify 160351 cycles 160181 cycles 1.00
ML-DSA-87 keypair 255233 cycles 255166 cycles 1.00
ML-DSA-87 sign 720046 cycles 720821 cycles 1.00
ML-DSA-87 verify 264335 cycles 263865 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 112198 cycles 112399 cycles 1.00
ML-DSA-44 sign 353753 cycles 353820 cycles 1.00
ML-DSA-44 verify 117415 cycles 117171 cycles 1.00
ML-DSA-65 keypair 194605 cycles 194786 cycles 1.00
ML-DSA-65 sign 584072 cycles 584013 cycles 1.00
ML-DSA-65 verify 193382 cycles 193015 cycles 1.00
ML-DSA-87 keypair 320883 cycles 320852 cycles 1.00
ML-DSA-87 sign 747307 cycles 747201 cycles 1.00
ML-DSA-87 verify 318179 cycles 318645 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 52008 cycles 51879 cycles 1.00
ML-DSA-44 sign 154462 cycles 155064 cycles 1.00
ML-DSA-44 verify 54166 cycles 54259 cycles 1.00
ML-DSA-65 keypair 89427 cycles 89685 cycles 1.00
ML-DSA-65 sign 253071 cycles 254754 cycles 0.99
ML-DSA-65 verify 89176 cycles 89439 cycles 1.00
ML-DSA-87 keypair 143367 cycles 142352 cycles 1.01
ML-DSA-87 sign 310438 cycles 311341 cycles 1.00
ML-DSA-87 verify 138807 cycles 139359 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a) (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 118639 cycles 118234 cycles 1.00
ML-DSA-44 sign 458559 cycles 458712 cycles 1.00
ML-DSA-44 verify 130837 cycles 131121 cycles 1.00
ML-DSA-65 keypair 201470 cycles 200822 cycles 1.00
ML-DSA-65 sign 743521 cycles 747583 cycles 0.99
ML-DSA-65 verify 209873 cycles 209481 cycles 1.00
ML-DSA-87 keypair 331170 cycles 332858 cycles 0.99
ML-DSA-87 sign 935608 cycles 936652 cycles 1.00
ML-DSA-87 verify 343114 cycles 343994 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a) (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 133187 cycles 134282 cycles 0.99
ML-DSA-44 sign 518022 cycles 520818 cycles 0.99
ML-DSA-44 verify 146634 cycles 147777 cycles 0.99
ML-DSA-65 keypair 224332 cycles 224694 cycles 1.00
ML-DSA-65 sign 843960 cycles 843127 cycles 1.00
ML-DSA-65 verify 234144 cycles 233924 cycles 1.00
ML-DSA-87 keypair 367057 cycles 367125 cycles 1.00
ML-DSA-87 sign 1058600 cycles 1057892 cycles 1.00
ML-DSA-87 verify 380447 cycles 380252 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2 (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 212129 cycles 212335 cycles 1.00
ML-DSA-44 sign 760693 cycles 761210 cycles 1.00
ML-DSA-44 verify 229484 cycles 229974 cycles 1.00
ML-DSA-65 keypair 375440 cycles 376324 cycles 1.00
ML-DSA-65 sign 1248106 cycles 1248564 cycles 1.00
ML-DSA-65 verify 371579 cycles 372377 cycles 1.00
ML-DSA-87 keypair 600041 cycles 601386 cycles 1.00
ML-DSA-87 sign 1586041 cycles 1607791 cycles 0.99
ML-DSA-87 verify 615917 cycles 617602 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 61706 cycles 61819 cycles 1.00
ML-DSA-44 sign 187966 cycles 188851 cycles 1.00
ML-DSA-44 verify 66228 cycles 66399 cycles 1.00
ML-DSA-65 keypair 109427 cycles 110427 cycles 0.99
ML-DSA-65 sign 311800 cycles 313769 cycles 0.99
ML-DSA-65 verify 109443 cycles 111323 cycles 0.98
ML-DSA-87 keypair 170673 cycles 173116 cycles 0.99
ML-DSA-87 sign 379405 cycles 385429 cycles 0.98
ML-DSA-87 verify 170627 cycles 174422 cycles 0.98

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i) (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 154055 cycles 154452 cycles 1.00
ML-DSA-44 sign 587101 cycles 588162 cycles 1.00
ML-DSA-44 verify 169059 cycles 168998 cycles 1.00
ML-DSA-65 keypair 261730 cycles 262562 cycles 1.00
ML-DSA-65 sign 964081 cycles 965903 cycles 1.00
ML-DSA-65 verify 271336 cycles 272372 cycles 1.00
ML-DSA-87 keypair 431609 cycles 431929 cycles 1.00
ML-DSA-87 sign 1212961 cycles 1211899 cycles 1.00
ML-DSA-87 verify 447991 cycles 447459 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-M55 (NUCLEO-N657X0-Q) benchmarks (opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 1690331 cycles 2328440 cycles 0.73
ML-DSA-44 sign 13600033 cycles 17851737 cycles 0.76
ML-DSA-44 verify 1806518 cycles 2398876 cycles 0.75
ML-DSA-65 keypair 2895358 cycles 4025122 cycles 0.72
ML-DSA-65 sign 11228232 cycles 14793522 cycles 0.76
ML-DSA-65 verify 2973684 cycles 4012792 cycles 0.74
ML-DSA-87 keypair 4888795 cycles 6864887 cycles 0.71
ML-DSA-87 sign 18455928 cycles 24785157 cycles 0.74
ML-DSA-87 verify 5009804 cycles 6864412 cycles 0.73

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-M55 (NUCLEO-N657X0-Q) benchmarks (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 2328440 cycles 2328440 cycles 1
ML-DSA-44 sign 17851737 cycles 17851737 cycles 1
ML-DSA-44 verify 2398876 cycles 2398876 cycles 1
ML-DSA-65 keypair 4025122 cycles 4025122 cycles 1
ML-DSA-65 sign 14793522 cycles 14793522 cycles 1
ML-DSA-65 verify 4012792 cycles 4012792 cycles 1
ML-DSA-87 keypair 6864887 cycles 6864887 cycles 1
ML-DSA-87 sign 24785157 cycles 24785157 cycles 1
ML-DSA-87 verify 6864412 cycles 6864412 cycles 1

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 272258 cycles 271067 cycles 1.00
ML-DSA-44 sign 811182 cycles 812320 cycles 1.00
ML-DSA-44 verify 273722 cycles 273045 cycles 1.00
ML-DSA-65 keypair 467921 cycles 467726 cycles 1.00
ML-DSA-65 sign 1374610 cycles 1341160 cycles 1.02
ML-DSA-65 verify 452328 cycles 454368 cycles 1.00
ML-DSA-87 keypair 794593 cycles 803830 cycles 0.99
ML-DSA-87 sign 1852551 cycles 1831323 cycles 1.01
ML-DSA-87 verify 781142 cycles 778295 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 465428 cycles 465494 cycles 1.00
ML-DSA-44 sign 2134682 cycles 2144256 cycles 1.00
ML-DSA-44 verify 556859 cycles 559936 cycles 0.99
ML-DSA-65 keypair 784546 cycles 783372 cycles 1.00
ML-DSA-65 sign 3502813 cycles 3496274 cycles 1.00
ML-DSA-65 verify 867806 cycles 869188 cycles 1.00
ML-DSA-87 keypair 1265637 cycles 1266792 cycles 1.00
ML-DSA-87 sign 4348278 cycles 4317269 cycles 1.01
ML-DSA-87 verify 1387880 cycles 1394437 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 218060 cycles 216344 cycles 1.01
ML-DSA-44 sign 597336 cycles 593807 cycles 1.01
ML-DSA-44 verify 217728 cycles 216553 cycles 1.01
ML-DSA-65 keypair 380946 cycles 378995 cycles 1.01
ML-DSA-65 sign 983605 cycles 982529 cycles 1.00
ML-DSA-65 verify 363319 cycles 363919 cycles 1.00
ML-DSA-87 keypair 636094 cycles 636641 cycles 1.00
ML-DSA-87 sign 1312206 cycles 1318452 cycles 1.00
ML-DSA-87 verify 617779 cycles 619777 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 301299 cycles 301210 cycles 1.00
ML-DSA-44 sign 1140771 cycles 1141652 cycles 1.00
ML-DSA-44 verify 333412 cycles 330898 cycles 1.01
ML-DSA-65 keypair 550692 cycles 543839 cycles 1.01
ML-DSA-65 sign 1876845 cycles 1856990 cycles 1.01
ML-DSA-65 verify 533547 cycles 525059 cycles 1.02
ML-DSA-87 keypair 845041 cycles 846962 cycles 1.00
ML-DSA-87 sign 2351108 cycles 2362668 cycles 1.00
ML-DSA-87 verify 874370 cycles 876584 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SpacemiT K1 8 (Banana Pi F3) benchmarks (no-opt)

Details
Benchmark suite Current: 91039f7 Previous: 27aebe7 Ratio
ML-DSA-44 keypair 760366 cycles 760367 cycles 1.00
ML-DSA-44 sign 3142358 cycles 3140398 cycles 1.00
ML-DSA-44 verify 859639 cycles 859497 cycles 1.00
ML-DSA-65 keypair 1287922 cycles 1289584 cycles 1.00
ML-DSA-65 sign 5081104 cycles 5091406 cycles 1.00
ML-DSA-65 verify 1367411 cycles 1368454 cycles 1.00
ML-DSA-87 keypair 2109891 cycles 2109786 cycles 1.00
ML-DSA-87 sign 6360762 cycles 6373090 cycles 1.00
ML-DSA-87 verify 2224537 cycles 2226888 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@bremoran

bremoran commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Added metadata YAML to the input source file so that slothy will preserve it and autogen will pass. This is a comment-only change and benchmarks do not need to be re-run.

@bremoran
bremoran marked this pull request as ready for review August 6, 2026 08:14

@mkannwischer mkannwischer left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @bremoran.

I have a number of comments. Please look at the LICENSE first as that will likely block this PR.

Comment thread mldsa/src/fips202/keccakf1600.c Outdated
Comment on lines 43 to 51
#if defined(MLD_USE_NATIVE_FIPS202_X1_EXTRACT_BYTES)
if (mld_keccakf1600_extract_bytes_x1_native(state, data, offset, length) ==
MLD_NATIVE_FUNC_SUCCESS)
{
return;
}
#endif /* MLD_USE_NATIVE_FIPS202_X1_EXTRACT_BYTES */
#if defined(MLD_SYS_LITTLE_ENDIAN)
uint8_t *state_ptr = (uint8_t *)state + offset;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This breaks C90 on platforms impkementing exract bytes due to the state_ptr declaration a little lower. that needs to be fixed. Also maybe worth switching to C90 to catch this in the future (unless this leads to more problems, then please open an issue).

Comment on lines +26 to +54
/* The Adomnicai x1 core consumes 24 bit-interleaved round constants followed
* by a 0xff loop terminator. Keep this table private to this backend. */
static MLD_ALIGN const uint32_t mld_keccakf1600_round_constants_x1[49] = {
0x00000001, 0x00000000, /* RC0 */
0x00000000, 0x00000089, /* RC1 */
0x00000000, 0x8000008b, /* RC2 */
0x00000000, 0x80008080, /* RC3 */
0x00000001, 0x0000008b, /* RC4 */
0x00000001, 0x00008000, /* RC5 */
0x00000001, 0x80008088, /* RC6 */
0x00000001, 0x80000082, /* RC7 */
0x00000000, 0x0000000b, /* RC8 */
0x00000000, 0x0000000a, /* RC9 */
0x00000001, 0x00008082, /* RC10 */
0x00000000, 0x00008003, /* RC11 */
0x00000001, 0x0000808b, /* RC12 */
0x00000001, 0x8000000b, /* RC13 */
0x00000001, 0x8000008a, /* RC14 */
0x00000001, 0x80000081, /* RC15 */
0x00000000, 0x80000081, /* RC16 */
0x00000000, 0x80000008, /* RC17 */
0x00000000, 0x00000083, /* RC18 */
0x00000000, 0x80008003, /* RC19 */
0x00000001, 0x80008088, /* RC20 */
0x00000000, 0x80000088, /* RC21 */
0x00000001, 0x00008000, /* RC22 */
0x00000000, 0x80008082, /* RC23 */
0x000000ff, /* loop terminator */
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please auto generate in autogen this in a similar way it is autogenerated for other platforms.

unsigned offset,
unsigned length);

/* The Adomnicai x1 core consumes 24 bit-interleaved round constants followed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's acknowledge the source of this implementation once in the assembly and then keep all other comments general, i.e., do not talk about the Adomnicai x1 core - it only leads to confusion later on.

*/

/*
* This helper is derived from the public-domain XKCP implementation and the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Be more specific here: CC0. public domain means different things for different people.


/*
* This helper is derived from the public-domain XKCP implementation and the
* Armv7-M optimizations described in [ADOMNICAI23]. See the backend README,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please cite the implementation directly

Comment thread test/zephyr/platform.mk
Comment on lines -36 to +37
ZEPHYR_FIPS202_BACKEND_mps3-an547 := fips202/native/armv81m/mve.h
ZEPHYR_FIPS202_BACKEND_nucleo-n657x0-q := fips202/native/armv81m/mve.h
ZEPHYR_FIPS202_BACKEND_mps3-an547 := fips202/native/armv81m/mve_x1.h
ZEPHYR_FIPS202_BACKEND_nucleo-n657x0-q := fips202/native/armv81m/mve_x1.h

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This swaps out the backend. Does this introduce a test gap as the x4 Keccak is no longer tested? We should not be regressing test coverage, so we need to at least add a x4 unit test to CI.


#define mld_keccak_f1600_x1_native_impl \
MLD_NAMESPACE(keccak_f1600_x1_native_impl)
int mld_keccak_f1600_x1_native_impl(uint64_t *state);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These should be marked MLD_INTERNAL_API

Comment thread test/zephyr/platform.mk Outdated
QEMU_TIMEOUT ?= 300
export QEMU_TIMEOUT

# Native backends are an OPT=1 feature (an547 builds the Armv8.1-M MVE backend).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stale

Comment thread test/src/test_unit.c
@@ -65,48 +65,75 @@ unsigned int mld_rej_eta_c(int32_t *a, unsigned int target, unsigned int offset,
void mld_keccakf1600_permute_c(uint64_t *state);

#if defined(MLD_USE_NATIVE_FIPS202_X1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CI does not run unit tests for Armv8.1-M. Please enable it.

Comment thread LICENSE
The Armv8.1-M Keccak implementation contains portions adapted from SLOTHY
examples distributed under the MIT license. The applicable SLOTHY copyright
notices are included below.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Which portions are MIT licensed? I did not see any file in this PR being MIT-only licensed and we wouldn't be able to take it for anything under mldsa/. Maybe this is an outdated notice that can be removed?

The lazy/eager polyvector unit test allocates both ML-DSA-87
representations and their scratch space through MLD_ALLOC. The default
MLD_ALLOC implementation expands to aligned automatic arrays, so this
single test exceeds the Zephyr test thread stack on the Cortex-M55 test
configuration.

Use test-local, aligned static buffers for this workspace. The
TEST_STATIC_ALLOC and TEST_STATIC_FREE helpers are deliberately scoped
to test_unit.c: they preserve cleanup by zeroizing every buffer and
clearing its pointer, without changing the allocator used by production
code or other tests.

Only test/src/test_unit.c changes. The test inputs, comparisons, and
coverage remain the same; this commit changes where its temporary
workspace is stored.

Signed-off-by: Brendan Moran <brendan.moran@arm.com>
Generated ABI checks normally fill every assembly argument buffer with
random bytes. That is unsuitable for interfaces containing control data.
The Armv8.1-M Keccak x1 permutation consumes 49 round-constant words and
requires the final word to be 0x000000ff as a loop terminator. Leaving
that word random can make the checker read beyond the supplied buffer.

Add an optional test_bytes mapping to buffer entries in the assembly ABI
YAML. scripts/autogen validates that offsets are integers within the
buffer and that values are bytes, sorts the overrides, and emits them
after randombytes initializes the rest of the buffer. Existing ABI
metadata without test_bytes retains its current behaviour.

Document the new metadata in test/abicheck/README.md. The generic
facility is introduced here before its Keccak x1 consumer so the
following backend commit contains only feature-specific metadata and
generated checks.

Signed-off-by: Brendan Moran <brendan.moran@arm.com>
Add a scalar Armv8.1-M Keccak backend using the Armv7-M bit-interleaved representation and an M7-scheduled permutation. Keep states interleaved across permutations and convert only the lanes crossing the byte interface.

Include the clean source, final SLOTHY scheduling driver, generated production sources, x1 xor/extract helpers, native integration, focused unit coverage, and AAPCS32 ABI checks. Preserve the assembly metadata and integration guards through regeneration, and keep the existing x4 implementation unchanged.

Document the XKCP, Adomnicai, and SLOTHY provenance and licensing.

Signed-off-by: Brendan Moran <brendan.moran@arm.com>
Run the full functional, KAT, ACVP, and unit suites for the retained x4 backend on mps3-an547 instead of relying on the new x1 default. Keep the ordinary unit suite enabled for all three ML-DSA parameter sets and reduce only its QEMU randomized budgets to fit the 300-second execution deadline.

Until the upstream unaligned-access fallback is available, constrain only the Armv8.1-M x4 xor/extract test buffers and offsets to the implementation's supported alignment. The x1 path and other backends retain arbitrary-alignment coverage.

Signed-off-by: Brendan Moran <brendan.moran@arm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Armv8.1-M: add regular and SLOTHY-optimized Armv7-M Keccak x1

3 participants