Skip to content

About

Reproducible INT8-vs-float32 benchmark for TensorFlow Lite Micro on a simulated nRF52840 (Zephyr + Renode): 3.3x smaller, 1.8x faster, -0.1pp accuracy — with exact linker footprints and honest caveats.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

INT8 quantization on an nRF52840: what it actually costs and saves

A reproducible, end-to-end measurement of full-integer INT8 quantization for a small CNN running under TensorFlow Lite Micro on a Cortex-M4 class MCU — trained, quantized, compiled into Zephyr firmware, and executed in the Renode simulator.

This is an engineering benchmark, not novel research. The model is a standard small CNN on MNIST and the technique is off-the-shelf post-training quantization. What the repo provides is the measurement: honest, reproducible numbers for the trade-off, produced by scripts you can re-run.

Results

Metric float32 INT8 Improvement
TFLite model size 111,172 B 33,360 B 3.33x
Firmware flash (linker) 193,256 B 115,432 B 1.67x
Static RAM (linker) 46,624 B 23,496 B 1.98x
Tensor arena used 33,280 B 10,136 B 3.28x
Inference latency (sim) 514,422 us 281,884 us 1.82x
Test accuracy (2,000 images) 98.30% 98.20% -0.10 pp

Both variants classify the same held-out MNIST digit correctly on-device (predicted=7, expected=7), so these are timings of real inference, not of a stub.

INT8 bought roughly 3.3x smaller weights, 2x less RAM and 1.8x faster inference for 0.1 percentage points of accuracy.

Read the numbers with these caveats

  • Simulated, not silicon. Everything runs in Renode on a simulated nRF52840 (the MCU in the Arduino Nano 33 BLE Sense). Flash and static RAM come from the linker and are exact. Latency is cycle-approximate — Renode models the RTC, not true instruction timing. Treat the ratio as meaningful and the absolute milliseconds as indicative until re-measured on hardware.
  • Reference kernels, no CMSIS-NN. TFLM's portable kernels are used, not the ARM CMSIS-NN optimized ones. That is why absolute latency is high (~0.3-0.5 s/inference). CMSIS-NN would cut both figures substantially and would likely widen the INT8 margin, since it optimizes integer paths hardest.
  • Accuracy is measured on the TFLite interpreter, not on the Keras model, on 2,000 held-out test images — so the INT8 figure is what the MCU actually computes.
  • Static RAM is a fair comparison because the harness sizes the tensor arena to what each model really needs (see Two-pass measurement below) rather than padding both to the same generous constant.
  • Single fixed seed (1337), single architecture. No repeated runs or confidence intervals — Renode is deterministic, and re-running gave 281,884 vs 281,875 us across different iteration counts, but that is reproducibility, not statistical power.

Method

train.py            small CNN on MNIST (26,698 params), fixed seed
   |
convert.py          -> model_f32.tflite   (float32 baseline)
                    -> model_int8.tflite  (full-integer: int8 weights AND activations,
                                           int8 in/out, 500-image representative set)
                    -> accuracy of each, measured on the TFLite interpreter
                    -> C arrays for embedding
   |
emit_input.py       one real MNIST digit as a C array, so on-device inference is genuine
   |
bench.sh            Zephyr build -> Renode -> parse UART -> results/<variant>.json
   |
summarize.py        results/table.md  (the table above is generated, not hand-typed)

Two-pass measurement

Static RAM only means something if the tensor arena is sized honestly, so bench.sh builds each variant twice:

  1. Build with a deliberately generous arena, run it, and read back interpreter->arena_used_bytes() — what the interpreter genuinely required.
  2. Rebuild with the arena set to that value, and take flash/RAM from the linker.

Without this, both variants would report whatever constant was hard-coded, and the RAM column would be meaningless.

Reproducing

Requires a Zephyr workspace + SDK, west, and Renode on PATH.

uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -r requirements.txt

.venv/bin/python scripts/train.py --epochs 8      # ~15 s on an M-series Mac
.venv/bin/python scripts/convert.py               # exports both variants + accuracy
.venv/bin/python scripts/emit_input.py

./scripts/bench.sh int8  results/int8.json
RUNFOR=60 ./scripts/bench.sh f32 results/f32.json  # float32 needs more virtual time

.venv/bin/python scripts/summarize.py

Override ZEPHYR_BASE, ZEPHYR_SDK_INSTALL_DIR, WEST, RENODE_BIN if yours live elsewhere.

Layout

scripts/     train.py  convert.py  emit_input.py  bench.sh  summarize.py
zephyr/
  src/       main.c  main_functions.cpp  test_input.cpp   (the TFLM runner)
  models/    model_f32.cpp  model_int8.cpp               (generated C arrays)
results/     f32.json  int8.json  table.md               (measured output)

Known limitations / next steps

  • Re-measure latency on a physical Nano 33 BLE Sense; the simulator figure needs confirming.
  • Enable CMSIS-NN and re-run — the most interesting missing comparison.
  • MNIST is a placeholder task. The same harness applied to a sensor workload (keyword spotting, accelerometer activity, anomaly detection) would say more.
  • Only post-training quantization. Quantization-aware training would likely close the 0.1 pp gap and is the obvious follow-up.

Licence

Apache-2.0. The TFLite Micro runner in zephyr/src/ is written against the TensorFlow Lite Micro API; TFLM itself is vendored by Zephyr, not by this repo.

About

Reproducible INT8-vs-float32 benchmark for TensorFlow Lite Micro on a simulated nRF52840 (Zephyr + Renode): 3.3x smaller, 1.8x faster, -0.1pp accuracy — with exact linker footprints and honest caveats.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages