A reproducible, end-to-end measurement of full-integer INT8 quantization for a small CNN running under TensorFlow Lite Micro on a Cortex-M4 class MCU — trained, quantized, compiled into Zephyr firmware, and executed in the Renode simulator.
This is an engineering benchmark, not novel research. The model is a standard small CNN on MNIST and the technique is off-the-shelf post-training quantization. What the repo provides is the measurement: honest, reproducible numbers for the trade-off, produced by scripts you can re-run.
| Metric | float32 | INT8 | Improvement |
|---|---|---|---|
| TFLite model size | 111,172 B | 33,360 B | 3.33x |
| Firmware flash (linker) | 193,256 B | 115,432 B | 1.67x |
| Static RAM (linker) | 46,624 B | 23,496 B | 1.98x |
| Tensor arena used | 33,280 B | 10,136 B | 3.28x |
| Inference latency (sim) | 514,422 us | 281,884 us | 1.82x |
| Test accuracy (2,000 images) | 98.30% | 98.20% | -0.10 pp |
Both variants classify the same held-out MNIST digit correctly on-device
(predicted=7, expected=7), so these are timings of real inference, not of a stub.
INT8 bought roughly 3.3x smaller weights, 2x less RAM and 1.8x faster inference for 0.1 percentage points of accuracy.
- Simulated, not silicon. Everything runs in Renode on a simulated nRF52840 (the MCU in the Arduino Nano 33 BLE Sense). Flash and static RAM come from the linker and are exact. Latency is cycle-approximate — Renode models the RTC, not true instruction timing. Treat the ratio as meaningful and the absolute milliseconds as indicative until re-measured on hardware.
- Reference kernels, no CMSIS-NN. TFLM's portable kernels are used, not the ARM CMSIS-NN optimized ones. That is why absolute latency is high (~0.3-0.5 s/inference). CMSIS-NN would cut both figures substantially and would likely widen the INT8 margin, since it optimizes integer paths hardest.
- Accuracy is measured on the TFLite interpreter, not on the Keras model, on 2,000 held-out test images — so the INT8 figure is what the MCU actually computes.
- Static RAM is a fair comparison because the harness sizes the tensor arena to what each model really needs (see Two-pass measurement below) rather than padding both to the same generous constant.
- Single fixed seed (1337), single architecture. No repeated runs or confidence intervals — Renode is deterministic, and re-running gave 281,884 vs 281,875 us across different iteration counts, but that is reproducibility, not statistical power.
train.py small CNN on MNIST (26,698 params), fixed seed
|
convert.py -> model_f32.tflite (float32 baseline)
-> model_int8.tflite (full-integer: int8 weights AND activations,
int8 in/out, 500-image representative set)
-> accuracy of each, measured on the TFLite interpreter
-> C arrays for embedding
|
emit_input.py one real MNIST digit as a C array, so on-device inference is genuine
|
bench.sh Zephyr build -> Renode -> parse UART -> results/<variant>.json
|
summarize.py results/table.md (the table above is generated, not hand-typed)
Static RAM only means something if the tensor arena is sized honestly, so bench.sh
builds each variant twice:
- Build with a deliberately generous arena, run it, and read back
interpreter->arena_used_bytes()— what the interpreter genuinely required. - Rebuild with the arena set to that value, and take flash/RAM from the linker.
Without this, both variants would report whatever constant was hard-coded, and the RAM column would be meaningless.
Requires a Zephyr workspace + SDK, west, and Renode on PATH.
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -r requirements.txt
.venv/bin/python scripts/train.py --epochs 8 # ~15 s on an M-series Mac
.venv/bin/python scripts/convert.py # exports both variants + accuracy
.venv/bin/python scripts/emit_input.py
./scripts/bench.sh int8 results/int8.json
RUNFOR=60 ./scripts/bench.sh f32 results/f32.json # float32 needs more virtual time
.venv/bin/python scripts/summarize.pyOverride ZEPHYR_BASE, ZEPHYR_SDK_INSTALL_DIR, WEST, RENODE_BIN if yours live elsewhere.
scripts/ train.py convert.py emit_input.py bench.sh summarize.py
zephyr/
src/ main.c main_functions.cpp test_input.cpp (the TFLM runner)
models/ model_f32.cpp model_int8.cpp (generated C arrays)
results/ f32.json int8.json table.md (measured output)
- Re-measure latency on a physical Nano 33 BLE Sense; the simulator figure needs confirming.
- Enable CMSIS-NN and re-run — the most interesting missing comparison.
- MNIST is a placeholder task. The same harness applied to a sensor workload (keyword spotting, accelerometer activity, anomaly detection) would say more.
- Only post-training quantization. Quantization-aware training would likely close the 0.1 pp gap and is the obvious follow-up.
Apache-2.0. The TFLite Micro runner in zephyr/src/ is written against the TensorFlow
Lite Micro API; TFLM itself is vendored by Zephyr, not by this repo.