Deployment¶
Deploy a quantized TFLite model to the STM32N6570-DK development board using ST's X-CUBE-AI toolchain.
Prerequisites¶
| Tool | Version | Download |
|---|---|---|
| X-CUBE-AI | 10.2.0+ | ST website |
| STM32CubeProgrammer | 2.20+ | ST website |
| STM32CubeIDE | 1.19+ | ST website |
| ARM GNU Toolchain | 14.3+ | ARM Developer |
Overview¶
The deployment pipeline has three stages:
flowchart LR
A[".tflite\nquantized model"] --> B["stedgeai generate\nN6-optimized binary"]
B --> C["n6_loader.py\nserial flash to board"]
C --> D["stedgeai validate\non-device inference"]
D --> E["Validation report\ncosine sim + latency"]
(Image source: STM32ai)
Step 1: Install X-CUBE-AI¶
unzip x-cube-ai-linux-v10.2.0.zip X-CUBE-AI.10.2.0
cd X-CUBE-AI.10.2.0
unzip stedgeai-linux-10.2.0.zip
Directory structure after extraction:
X-CUBE-AI.10.2.0/
├── Utilities/
│ └── linux/
│ └── stedgeai # CLI tool
├── Middlewares/
└── Projects/
Step 2: Install ARM GNU Toolchain¶
wget https://developer.arm.com/-/media/Files/downloads/gnu/14.3.rel1/binrel/arm-gnu-toolchain-14.3.rel1-x86_64-arm-none-eabi.tar.xz
tar xf arm-gnu-toolchain-14.3.rel1-x86_64-arm-none-eabi.tar.xz
export PATH=$PWD/arm-gnu-toolchain-14.3.rel1-x86_64-arm-none-eabi/bin:$PATH
Verify:
Step 3: Install STM32CubeProgrammer¶
Download and run the installer:
Add to PATH and configure permissions:
export PATH=$PATH:/path/to/STM32Cube/STM32CubeProgrammer/bin
sudo usermod -aG plugdev $USER
sudo usermod -aG dialout $USER
# Install udev rules
sudo cp /path/to/STM32CubeProgrammer/Drivers/rules/*.* /etc/udev/rules.d/
sudo udevadm control --reload-rules && sudo udevadm trigger
Unplug and replug the board (or reboot) to apply the new rules.
Verify:
Step 4: Generate model files¶
Navigate to the X-CUBE-AI utilities directory and run:
./stedgeai generate \
--model /path/to/checkpoints/my_model_quantized.tflite \
--target stm32n6 \
--st-neural-art \
--output /path/to/birdnet-stm32/validation/st_ai_output \
--workspace /path/to/birdnet-stm32/validation/st_ai_ws \
--verbose
Analyze first
Run stedgeai analyze instead of generate to get detailed model metrics
(size, memory, per-layer info) without generating output files. Always
analyze new model architectures to verify N6 NPU operator compatibility.
The output includes network_generate_report.txt with model size and compute
requirements.
Step 5: Configure and flash the board¶
Set board to DEV mode¶
- Disconnect the board from USB.
- Set BOOT0 to right.
- Set BOOT1 to left.
- Set JP2 to position 1-2.
- Reconnect the board.
(Image source: ST Community)
Create configuration files¶
Copy the example config and fill in your local paths:
Edit config.json with your machine-local paths:
{
"compiler_type": "gcc",
"cubeide_path": "/path/to/stm32cubeide",
"x_cube_ai_path": "/path/to/X-CUBE-AI.10.2.0",
"model_path": "checkpoints/best_model_quantized.tflite",
"output_dir": "validation/st_ai_output",
"workspace_dir": "validation/st_ai_ws",
"n6_loader_config": "config_n6l.json"
}
Create config_n6l.json in the project root (required by ST's n6_loader):
{
"network.c": "/path/to/birdnet-stm32/validation/st_ai_output/network.c",
"project_path": "/path/to/X-CUBE-AI.10.2.0/Projects/STM32N6570-DK/Applications/NPU_Validation",
"project_build_conf": "N6-DK",
"skip_external_flash_programming": false,
"skip_ram_data_programming": false,
"objcopy_binary_path": "/usr/bin/arm-none-eabi-objcopy"
}
Warning
Both config files contain machine-local paths. They are listed in
.gitignore — do not commit them. Use config.example.json as a
reference template.
Run the full deploy pipeline¶
The CLI reads all paths from config.json and runs generate → flash → validate:
You can override any path via CLI arguments:
Or via environment variables:
Priority order: CLI arguments > environment variables > config.json values.
Verify the board is connected:
You may need serial port permissions:
Step 6: Validate on-device¶
The deploy command runs validation automatically. To run validation separately
with additional options (e.g., --valinput for specific test data):
/path/to/X-CUBE-AI.10.2.0/Utilities/linux/stedgeai validate \
--model checkpoints/my_model_quantized.tflite \
--target stm32n6 \
--mode target \
--desc serial:921600 \
--output /path/to/birdnet-stm32/validation/st_ai_output \
--workspace /path/to/birdnet-stm32/validation/st_ai_ws \
--valinput /path/to/checkpoints/my_model_quantized_validation_data.npz \
--classifier \
--verbose
The validation runs inference on the physical board and compares results to the
reference model. Results are saved to network_validate_report.txt in the
output directory.
Read the cross-accuracy block in that report, not just the exit status:
cos is the number that matters — 1.000000 means the device reproduces the
host exactly. A model can flash, run, report plausible timings and still be
numerically wrong, so this is the only check that proves the deployed artifact
computes what the host does. Compare m_outputs (host reference) against
c_outputs (target): if the target's range is visibly compressed relative to
the reference, the graph is not computing correctly regardless of timing.
Check mae as well. Cosine is blind to a uniform gain error or a constant
offset, so a layer can score cos 0.95 while being clearly wrong; the mean
absolute error is not. A correct INT8 model stays under one output LSB
(1/256 ≈ 0.0039) — the fixed 100-class raw model measures 0.0017. Prefer it
to l2r, which divides by the reference norm and so reads high on sparse
multi-label outputs even when the device is right (0.108 for that same model).
To localise a mismatch, cut the graph at a node index and validate the partial
model — --cut-output-layers N (and --cut-input-layers N to start there
instead). Bisecting on the first node whose cos falls away names the layer.
n6_loader.py hangs with ST's stock NPU_Validation app
--mode target needs ST's stock validation app on the board, and flashing
it with n6_loader.py hangs at "Loading internal memories & Running the
program".
RISAF_Config() programs RISAF4_S/RISAF5_S, the NPU master ports, and
its own comment notes that an IP must be clocked before its RISAF is set.
ST's stock Core/Src/main.c calls RISAF_Config() before anything clocks
the NPU, so the write stalls the bus, the app never reaches
aiValidationInit(), and the temporary GDB breakpoint n6_loader sets
there never hits — its continue then waits forever.
Fix it by calling NPU_Config() immediately before RISAF_Config() in
ST's Core/Src/main.c (keep a backup; it is a vendor file). This project's
own firmware already does this and is unaffected.
Two related traps:
ST-LINK_gdbserverreports "Target unknown error 32" if the probe has not been released after a CubeProgrammer session. Connect once withSTM32_Programmer_CLI -c port=SWD mode=UR, wait a few seconds, then start the gdb server; retry if needed.pkill -f ST-LINK_gdbserverkills the shell that runs it, because the pattern matches that shell's own command line. Kill by PID instead.
Demo application¶
The demo application is under development. The planned pipeline:
- Record audio using the on-board microphone.
- Run FFT on 512-sample frames, accumulating into a ring buffer.
- Run inference every second on the last 3 seconds of audio.
- Map prediction scores to labels using
labels.txt. - Log top-5 predictions to the serial console.
Board test¶
The board-test command runs a standalone inference test on the STM32N6570-DK.
The firmware reads WAV files from the SD card, computes the STFT on the
Cortex-M55, runs the model on the NPU, and streams results over UART. This
verifies the entire on-device pipeline end-to-end.
SD card preparation¶
- Format a micro-SD card as FAT32.
- Create an
audio/directory at the root. - Copy mono or stereo 16-bit PCM
.wavfiles intoaudio/. Each file should be at least as long as the model's chunk duration (default 3 s). The sample rate must match the model's (printed in_model_config.json). - Insert the card into the STM32N6570-DK slot.
Board-test arguments¶
| Argument | Default | Description |
|---|---|---|
--model_path |
(from config) | Path to quantized .tflite model |
--model_config |
(inferred) | Path to _model_config.json |
--labels |
(inferred) | Path to _labels.txt |
--serial_port |
/dev/ttyACM0 |
Serial port for UART capture |
--top_k |
5 | Top-K predictions per file |
--score_threshold |
0.01 | Minimum score to display |
--config |
config.json |
Deploy configuration JSON |
--timeout |
300 | Max seconds to wait for firmware response |
--host_audio_dir |
None | Local copy of the SD card's audio/ folder; enables the host x board parity check |
--parity_tolerance |
0.05 | Top-1 tie margin and borderline margin around the detection threshold |
--detection_threshold |
0.5 | Score at which a detection is reported; board and host must agree |
--save_results |
None | Save results summary to a CSV file (with host columns when parity is checked) |
Host x board parity¶
The board test is only evidence if the board computes what the host computes.
Pass --host_audio_dir a local copy of the files on the SD card and the same
.tflite is run on the host over the same files, through the host's own
evaluation preprocessing, on the first chunk of each file — the chunk the
firmware reads. Every file is then checked:
- the board's top-1 must be the host's top-1, or a label the host scores within
--parity_toleranceof its own top-1 (reported as a tie); - board and host must make the same call at
--detection_threshold(0.5), unless the host's score is within--parity_toleranceof it — NPU rounding can tip a score that close, so the file passes but is flaggedborderline.
Score differences larger than the tolerance that change neither are flagged
drift, not failed. On the V12 raw model the largest was 0.054, on a
mid-range score where the sigmoid is steepest.
The command prints a per-file table and a PARITY PASS/PARITY FAIL line, and
exits nonzero on a failure. If a manifest.csv with file and true_species
columns sits in or beside the audio folder, it also counts correct top-1
detections for board and host.
Use audio the model classifies confidently, one file per species, with the vocalization in the first chunk. Background-only audio proves nothing: a broken model and a working one both return low scores on it, which is how the raw frontend's NPU defects went unnoticed.
Board test is standalone
The board-test command deploys real firmware that does all processing on
the board: read WAV from SD card → apply frontend-specific preprocessing
(raw normalization, hybrid STFT, or librosa STFT + mel) → run NPU
inference → stream results over UART. Do not precompute test inputs on the
host; that would bypass the integration path being tested.
Firmware documentation¶
For detailed documentation on the board firmware — hardware specs, build system, configuration, source module reference, UART protocol, and troubleshooting — see the Firmware section.