Implementation Notes¶
Design rationale for key architectural decisions in BirdNET-STM32.
Why DS-CNN?¶
Depthwise-separable convolutions (DS-CNN) are the backbone because:
- Parameter efficiency: a depthwise-separable block uses ~8-9× fewer parameters than a standard convolution with the same receptive field.
- NPU compatibility: the STM32N6 NPU natively supports
DepthwiseConv2DandConv2D(pointwise). No custom ops needed. - Proven track record: MobileNetV1/V2-style architectures are the de facto standard for on-device audio classification (Google's keyword spotting, ARM ML Zoo, etc.).
The 4-stage design with stride-2 downsampling gives a 16× spatial reduction, which is sufficient for mel spectrograms of typical size (64×256).
Plain DS blocks are the only option the builder offers. Inverted-residual and squeeze-and-excite variants were available once and were removed: their extra requantization boundaries put INT8 parity out of reach of the release gates. See Removed: inverted residual and squeeze-and-excite blocks.
Why PWL magnitude scaling?¶
| Scaling | Quantization behavior | N6 compatibility | Notes |
|---|---|---|---|
| PWL (piecewise-linear) | Excellent — depthwise conv + ReLU only | Full | Default |
| none | Exact — pass-through | Full | Ablation baseline |
PWL uses only operations that quantize cleanly to INT8, and its learned breakpoints adapt to the dataset's dynamic range during training. PCEN and dB were removed in 1.2.0: dB's log op creates exactly the wide dynamic range INT8 cannot hold, and PCEN was never used by a release.
Why float32 I/O?¶
Audio spectrograms are continuous-valued signals with meaningful precision at small magnitudes. Quantizing model inputs to INT8 would:
- Destroy quiet details: bird calls often have low-energy harmonics that fall below INT8 resolution.
- Waste quantization range: spectrogram values are not uniformly distributed — most energy concentrates in a few frequency bands.
- Complicate preprocessing: the STM32 firmware would need to quantize float STFT output to INT8 before feeding the NPU, adding complexity and latency.
The pipeline enforces float32 inputs and outputs with INT8 internal weights and activations. This is the standard approach for audio/speech models on edge devices.
N6 NPU operator coverage¶
The STM32N6 Neural-ART NPU supports a subset of TFLite operators. Verified compatible ops (as of X-CUBE-AI 10.2):
| Category | Supported operators |
|---|---|
| Convolution | Conv2D, DepthwiseConv2D |
| Normalization | BatchNormalization (fused into conv) |
| Activation | ReLU, ReLU6, Sigmoid |
| Pooling | GlobalAveragePooling2D, AveragePooling2D, MaxPooling2D |
| Arithmetic | Add, Multiply |
| Reshape | Reshape, Flatten |
| Linear | Dense (MatMul + BiasAdd) |
| Other | Concatenate, Pad |
Always verify with stedgeai
This table is a guideline. Always run stedgeai analyze on your TFLite
model before attempting deployment. Op support can change between
X-CUBE-AI versions.
Known unsupported ops¶
Softmax— useSigmoidfor multi-label classification insteadLayerNormalization— useBatchNormalizationGRU/LSTM— no recurrent op supportResizeBilinear/ResizeNearestNeighbor— no upsamplingExp,Log,Pow— no transcendental math (this is whydbscaling is problematic)
Channel alignment¶
The N6 NPU vectorizes computation in groups of 8 channels. The model builder
enforces this via _make_divisible(channels, 8) in
birdnet_stm32/models/blocks.py.
When alpha=0.25, stage 1 gets 64 × 0.25 = 16 channels (aligned).
When alpha=0.1, stage 1 would get 64 × 0.1 = 6.4 → rounded to 8.
Misaligned channels either waste compute (the NPU pads to the next multiple of 8) or fail compilation entirely.
QAT implementation¶
The QAT implementation uses a native Keras 3 training graph rather than TensorFlow Model Optimization Toolkit (tfmot), because:
- tfmot is incompatible with Keras 3 (as of 2026).
- tfmot injects FakeQuant ops that may not be supported by the N6 NPU.
- The deployment graph stays clean: the QAT graph applies differentiable per-channel kernel and per-tensor activation INT8 grids. Standard backbone variables are shared; separately cloned custom-frontend weights are synced into the clean model immediately before every checkpoint.
The simulation covers the quantized waveform input, fused outer-graph
activation boundaries, and the kernel and elementwise boundaries hidden inside
AudioFrontendLayer and MagnitudeScalingLayer. BatchNorm boundaries followed
by ReLU are not quantized twice because LiteRT folds those sequences into one
operator. Boundary discovery follows through Dropout and SpatialDropout layers,
which disappear during inference, so training-only regularization cannot create
a fake requantization boundary that is absent from the deployed graph.
The activation ranges come from the converter's exact deterministic, class-stratified calibration manifest and preprocessing path. The untouched float checkpoint is also loaded as a frozen teacher. Training adds per-output Bernoulli KL divergence and both mean and worst-sample per-sample cosine distance to the label loss. The default tail objective targets the worst 10% of each batch. Background and low-confidence probabilities remain calibrated while the lower parity tail is optimized directly; hard-label BCE alone barely penalizes those errors.
Which epoch is kept is a separate decision from how the loss is weighted, and the two can pull apart: the tail objective is a parity measure, so selecting on it once parity is comfortable keeps the epoch that drifted furthest from the float teacher. Selection therefore ignores the loss and scores the actual converted INT8 model each epoch on exact file cMAP. See Selection and artifacts.
The saved .keras model contains only standard float32 weights — no FakeQuant
nodes. Standard PTQ then calibrates and quantizes the hardened deployment
graph; held-out parity and task-level accuracy remain mandatory gates.
See birdnet_stm32/training/qat.py for the implementation, and
birdnet_stm32/training/distillation.py for the teacher-consistency losses it
optimizes alongside the label loss.