Eigen RadarAI
Analysis

Quantized vision models lose apparent map collapse when the right layer is chosen

Sunzil Khandaker’s vision-model comparison found that a layer-selection shortcut showed apparent explanation collapse. Selecting layers with usable spatial information removed that failure across 3,240 comparisons, while differences between full-precision and compressed models remained. The study also found that calibration and training choices affected architectures differently, with different effects during conversion to a lower-precision representation.

Artificial Intelligence··Night
A wide metal aperture above rice leaves with brown lesions.

Layer selection removes apparent collapse

Sunzil Khandaker’s comparison of compressed vision models identified a failure in the way explanation tools selected their input layer. A shortcut picked a representation with a spatial extent of 1×1 in three architectures. Apparent explanation collapse reached 61.7 per cent in some settings. After architecture-appropriate spatial-layer selection, the experiment showed none of that collapse across 3,240 paired comparisons.[1]

Compression still changes highlighted regions

The experiment compared EfficientNetV2-S, ResNet50 and MobileNetV3-Large on 320 class-balanced images from Paddy Disease, a plant-disease image dataset. It evaluated four explanation methods, which identify image regions associated with a classification. After correcting layer selection, important-region overlap ranged from 0.366 at the lower end and 0.790 at the upper end; the random baseline was 0.081. Integrated Gradients preserved its maps least well, and MobileNetV3-Large was the most affected architecture.[1]

Calibration changes prediction agreement

Calibration, the choice of numerical ranges during conversion, also altered the outcomes measured with the ONNX Runtime model-execution system. MinMax calibration gave EfficientNet agreement of 0.644 and MobileNet agreement of 0.641. With percentile calibration, those values reached 0.927 and 0.915 respectively. With training that aligned explanation maps between teacher and student models, MobileNet overlap rose while EfficientNet overlap fell. The peer-reviewed version expanded an earlier 20-image evaluation to 320 images.[1]

References

  1. News sourceScientific ReportsA comparison framework measures explanation drift after quantization↩1↩2↩3