Quantized vision models lose apparent map collapse when the right layer is chosen
Sunzil Khandaker’s vision-model comparison found that a layer-selection shortcut showed apparent explanation collapse. Selecting layers with usable spatial information removed that failure across 3,240 comparisons, while differences between full-precision and compressed models remained. The study also found that calibration and training choices affected architectures differently, with different effects during conversion to a lower-precision representation.
Artificial Intelligence··Night
Layer selection removes apparent collapse
Sunzil Khandaker’s comparison of compressed vision models identified a failure in the way explanation tools selected their input layer. A shortcut picked a representation with a spatial extent of 1×1 in three architectures. Apparent explanation collapse reached 61.7 per cent in some settings. After architecture-appropriate spatial-layer selection, the experiment showed none of that collapse across 3,240 paired comparisons.[1]
Compression still changes highlighted regions
The experiment compared EfficientNetV2-S, ResNet50 and MobileNetV3-Large on 320 class-balanced images from Paddy Disease, a plant-disease image dataset. It evaluated four explanation methods, which identify image regions associated with a classification. After correcting layer selection, important-region overlap ranged from 0.366 at the lower end and 0.790 at the upper end; the random baseline was 0.081. Integrated Gradients preserved its maps least well, and MobileNetV3-Large was the most affected architecture.[1]
Calibration changes prediction agreement
Calibration, the choice of numerical ranges during conversion, also altered the outcomes measured with the ONNX Runtime model-execution system. MinMax calibration gave EfficientNet agreement of 0.644 and MobileNet agreement of 0.641. With percentile calibration, those values reached 0.927 and 0.915 respectively. With training that aligned explanation maps between teacher and student models, MobileNet overlap rose while EfficientNet overlap fell. The peer-reviewed version expanded an earlier 20-image evaluation to 320 images.[1]
Related columns
For more information on this topic, you can read the related columns.