Interpretability of Vision Transformer in bacterial classification on SEM images


Interpretability of Vision Transformer in bacterial classification on SEM images

Gridin V.N. (DITC RAS, Odintsovo, Russia)
Salem B.R. (DITC RAS, Odintsovo, Russia)
Solodovnikov V.I. (DITC RAS, Odintsovo, Russia)
Novikov I.A. (DITC RAS, Odintsovo, Russia; NIIGB, Moscow, Russia)
Rodina E.S. (NIIGB, Moscow, Russia)

Abstract

The paper addresses multiclass single-label classification of bacterial microorganisms on single-channel images obtained by scanning electron microscopy with lanthanide contrast enhancement. The study compares three approaches to interpreting Vision Transformer decisions: the baseline attention map of the last block (AttentionCAM, CLS-to-patches), the class-specific CDAM map, and the R-Cut method, which yields a connected localization region. The analysis was carried out for the current version of the ViT-Base/16 model operating on 384×384 images. An independent test subsample of 371 images was used to evaluate classification performance. The following integrated performance indicators were obtained: a correct classification rate of 0.784, a macro-averaged F1 score of 0.790, and a weighted F1 score of 0.785. For quantitative comparison of explanations without manual mask annotation, a stratified subset of 160 correctly classified images was used. On this subset, AttentionCAM showed the best mean values in terms of the area under the deletion and insertion curves, the average decrease in the target response, and the proportion of map intensity inside the automatically extracted object region. At the same time, CDAM and R-Cut provide a different type of interpretation: the former highlights local discriminative regions, whereas the latter tends to form a more connected localization area, although it is sensitive to the choice of output normalization and the parameter φ.

Keywords

Vision Transformer; interpretability; attention maps; class-specific attribution; connected localization; bacterial classification; scanning electron microscopy.

Edition

Proceedings of the Institute for System Programming, vol. 38, issue 6, part 1, 2026, pp. 311-322

ISSN 2220-6426 (Online), ISSN 2079-8156 (Print).

DOI: 10.15514/ISPRAS-2026-38(6)-20

For citation

Gridin V.N., Salem B.R., Solodovnikov V.I., Novikov I.A., Rodina E.S. Interpretability of Vision Transformer in bacterial classification on SEM images. Proceedings of the Institute for System Programming, vol. 38, issue 6, part 1, 2026, pp. 311-322 DOI: 10.15514/ISPRAS-2026-38(6)-20.

Full text of the paper in pdf (in Russian) Back to the contents of the volume