Multilevel saliency-guided self-supervised learning for image anomaly detection

Jianjian Qin; Chunzhi Gu; Jun Yu; Chao Zhang

doi:10.1007/s11760-024-03320-z

Multilevel saliency-guided self-supervised learning for image anomaly detection

Jianjian Qin, Chunzhi Gu, Jun Yu, Chao Zhang^*

^*Corresponding author for this work

Intellectual Information Engineering

Research output: Contribution to journal › Article › peer-review

Abstract

Anomaly detection (AD) is a fundamental task in computer vision. It aims to identify incorrect image data patterns which deviate from the normal ones. Conventional methods generally address AD by preparing augmented negative samples to enforce self-supervised learning. However, these techniques typically do not consider semantics during augmentation, leading to the generation of unrealistic or invalid negative samples. Consequently, the feature extraction network can be hindered from embedding critical features. In this study, inspired by visual attention learning approaches, we propose CutSwap, which leverages saliency guidance to incorporate semantic cues for augmentation. Specifically, we first employ LayerCAM to extract multilevel image features as saliency maps and then perform clustering to obtain multiple centroids. To fully exploit saliency guidance, on each map, we select a pixel pair from the cluster with the highest centroid saliency to form a patch pair. Such a patch pair includes highly similar context information with dense semantic correlations. The resulting negative sample is created by swapping the locations of the patch pair. Compared to prior augmentation methods, CutSwap generates more subtle yet realistic negative samples to facilitate quality feature learning. Extensive experimental and ablative evaluations demonstrate that our method achieves state-of-the-art AD performance on two mainstream AD benchmark datasets.

Original language	English
Pages (from-to)	6339-6351
Number of pages	13
Journal	Signal, Image and Video Processing
Volume	18
Issue number	8-9
DOIs	https://doi.org/10.1007/s11760-024-03320-z
State	Published - 2024/09

Keywords

Anomaly detection
Data augmentation
Representation learning
Self-supervised learning

ASJC Scopus subject areas

Signal Processing
Electrical and Electronic Engineering

Access to Document

10.1007/s11760-024-03320-z

Cite this

@article{c93d9fdc8b8e4ebb8bcd9f3f33970587,

title = "Multilevel saliency-guided self-supervised learning for image anomaly detection",

abstract = "Anomaly detection (AD) is a fundamental task in computer vision. It aims to identify incorrect image data patterns which deviate from the normal ones. Conventional methods generally address AD by preparing augmented negative samples to enforce self-supervised learning. However, these techniques typically do not consider semantics during augmentation, leading to the generation of unrealistic or invalid negative samples. Consequently, the feature extraction network can be hindered from embedding critical features. In this study, inspired by visual attention learning approaches, we propose CutSwap, which leverages saliency guidance to incorporate semantic cues for augmentation. Specifically, we first employ LayerCAM to extract multilevel image features as saliency maps and then perform clustering to obtain multiple centroids. To fully exploit saliency guidance, on each map, we select a pixel pair from the cluster with the highest centroid saliency to form a patch pair. Such a patch pair includes highly similar context information with dense semantic correlations. The resulting negative sample is created by swapping the locations of the patch pair. Compared to prior augmentation methods, CutSwap generates more subtle yet realistic negative samples to facilitate quality feature learning. Extensive experimental and ablative evaluations demonstrate that our method achieves state-of-the-art AD performance on two mainstream AD benchmark datasets.",

keywords = "Anomaly detection, Data augmentation, Representation learning, Self-supervised learning",

author = "Jianjian Qin and Chunzhi Gu and Jun Yu and Chao Zhang",

note = "Publisher Copyright: {\textcopyright} The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2024.",

year = "2024",

month = sep,

doi = "10.1007/s11760-024-03320-z",

language = "英語",

volume = "18",

pages = "6339--6351",

journal = "Signal, Image and Video Processing",

issn = "1863-1703",

publisher = "Springer London",

number = "8-9",

}

TY - JOUR

T1 - Multilevel saliency-guided self-supervised learning for image anomaly detection

AU - Qin, Jianjian

AU - Gu, Chunzhi

AU - Yu, Jun

AU - Zhang, Chao

N1 - Publisher Copyright: © The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2024.

PY - 2024/9

Y1 - 2024/9

N2 - Anomaly detection (AD) is a fundamental task in computer vision. It aims to identify incorrect image data patterns which deviate from the normal ones. Conventional methods generally address AD by preparing augmented negative samples to enforce self-supervised learning. However, these techniques typically do not consider semantics during augmentation, leading to the generation of unrealistic or invalid negative samples. Consequently, the feature extraction network can be hindered from embedding critical features. In this study, inspired by visual attention learning approaches, we propose CutSwap, which leverages saliency guidance to incorporate semantic cues for augmentation. Specifically, we first employ LayerCAM to extract multilevel image features as saliency maps and then perform clustering to obtain multiple centroids. To fully exploit saliency guidance, on each map, we select a pixel pair from the cluster with the highest centroid saliency to form a patch pair. Such a patch pair includes highly similar context information with dense semantic correlations. The resulting negative sample is created by swapping the locations of the patch pair. Compared to prior augmentation methods, CutSwap generates more subtle yet realistic negative samples to facilitate quality feature learning. Extensive experimental and ablative evaluations demonstrate that our method achieves state-of-the-art AD performance on two mainstream AD benchmark datasets.

AB - Anomaly detection (AD) is a fundamental task in computer vision. It aims to identify incorrect image data patterns which deviate from the normal ones. Conventional methods generally address AD by preparing augmented negative samples to enforce self-supervised learning. However, these techniques typically do not consider semantics during augmentation, leading to the generation of unrealistic or invalid negative samples. Consequently, the feature extraction network can be hindered from embedding critical features. In this study, inspired by visual attention learning approaches, we propose CutSwap, which leverages saliency guidance to incorporate semantic cues for augmentation. Specifically, we first employ LayerCAM to extract multilevel image features as saliency maps and then perform clustering to obtain multiple centroids. To fully exploit saliency guidance, on each map, we select a pixel pair from the cluster with the highest centroid saliency to form a patch pair. Such a patch pair includes highly similar context information with dense semantic correlations. The resulting negative sample is created by swapping the locations of the patch pair. Compared to prior augmentation methods, CutSwap generates more subtle yet realistic negative samples to facilitate quality feature learning. Extensive experimental and ablative evaluations demonstrate that our method achieves state-of-the-art AD performance on two mainstream AD benchmark datasets.

KW - Anomaly detection

KW - Data augmentation

KW - Representation learning

KW - Self-supervised learning

UR - http://www.scopus.com/inward/record.url?scp=85196271394&partnerID=8YFLogxK

U2 - 10.1007/s11760-024-03320-z

DO - 10.1007/s11760-024-03320-z

M3 - 学術論文

AN - SCOPUS:85196271394

SN - 1863-1703

VL - 18

SP - 6339

EP - 6351

JO - Signal, Image and Video Processing

JF - Signal, Image and Video Processing

IS - 8-9

ER -

Multilevel saliency-guided self-supervised learning for image anomaly detection

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Fingerprint

Cite this