Hybrid Vision Transformer Model for Intelligent Surface Defect Detection in Manufacturing
Main Article Content
Abstract
The rapid advancement of Industry 4.0 and intelligent manufacturing has significantly increased the demand for automated surface defect inspection systems capable of ensuring high product quality and production efficiency. Surface defects such as scratches, cracks, dents, corrosion, inclusions, and texture irregularities directly affect the structural integrity, functionality, and commercial value of manufactured products. Conventional manual inspection methods are often time-consuming, subjective, and inconsistent, while traditional computer vision techniques struggle to detect complex defects under varying illumination, background noise, and texture conditions. Recent advances in deep learning, particularly Vision Transformers (ViTs), have demonstrated superior capability in capturing long-range dependencies and global contextual information. However, standalone Vision Transformer models often exhibit limitations in learning fine local features that are essential for accurate defect localization. Consequently, hybrid architectures integrating Convolutional Neural Networks (CNNs) and Vision Transformers have emerged as promising solutions for intelligent surface defect detection. Recent hybrid transformer models have demonstrated improved performance by effectively combining local spatial feature extraction with global contextual representation for industrial inspection tasks.