Intelligent Fault Detection for Belt Conveyor Idlers Using Machine Learning

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

Saudi Digital Library

Abstract

Conveyor belt systems are essential components of modern mining, logistics, and manufacturing operations. However, faults in idlers, which are the rollers that support and guide the belt, can reduce system reliability, cause unplanned downtime, increase maintenance costs, and create safety risks. Conventional inspection methods, such as visual assessment and manual listening, are widely used to identify abnormal idler behaviour, but they are labour intensive, subjective, and difficult to apply across large conveyor systems. Acoustic monitoring offers a promising contactless alternative because developing faults often produce changes in sound before complete failure occurs. However, these fault related acoustic signatures can be subtle and may be masked by environmental and operational noise, making them difficult to interpret using traditional signal processing methods alone. Machine learning (ML) has therefore been increasingly applied to acoustic signals to automate feature extraction, fault detection, and classification. Nevertheless, existing ML approaches still face several important challenges, including the limited availability of labelled fault data, inadequate generalisation across recording conditions and sensing platforms, and insufficient modelling of the spatial and temporal characteristics of acoustic signals. This thesis addresses these challenges by developing intelligent fault detection (IFD) models that automatically analyse acoustic signals from belt conveyor idlers. The research follows a progressive methodology comprising supervised transfer learning, semi-supervised anomaly detection, task-specific spatial–temporal deep learning, and cross-domain Convolutional Neural Network (CNN)–Transformer adaptation. To support these investigations, acoustic datasets were collected using handheld microphones and drone-mounted recorders under multiple idler operating and fault conditions. The first study investigated supervised fault classification using the original handheld acoustic dataset, which contained 255 four-second samples representing Normal, Stage 1, Stage 2, and Stage 3 operating conditions. Embeddings were extracted using YAMNet, a pre-trained audio neural network, and processed using Bidirectional Long Short-Term Memory (BiLSTM) and Bidirectional Gated Recurrent Unit (BiGRU) networks to model contextual relationships in both forward and backward temporal directions. Attention mechanisms and an Extreme Gradient Boosting (XGBoost) classifier were also evaluated. The best-performing YAMNet–BiLSTM configuration achieved an accuracy of 90.59% and an F1-score of 90.57% for four-class fault-stage classification. Although this study established a strong supervised baseline, it required labelled examples from every operating condition and fault stage. This dependence limits practical application because faulty-idler recordings are relatively rare, costly to collect, and difficult to label in industrial environments. To address this limitation, the second study developed CASSAD (Chroma-Augmented Semi-Supervised Anomaly Detection), which was trained exclusively on normal operating samples during training. Following the first study, additional handheld recordings were collected, increasing the dataset from 255 to 468 acoustic samples. This larger dataset enabled CASSAD to be evaluated using a broader set of normal and abnormal operating recordings. CASSAD combines Chroma-STFT, Chroma-CQT, and Chroma-CENS representations with filtering, statistical aggregation, and a one-class support vector machine (OC-SVM). On the expanded 468-sample handheld dataset, the best-performing CASSAD configuration achieved an accuracy of 90.59%, a positive-class F1-score of 92.42%, and an Area Under the Receiver Operating Characteristic Curve of 96.29%. To enable a direct comparison with the YAMNet-based models, CASSAD was also evaluated on the original 255-sample dataset after the three fault stages had been combined into a single Abnormal class. In this binary evaluation, CASSAD achieved an accuracy of 93.00% and an F1-score of 93.25%, compared with an accuracy of 92.18% and an F1-score of 93.00% for the strongest YAMNet-based configuration. These results demonstrate that competitive anomaly-detection performance can be achieved without labelled abnormal samples during training. However, CASSAD provides only binary Normal–Abnormal decisions and uses temporally aggregated features, limiting its ability to distinguish fault severity and model changes in acoustic behaviour over time. To overcome these limitations, the third study developed TD-CLNet, a Time-Distributed CNN–Long Short-Term Memory (LSTM) architecture designed to perform multi-stage fault classification while learning spatial and temporal representations from the acoustic data. The expanded 468-sample handheld dataset was re-segmented into one-second samples and converted into log-Mel-spectrogram frames. TD-CLNet applies a shared CNN feature extractor to each frame and then uses an LSTM to model the resulting temporal feature sequence. This design enables the model to distinguish among Normal, Stage 1, Stage 2, and Stage 3 conditions while preserving temporal information that was reduced through the statistical aggregation used in CASSAD. Under four-fold cross-validation, TD-CLNet achieved a mean accuracy, precision, recall, and weighted F1-score of 92.1%. It outperformed the evaluated conventional CNN–LSTM configurations and provided a small improvement over the strongest YAMNet model re-evaluated on the same expanded dataset. Nevertheless, the sequential processing used by LSTM networks limits parallel computation and provides less direct access to broader global relationships within the acoustic feature sequence. To address these limitations and investigate cross-domain generalisation, the fourth study developed hybrid CNN–Transformer models using four pre-trained CNN backbones: ResNet-18, DenseNet-121, EfficientNet-B0, and ShuffleNet-V2. Two acoustic feature representations were investigated: Mel-spectrograms, which represent the distribution of signal energy across perceptually scaled frequency bands over time, and Mel-Frequency Cepstral Coefficients (MFCCs), which provide a compact representation of the short-term spectral envelope. The CNN backbones extracted local spectral representations, while Transformer encoders modelled broader contextual relationships within the feature sequences. In the first phase, the models were trained and evaluated using 0.5-second segments derived from the handheld source-domain recordings. The ResNet-18–Transformer models achieved a cross-fold mean accuracy of 96.9%, with a 95% confidence interval of 96.4%–97.6%, while a ResNet-18–Transformer ensemble increased the handheld-domain accuracy to 98.0%. In the second phase, the models trained on the handheld source-domain dataset were adapted to the drone-acquired target-domain dataset. The drone recordings represented a more challenging sensing environment because they were affected by rotor noise, changing recording distances, varying microphone positions, and environmental interference. Multiple fine-tuning strategies were evaluated, included full fine-tuning (Full-FT), freezing the CNN backbone (CNN-Frozen), freezing the Transformer encoder (TR-Frozen), and 𝐿2-SP regularisation. Mean–Covariance Alignment (MCA) was compared with a cross-entropy (CE) baseline and several established domain-adaptation methods. MFCCs produced the strongest CE baseline in several drone-domain experiments, achieving an accuracy of 90.5% under 𝐿2-SP. However, MFCC performance generally decreased when MCA was applied. In contrast, MCA improved the Mel-spectrogram Full-FT configuration, increasing accuracy from 84.7% to 86.0% and the F1-score from 83.4% to 85.6%. These findings demonstrate that domain-adaptation performance depends on the interaction among the acoustic representation, fine-tuning strategy, model architecture, and alignment objective; no single adaptation method was uniformly optimal across all evaluated configurations. In summary, this thesis establishes a connected research pathway for acoustic fault detection in belt conveyor idlers. It provides handheld and drone-acquired acoustic datasets, establishes supervised transfer-learning baselines, introduces CASSAD for anomaly detection without labelled abnormal training samples, develops TD-CLNet for task-specific spatial–temporal fault-stage classification, and proposes a two-phase CNN–Transformer framework for handheld-to-drone domain adaptation. Collectively, these contributions advance acoustic idler monitoring towards more accurate, data-efficient, and adaptable fault detection while identifying the need for further validation across additional industrial sites, conveyor configurations, sensing platforms, and operating conditions before large-scale real-world deployment.

Description

Keywords

Anomaly detection, Acoustic monitoring, Acoustic fault diagnosis, Belt conveyor idlers, Industrial condition monitoring, Predictive maintenance, Machine learning, Deep learning, Fault detection, Transfer learning, Semi-supervised anomaly detection, Spatial-temporal modelling, CNN-LSTM, CNN-Transformer, Domain adaptation, Drone-based inspection, Mel-spectrograms, Chroma features.

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By

Copyright owned by the Saudi Digital Library (SDL) © 2026