Preprint / Version 1

Dissolved Gas Analysis-Based Power Transformer Fault Diagnosis Using Random Forest

A Comparative Evaluation Against K-Nearest Neighbors, Support Vector Machine, and Artificial Neural Network Classifiers

##article.authors##

DOI:

https://doi.org/10.31224/8150

Keywords:

Dissolved Gas Analysis, Power Transformers, Random Forest, Fault Diagnosis, Machine Learning, Domain Generalization, Predictive Maintenance

Abstract

Power transformers are among the most critical and capital-intensive assets in an electrical power system, and their unplanned failure carries significant economic and reliability consequences. Dissolved Gas Analysis (DGA) remains the most widely adopted diagnostic technique for detecting incipient faults in oil-immersed transformers, yet classical interpretation schemes such as the Rogers Ratio, IEC 60599, and Duval Triangle methods are rule-based and frequently produce inconclusive or borderline diagnoses. This study investigates the application of Random Forest, an ensemble machine-learning algorithm, to DGA-based fault classification, and rigorously benchmarks it against three widely used alternatives — K-Nearest Neighbors (KNN), Support Vector Machine (SVM), and Artificial Neural Network (ANN) — trained and evaluated on an identical, stratified data split to ensure a fair comparison. A curated dataset of 589 field-verified DGA records spanning six IEC 60599 fault classes was enriched with domain-informed features, including classical diagnostic gas ratios and Duval-triangle percentages, and class imbalance was addressed using a SMOTE-style synthetic oversampling procedure applied strictly within training folds. Following exhaustive hyperparameter tuning via grid search and repeated stratified cross-validation, Random Forest achieved the highest held-out test accuracy of 83.76%, outperforming KNN (76.07%), SVM (76.07%), and ANN (70.94%). Beyond same-dataset evaluation, this study addresses a gap consistently left open in prior DGA literature: generalization across independent data sources. When the Random Forest model trained on the curated dataset was tested against two independent multi-source DGA datasets, accuracy fell to 66.48% and 67.12% respectively, revealing a substantial domain-shift problem that single-dataset accuracy claims in the literature do not capture. These findings suggest that while Random Forest is a strong candidate for DGA-based fault diagnosis, cross-source generalization — not same-dataset accuracy — should be the benchmark by which practical deployability is judged. The study concludes with recommendations for domain adaptation and standardized multi-utility datasets as directions for future work.

Downloads

Download data is not yet available.

Downloads

Posted

2026-09-06