Novel anomaly-aware machine learning framework for accurate higher heating value prediction of biomass and wastes


İNSEL M. A., Hallak M., SADIKOĞLU H., Yucel O.

Journal of Thermal Analysis and Calorimetry, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s10973-026-16051-9
  • Dergi Adı: Journal of Thermal Analysis and Calorimetry
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Chemical Abstracts Core, Chimica, Compendex, Index Islamicus, INSPEC, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
  • Anahtar Kelimeler: Anomaly detection, Higher heating value, Local outlier factor, Machine learning, Waste to energy
  • Yıldız Teknik Üniversitesi Adresli: Evet

Özet

With the mounting global demand for sustainable energy alternatives, biomass and waste-derived fuels have emerged as pivotal resources for achieving carbon neutrality. Accurate prediction of the higher heating value (HHV) of biomass and waste-derived fuels is essential for data-driven energy applications; however, heterogeneous datasets often contain anomalous observations that degrade machine learning (ML) model performance. This study proposes an anomaly-aware ML framework to enhance HHV prediction using a dataset of 1526 samples comprising biomass, sludge, coal, and waste-derived fuels. Four unsupervised anomaly detection algorithms (Isolation Forest, Local Outlier Factor, One-Class SVM, and Elliptic Envelope) were systematically evaluated under two integration strategies: anomaly flagging and data reduction via outlier removal. Subsequently, nine regression models were trained and assessed using 10-fold cross-validation. The results demonstrate that anomaly-aware preprocessing significantly improves predictive performance. Specifically, the integration of the Local Outlier Factor approach with Random Forest Regression achieved the optimal balance between training noise reduction and data representativeness, yielding the highest predictive capability with an R2 score of 0.8986 and a reduced RMSE of 1.3641 MJ kg−1, thereby successfully filtering out detrimental training noise while preserving core distribution features. Principal Component Analysis and Van Krevelen diagrams confirm that detected anomalies correspond to statistical outliers representing non-standard or peripheral fuel chemistries, supporting framework interpretability. The proposed transferable methodology optimizes ML performance in high-variability engineering datasets. Ultimately, these findings of this study help for better understanding of fuel-specific compositional anomalies, significantly enhancing the reliability of large-scale bioenergy yield assessments, techno-economic evaluations, and thermal conversion system optimization.