Interpretable Acoustic Domains in Smoking-Status Prediction From Sustained Phonation: A Speaker-Independent Secondary Analysis


Aydoğan Y., CANTÜRK İ.

Journal of Voice, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1016/j.jvoice.2026.08.044
  • Dergi Adı: Journal of Voice
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Periodicals Index Online, CINAHL, Communication Abstracts, EMBASE, International Bibliography of Theatre & Dance (IBTD), MEDLINE, Music Index, Music Periodicals Database, EBSCO Communication Source, Communication Source (EBSCO), Health Research Premium Collection (ProQuest), Music & Performing Arts Collection (ProQuest)
  • Anahtar Kelimeler: Acoustic domains, Counterfactual explanation, Interpretable machine learning, Leave-one-speaker-out validation, Smoking status, Voice acoustics
  • Yıldız Teknik Üniversitesi Adresli: Evet

Özet

Objectives: To determine whether a previously developed 208-variable smoking-status voice model could be compressed into interpretable acoustic domains without materially reducing speaker-independent discrimination, and to assess which domains remained stable across model, perturbation, and counterfactual analyses. Methods: A secondary analysis was performed on the previously reported cohort of 64 speakers (30 smokers and 34 nonsmokers), using one primary smartphone-recorded sustained /a/ phonation per speaker. The exact 208-variable prosody–spectral feature cache was grouped into 16 prespecified acoustic domains. Domain scores were constructed independently within each leave-one-speaker-out fold using training-only imputation, robust scaling, direction alignment, and median aggregation. A balanced logistic model was evaluated from raw out-of-fold scores. Full-pipeline permutation tests, demographic residualization, domain-only models, demographic-conditional domain replacement, one-speaker jackknife analysis, a secondary Explainable Boosting Machine, and empirical counterfactual searches were performed. Results: The 16-domain model achieved an area under the receiver operating characteristic curve (AUC) of 0.753 (95% confidence interval [CI], 0.624–0.864), compared with 0.769 (95% CI, 0.645–0.878) for the recovered 208-variable model; the paired difference was −0.016 (95% CI, −0.129 to 0.095). Full-pipeline permutation tests were significant under both unrestricted permutation (P = 0.00599) and permutation restricted within age-by-gender strata (P = 0.00500). Spectral centroid, formant structure, and harmonicity/noise were significant as standalone domains after Holm correction. No single conditional-replacement test remained significant after correction. Spectral centroid ranked first in 49 of 64 jackknife analyses and appeared in 57.1% of minimal plausible counterfactuals. A plausible counterfactual was found for 98.4% of speakers, with one domain sufficient for 84.4%. Conclusions: Smoking-associated voice discrimination could be represented compactly through interpretable acoustic domains. The evidence favored a distributed, partially redundant spectral and resonance pattern rather than a single indispensable biomarker. The results describe model behavior in this cohort and require external validation before clinical interpretation.