|
|
 |
| |
|
|
|
Multimodal Deep Neural Network for Suicide Risk Classification Using Audio and Facial Features |
|
|
|
PP: 1091-1101 |
|
|
doi:10.18576/amis/200418
|
|
|
|
Author(s) |
|
|
|
A. B. Mukhametzhanov,
A. Shoiynbek,
A. Nurzhas,
B. Meraliyev,
D. Kuanyshbay,
S. Sklyar,
P. Menezes,
|
|
|
|
Abstract |
|
|
| Suicide-risk assessment remains a major public-health priority, yet current clinical evaluations rely heavily on subjective judgment and self-disclosure. This study investigates the effectiveness of combining low-level audio features and deep visual embeddings for suicide classification using interview-style recordings. The dataset consists of 103 clinical interviews stratified into control and at-risk groups, from which Mel-Frequency Cepstral Coefficients (MFCCs) and VGG16-derived facial embeddings were extracted through a structured preprocessing and alignment pipeline. Unimodal and multimodal models were evaluated using classical machine learning algorithms and a shallow neural network. The best-performing multimodal model, a two-layer neural network, achieved 0.82 accuracy and 0.82 macro F1-score, outperforming logistic regression, gradient boosting, XGBoost, and naive Bayes baselines. The results demonstrate the predictive potential of facial embeddings for suicide-risk detection and provide a reproducible feature-alignment pipeline for future multimodal mental-health classification studies. |
|
|
|
|
 |
|
|