Exploración de los límites del muestreo selectivo para el reconocimiento de emociones del habla dependientes del hablante
Main Article Content
Resumen
El reconocimiento de emociones en el habla es crucial para diversas aplicaciones, desde el ámbito de la salud mental hasta la interacción con máquinas. Aunque los modelos basados en aprendizaje profundo han demostrado ser efectivos en este tipo de tareas, su rendimiento puede verse afectado cuando se enfrentan a nuevos individuos, incluso si se entrenaron con grandes y diversos conjuntos de datos. En este estudio, analizamos la relevancia de ajustar y personalizar estos modelos para mejorar la precisión en el reconocimiento emocional. A pesar de que los modelos generales ofrecen buenos resultados iniciales, notamos que su capacidad para predecir emociones disminuye cuando se aplican a nuevos sujetos. A través de este trabajo, mostramos cómo personalizar un modelo con una pequeña cantidad de muestras de un nuevo individuo puede aumentar considerablemente su desempeño, utilizando técnicas como el ajuste fino y la transferencia de aprendizaje. Finalmente, discutimos las implicaciones de estos resultados en aplicaciones reales y sugerimos posibles direcciones para futuras investigaciones en la detección de emociones mediante inteligencia artificial.
Article Details

Esta obra está bajo licencia internacional Creative Commons Reconocimiento-NoComercial-SinObrasDerivadas 4.0.
Citas
[2] N. Elsayed, Z. ElSayed, N. Asadizanjani, M. Ozer, A. Abdelgawad, & M. Bayoumi. (2022). “Speech emotion recognition using supervised deep recurrent system for mental health monitoring”. IEEE 8th World Forum on Internet of Things (WF-IoT), pp. 1–6. DOI:10.48550/arXiv.2208.12812
[3] A. Radford, J.W. Kim, T. Xu, G. Brockman, C. Mcleavey, & I. Sutskever. (2023). “Robust speech recognition via large-scale weak supervision”. Proceedings of the 40th International Conference on Machine Learning, vol. 202, pp. 28 492–28 518. DOI:10.48550/arXiv.2212.04356
[4] M. AbdelWahab & C. Busso. (2018). “Domain Adversarial for Acoustic Emotion Recognition”. IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 26(12), 2423–2435. DOI:10.1109/TASLP.2018.2867099
[5] J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, & B. W. Schuller. (2023). “Dawn of the transformer era in speech emotion recognition: Closing the valence gap”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, pp. 10 745–10 759. DOI:10.48550/arXiv.2203.07378
[6] M. Abdelwahab & C. Busso. (2017). “Incremental adaptation using active learning for acoustic emotion recognition”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5160–5164. DOI:10.1109/ICASSP.2017.7953140
[7] S. Madanian, T. Chen, O. Adeleye, J. M. Templeton, C. Poellabauer, D. Parry, & S. L. Schneider. (2023). Speech emotion recognition using machine learning — A systematic review. In Intelligent Systems with Applications (Vol. 20). DOI:10.1016/j.iswa.2023.200266
[8] A. Hashem, M. Arif, & M. Alghamdi. (2023). Speech emotion recognition approaches: A systematic review. In Speech Communication (Vol. 154). DOI:10.1016/j.specom.2023.102974
[9] C. Lu, Y. Zong, W. Zheng, Y. Li , C. Tang, & B. W. Schuller. (2022). “Domain Invariant Feature Learning for Speaker-Independent Speech Emotion Recognition”. IEEE/ACM Transactions on Audio Speech and Language Processing, 30, 2217–2230. DOI:10.1109/TASLP.2022.3178232
[10] M. Abdelwahab & C. Busso. (2015). “Supervised domain adaptation for emotion recognition from speech”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5058–5062. DOI:10.1109/ICASSP.2015.7178934
[11] C. Castorena, J. Lopez-Ballester, M. Cobos & F.J. Ferri. (2024). “Explorando la personalización de modelos en el reconocimiento de emociones en el habla”. Congreso Ibérico de Acústica – 55º Congreso Español de Acústica – Tecniacústica.
[12] C. Busso, M. Bulut, C. C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, & S. S. Narayanan. (2008). “Iemocap: Interactive emotional dyadic motion capture database”. Language Resources and Evaluation, vol. 42, pp. 335–359. DOI:10.1007/s10579-008-9076-6
[13] R. Lotfian & C. Busso. (2019). “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings”. IEEE Transactions on Affective Computing, vol. 10, pp. 471–483. DOI:10.1109/TAFFC.2017.2736999
[14] F. Burkhardt, W. F. Sendlmeier, F. Burkhardt, A. Paeschke, M. Rolfes, W. Sendlmeier, & B. Weiss. (2005). “A database of german emotional speech see profile a database of german emotional speech”. Proc. Interspeech 2005, pp. 1517–1520. DOI:10.21437/Interspeech.2005-446
[15] M. M. Duville, L. M. Alonso-Valerdi, & D. I. Ibarra-Zarate. (2021). “Mexican emotional speech database based on semantic, frequency, familiarity, concreteness, and cultural shaping of affective prosody”. Data 2021, Vol. 6, vol. 6, p. 130. DOI:10.3390/data6120130
[16] S. R. Livingstone & F. A. Russo. (2018) “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english”. PLOS ONE, vol. 13, p. e0196391. DOI:10.1371/journal.pone.0196391
[17] K. Zhou, B. Sisman, R. Liu, & H. Li. (2021). “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset”. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, vol. 2021-June, pp. 920–924. DOI:10.48550/arXiv.2010.14794
[18] Z. T. Liu, A. Rehman, M. Wu, W. H. Cao & M. Hao (2021). “Speech emotion recognition based on formant characteristics feature extraction and phoneme type convergence”. Information Sciences, p. 563. DOI:10.1016/j.ins.2021.02.016
[19] S. Vaijayanthi, & J. Arunnehru (2021). “Synthesis Approach for Emotion Recognition from Cepstral and Pitch Coefficients Using Machine Learning”. Lecture Notes in Electrical Engineering, p. 733. DOI:10.1007/978-981-33-4909-4_39
[20] R. Thirumuru, K. Gurugubelli & A. K. Vuppala (2022). “Novel feature representation using single frequency filtering and nonlinear energy operator for speech emotion recognition”. Digital Signal Processing: A Review Journal, p. 120. DOI:10.1016/j.dsp.2021.103293
[21] J. Ancilin & A. Milton (2021). “Improved speech emotion recognition with Mel frequency magnitude coefficient”. Applied Acoustics, p. 179. DOI:10.1016/j.apacoust.2021.108046
http://orcid.org/0000-0002-4202-1008