Comparación de modelos transformer para clasificar poesía en español
Main Article Content
Resumen
La poesía en español, por su riqueza estilística y el uso de recursos retóricos, plantea retos particulares para la clasificación automática. Cualquier sesgo en etapas clave, incluyendo la representación del texto, la arquitectura del modelo o el protocolo de evaluación, puede impactar el desempeño; por ello, resulta pertinente comparar familias de modelos bajo condiciones homogéneas. En este artículo se estudia la efectividad relativa de cuatro arquitecturas basadas en Transformers ampliamente utilizadas: BERT-base uncased (DCC), BERT-base cased (DCC), RoBERTa-base (BERTIN) y DistilBERT uncased (DCC). La metodología emplea un protocolo único de ajuste fino para cada modelo (mismos hiperparámetros iniciales, partición proporcional por clase y control de semilla), con evaluación mediante métricas estándar (accuracy, precision, recall, F1) y análisis cualitativo de embeddings antes y después del entrenamiento. El objetivo es identificar, en un escenario de datos limitados, qué arquitectura ofrece el mejor equilibrio entre rendimiento y costo computacional, así como caracterizar errores típicos asociados al estilo poético. Los hallazgos buscan ofrecer lineamientos reproducibles para la selección de modelos en corpus pequeños de poesía en lengua española.
Article Details

Esta obra está bajo licencia internacional Creative Commons Reconocimiento-NoComercial-SinObrasDerivadas 4.0.
Citas
[2] A. Perez Pozo, J. De La Rosa, S. Ros, E. Gonzalez-Blanco, L. Hernandez, and M. De Sisto, “A bridge too far for artificial intelligence?: Automatic classification of stanzas in Spanish poetry,” Journal of the Association for Information Science and Technology, vol. 73, no. 2, pp. 258–267, Feb. 2022. DOI: 10.1002/asi.24532
[3] J. M. Johnson and T. M. Khoshgoftaar, “Survey on deep learning with class imbalance,” Journal of Big Data, vol. 6, no. 1, p. 27, Dec. 2019. DOI: 10.1186/s40537-019-0192-5
[4] J. Cañete, S. Donoso, F. Bravo-Marquez, A. Carvallo, and V. Araujo, “Spanish pre-trained bert model and evaluation data,” in Proceedings of PML4DC at ICLR, 2020. [Online]. Available: http://ceur-ws.org/Vol-2421/
[5] A. Gutiérrez-Fandiño, J. Armengol-Estapé, M. Pàmies, J. Llop-Palao, J. Silveira-Ocampo, C. P. Carrino, A. Gonzalez-Agirre, C. Armentano-Oller, C. Rodriguez-Penagos, and M. Villegas, “Maria: Spanish language models,” arXiv preprint arXiv:2107.07253, 2021. DOI: 10.48550/arXiv.2107.07253. [Online]. Available: https://arxiv.org/abs/2107.07253
[6] PlanTL-GOB-ES, “PlanTL-GOB-ES/roberta-base-bne – model card,” Hugging Face, 2021, roBERTa entrenado con el corpus de la Biblioteca Nacional de España (BNE). [Online]. Available: https://huggingface.co/PlanTL-GOB-ES/roberta-base-bne
[7] J. Cañete, S. Donoso, F. Bravo-Marquez, A. Carvallo, and V. Araujo, “Albeto and distilbeto: Lightweight spanish language models,” in Proceedings of the 13th Language Resources and Evaluation Conference (LREC 2022). Marseille, France: European Language Resources Association (ELRA), 2022. [Online]. Available: https://aclanthology.org/2022.lrec-1.457/
[8] T. Ceccon Silveira, “Applying BERT language model to poem classification: A study on data imbalance issues,” Universidade Federal do Rio Grande do Sul (UFRGS), Tech. Rep., 2023. [Online]. Available: https://lume.ufrgs.br/bitstream/handle/10183/259959/001172384.pdf
[9] W. Antoun, F. Baly, and H. Hajj, “Arabert: Transformer-based model for arabic language understanding,” in Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT4). Marseille, France: European Language Resources Association (ELRA), may 2020. [Online]. Available: https://aclanthology.org/2020.osact-1.2/
[10] L. Martin, B. Muller, P. J. Ortiz Suarez, Y. Dupont, L. Romary, E. De La Clergerie, D. Seddah, and B. Sagot, “CamemBERT: a Tasty French Language Model,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020, pp. 7203–7219. DOI: 10.18653/v1/2020.acl-main.645
[11] R. Delmonte and N. Busetto, “Stress Test for Bert and Deep Models: Predicting Words from Italian Poetry,” International Journal on Natural Language Computing, vol. 11, no. 6, pp. 15–37, Dec. 2022. DOI: 10.5121/ijnlc.2022.11602. [Online]. Available: https://aircconline.com/ijnlc/V11N6/11622ijnlc02.pdf
[12] I. N. Santana, R. S. Oliveira, and E. G. S. Nascimento, “Text Classification of News Using Transformer-based Models for Portuguese,” Journal of Systemics, Cybernetics and Informatics, vol. 20, no. 5, pp. 33–59, Oct. 2022. DOI: 10.54808/JSCI.20.05.33. [Online]. Available: https://www.iiisci.org/DOIJSCI/SA702PC22
[13] M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Information Processing & Management, vol. 45, no. 4, pp. 427–437, 2009. DOI: 10.1016/j.ipm.2009.03.002
[14] Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, pp. 1798–1828, 2013. DOI: 10.1109/TPAMI.2013.50
[15] T. Mikolov, K. Chen, G. S. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in International Conference on Learning Representations, 2013. [Online]. Available: https://api.semanticscholar.org/CorpusID:5959482
[16] J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of EMNLP, 2014, pp. 1532–1543. [Online]. Available: https://aclanthology.org/D14-1162/
[17] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching word vectors with subword information,” Transactions of the Association for Computational Linguistics, vol. 5, pp. 135–146, 2017. [Online]. Available: https://aclanthology.org/Q17-1010/
[18] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” CoRR, vol. abs/1706.03762, 2017. [Online]. Available: http://arxiv.org/abs/1706.03762
[19] B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth, “Recent advances in natural language processing via large pre-trained language models: A survey,” ACM Comput. Surv., vol. 56, no. 2, sep 2023. DOI: 10.1145/3605943
[20] J. Devlin, M. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT, 2019, pp. 4171–4186. [Online]. Available: https://aclanthology.org/N19-1423/
[21] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter,” CoRR, vol. abs/1910.01108, 2019. [Online]. Available: http://arxiv.org/abs/1910.01108
[22] J. de la Rosa, E. G. Ponferrada, P. Villegas, P. González de Prado Salas, M. Romero, and M. Grandury, “Bertin: Efficient pre-training of a spanish language model using perplexity sampling,” arXiv preprint arXiv:2207.06814, 2022. [Online]. Available: https://arxiv.org/abs/2207.06814
[23] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in Proceedings of ACL, 2020, pp. 8440–8451. [Online]. Available: https://aclanthology.org/2020.acl-main.747/
[24] M. Schuster and K. Nakajima, “Japanese and korean voice search,” in Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), 2012, pp. 5149–5152. DOI: 10.1109/ICASSP.2012.6288945
[25] R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of ACL, 2016, pp. 1715–1725. [Online]. Available: https://aclanthology.org/P16-1162/
[26] T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of EMNLP: System Demonstrations, 2018, pp. 66–71. [Online]. Available: https://aclanthology.org/D18-2012/
[27] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Technical Report, 2019, gPT-2 tokenizer uses byte-level BPE. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
[28] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692, 2019. [Online]. Available: https://arxiv.org/abs/1907.11692
[29] J. Cañete, G. Chaperon, R. Fuentes, J.-H. Ho, H. Kang, and J. Pérez, “Spanish Pre-trained BERT Model and Evaluation Data,” 2023, version Number: 1. [Online]. Available: https://arxiv.org/abs/2308.02976
[30] A. Gutiérrez-Fandiño, J. Armengol-Estapé, M. Pàmies, J. Llop-Palao, J. Silveira-Ocampo, C. P. Carrino, C. Armentano-Oller, C. Rodriguez-Penagos, A. Gonzalez-Agirre, and M. Villegas, “MarIA: Spanish Language Models,” Procesamiento del Lenguaje Natural, pp. 39–60, 2022. DOI: 10.26342/2022-68-3
[31] J. Cañete, S. Donoso, F. Bravo-Marquez, A. Carvallo, and V. Araujo, “ALBETO and DistilBETO: Lightweight Spanish language models,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, J. Odijk, and S. Piperidis, Eds.Marseille, France: European Language Resources Association, Jun. 2022, pp. 4291–4298. [Online]. Available: https://aclanthology.org/2022.lrec-1.457/
http://orcid.org/0009-0002-6293-7870