The Effect of Normalization for Bi-directional Amharic-English Neural Machine Translation

Tadesse Destaw Belay, Atnafu Lambebo Tonja, Olga Kolesnikova, Seid Muhie Yimam, Abinew Ali Ayele, Silesh Bogale Haile, Grigori Sidorov, Alexander Gelbukh

Producción científica: Capítulo del libro/informe/acta de congresoContribución a la conferenciarevisión exhaustiva

3 Citas (Scopus)

Resumen

Machine translation (MT) is one of the prominent tasks in natural language processing whose objective is to translate texts automatically from one natural language to another. Nowadays, using deep neural networks for MT task has received a great deal of attention. These networks require lots of data to learn abstract representations of the input and store it in continuous vectors. This paper presents the first relatively large-scale Amharic-English parallel sentence dataset. Using these compiled data, we build bi-directional Amharic-English translation models by fine-tuning the existing Facebook M2M100 pre-trained model achieving a BLEU score of 37.79 in Amharic-English translation and 32.74 in English-Amharic translation. Additionally, we explore the effects of Amharic homophone normalization on the machine translation task. The results show that normalization of Amharic homophone characters increases the performance of Amharic-English machine translation in both directions.

Idioma originalInglés
Título de la publicación alojada2022 International Conference on Information and Communication Technology for Development for Africa, ICT4DA 2022
EditoresEsubalew Alemneh, Ethiopia Nigussie, Fisseha Mekuria
EditorialInstitute of Electrical and Electronics Engineers Inc.
Páginas84-89
Número de páginas6
ISBN (versión digital)9781665455879
DOI
EstadoPublicada - 2022
Evento2022 International Conference on Information and Communication Technology for Development for Africa, ICT4DA 2022 - Bahir Dar, Etiopía
Duración: 28 nov. 202230 nov. 2022

Serie de la publicación

Nombre2022 International Conference on Information and Communication Technology for Development for Africa, ICT4DA 2022

Conferencia

Conferencia2022 International Conference on Information and Communication Technology for Development for Africa, ICT4DA 2022
País/TerritorioEtiopía
CiudadBahir Dar
Período28/11/2230/11/22

Huella

Profundice en los temas de investigación de 'The Effect of Normalization for Bi-directional Amharic-English Neural Machine Translation'. En conjunto forman una huella única.

Citar esto