Development of a Transformer-Based Neural Machine Translation System for English to Ebira Language
DOI:
https://doi.org/10.33003/fjs-2026-1011-5311Keywords:
Neural Machine Translation, Low-Resource Languages, Transformer Architecture, Fine-Tuning, Language PreservationAbstract
Transformer-based neural machine translation (NMT) models have boosted translation accuracy for high-resource languages; however, research has largely bypassed unwritten and low-resource languages, particularly African languages such as Ebira. Ebira is an unwritten, low-resource language spoken by approximately 2.5 million people predominantly in Kogi State, Nigeria. Existing Ebira machine translation (MT) systems suffer from poor fluency, accuracy, and missed nuances, constrained by small datasets and rule-based methods. This study presents the development of a neural machine translation (NMT) system for English-to-Ebira translation using Google’s T5-base transformer model. A bilingual parallel corpus of 32,322 English-Ebira sentence pairs was compiled and used to fine-tune the model. The system achieved a corpus-level BLEU score of 40.95%, with 87% of evaluated sentences scoring 0.5 BLEU or higher, surpassing the prior rule-based system’s threshold result of 81.50%, corresponding to 6.75% relative improvement. Human evaluation by ten native Ebira speakers yielded a mean rating of 8.33/10 for fluency, accuracy, and cultural relevance. This research demonstrated that the application of transfer learning on transformer NMT model significantly improves the quality of (MT) systems; and also provides a foundational step for the development of computational resources for Ebira and supports the broader goal of linguistic inclusivity in artificial intelligence.
References
Abdulmusawir, T. M., Ayegba, S. F., Kayode, Y. M., & Eze, C. C. (2021). A system for machine translation from English to Ebira using the rule-based approach. Journal of Scientific Research and Reports, 27(11), 137–148. https://doi.org/10.9734/jsrr/2021/v27i1130465
Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. International Conference on Learning Representations (ICLR).
Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgements. ACL Workshop on Evaluation Measures for MT, 65–72.
Ekle, O. A., & Das, B. (2025). Low-resource neural machine translation using recurrent neural networks and transfer learning: A case study on English-to-Igbo. arXiv:2504.17252.
Etuk, S. O., Oyefolahan, I. O., Zubairu, H. A., Babakano, F. J., and Momoh, L. O. (2020). Development of an English-to-Ebira and Ebira-to-English machine translation system. International Journal of Information Processing and Communication, 6(1), 1–12.
Joshua Project. (2024). Ebira in Nigeria people group profile. https://joshuaproject.net/people_groups/12188
Lin, P.-J., Saeed, M., Chang, E., & Scholman, M. (2023). Low-resource cross-lingual adaptive training for Nigerian Pidgin. Interspeech 2023, 3950–3954.
Makoji, E., & Sani, F. (2025). Development and evaluation of an English-to-Igala NMT system using deep learning. International Journal of Innovative Science and Research Technology, 10(5). https://doi.org/10.38124/ijisrt/25may556
Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Kolawole, T., Akinadan, F., Adewole, D., Yamahata, K., & Hsu, C.-J. (2020). Participatory research for low-resourced machine translation: A case study in African languages. Findings of the Association for Computational Linguistics: EMNLP 2020, 2144–2160. https://doi.org/10.18653/v1/2020.findings-emnlp.195
Oluwatoki, T. G., Adetunmbi, A. O., Boyinbode, O. K., et al. (2021). Machine translation systems for Nigerian indigenous languages: A statistical overview. Nigerian Journal of Technological Research, 16(2), 41–52.
Orife, I. (2020). Neural machine translation for Edoid languages. In AfricaNLP Workshop, ICLR 2020.
Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of ACL 2002, 311–318.
Popovic, M. (2015). ChrF: Character n-gram F-score for automatic MT evaluation. Workshop on Statistical MT, 392–395.
Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67.
Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. ACL 2016, 86–96.
Tijani, M. A., Abubakar, A. H., & Donfack Kana, A. F. (2025). Leveraging data triangulation technique for the development of unwritten language parallel corpora for NLP applications. In Proceedings of ICCAIT2025, 234–241.
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. In Advances in NeurIPS 30, 5998–6008.
Xue, L., Constant, N., Roberts, A., et al. (2021). mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of NAACL-HLT 2021, 483–498.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Musari Abdulmusawir Tijani, Amina Hassan Abubakar, Armand Florentin Donfack Kana, Mohammed Abdullahi

This work is licensed under a Creative Commons Attribution 4.0 International License.