Resource Sufficient Data Generation Technique for Machine Translation Models Using Named Entity Transliteration and Cross- Lingual Word Embeddings

Authors

  • Rabiu Ibrahim Abdullahi Division of Agricultural Colleges, Ahmadu Bello University,
  • Georgina Nkolika Obunadike
  • Adam Abdullahi
  • Martins E. Irhebhude

DOI:

https://doi.org/10.33003/fjs-2026-1020-6074

Keywords:

Natural language processing, Low-resource language, Hausa-English, Machine Translation

Abstract

Machine translation offers speed and cost advantages, but human translation remains superior in quality and contextual understanding. For low-resource languages like Hausa, the severe scarcity of parallel data leads to poor model performance, overfitting, and inadequate handling of named entities. While existing synthetic data generation methods such as back-translation and data copying introduce noise, this study proposes a technique that integrates named entity transliteration with cross-lingual word embeddings to generate synthetic parallel sentences. The approach replaces the common practice of entity copying with transliteration to preserve target-language vocabulary integrity, while embeddings translate non-entity words. Evaluated on the FLORES-200 and MAFAND-MT benchmarks, the unsupervised setting achievedBilingual Evaluation Understudy  (BLEU) scores of 1.04 (Hausa→English) and 1.42 (English→Hausa) on FLORES-200, while the semi-supervised setting improved substantially to 15.56 (Hausa→English) and 12.42 (English→Hausa) BLEU on MAFAND-MT outperforming the monolingual data copying baseline by up to 5 BLEU points. These results demonstrate that the proposed technique offers an effective strategy for augmenting scarce parallel data in low-resource MT systems.

References

Adelani, D. I., Abbott, J., Neubig, G., Daniel, D., Kreutzer, J., Lignos, C., Palen-michel, C., Buzaaba, H., Rijhwani, S., Ruder, S., Mayhew, S., & Azime, I. A. (2021). MasakhaNER: Named entity recognition for African languages. Transactions of the Association for Computational Linguistics, 9, 1116–1131. https://doi.org/10.1162/tacl_a_00408

Adelani, D. I., Alabi, J. O., Fan, A., Kreutzer, J., Shen, X., Reid, M., Ruiter, D., Klakow, D., Nabende, P., Chang, E., Gwadabe, T. R., Sackey, F., Dossou, B. F. P., Emezue, C. C., Leong, C., Beukman, M., Muhammad, S. H., Jarso, G. D., Yousuf, O., ... Manthalu, S. (2022). A few thousand translations go a long way! Leveraging pre-trained models for African news translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3053–3070. https://doi.org/10.18653/v1/2022.naacl-main.223

Au, T. W. T., Lampos, V., & Cox, I. J. (2022). E-NER: An annotated named entity recognition corpus of legal text. Proceedings of the Natural Legal Language Processing Workshop 2022, 246–255. https://doi.org/10.18653/v1/2022.nllp-1.22

Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Grave, E., Auli, M., & Joulin, A. (2021). Beyond English-centric multilingual machine translation. Journal of Machine Learning Research, 22, 1–48.

Goyal, N., Gao, C., Chaudhary, V., Chen, P. J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., & Fan, A. (2022). The FLORES-200 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10, 522–538. https://doi.org/10.1162/tacl_a_00474

Ibrahim, R. A., & Abdulmumin, I. (2022). NECAT-CLWE: A simple but efficient parallel data generation approach for low resource. AfricaNLP Workshop at the International Conference on Learning Representations (ICLR) 2022, 1–8.

Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., & Carlini, N. (2022). Deduplicating training data makes language models better. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1, 8424–8445. https://doi.org/10.18653/v1/2022.acl-long.577

Li, J., Sun, A., Han, J., & Li, C. (2022). A survey on deep learning for named entity recognition. IEEE Transactions on Knowledge and Data Engineering, 34(1), 50–70. https://doi.org/10.1109/TKDE.2020.2981314

Mager, M., Bhatnagar, R., Neubig, G., Thang, N., & Katharina, V. (2023). Neural machine translation for the indigenous languages of the Americas. Proceedings of the Association for Computational Linguistics, 109–133.

Muhammad, P. F., Kusumaningrum, R., & Wibowo, A. (2021). Sentiment analysis using Word2vec and long short-term memory (LSTM) for Indonesian hotel reviews. Procedia Computer Science, 179, 728–735. https://doi.org/10.1016/j.procs.2021.01.061

Muhammad, S. H., Ahmad, I. S., Abdulmumin, I., Lawan, F. I., Imam, S. H., Aliyu, Y., Sani, S. A., Umar, A. U., Gwadabe, T., Church, K., & Marivate, V. (2025). HausaNLP: Current status, challenges and future directions for Hausa natural language processing. Proceedings of the Sixth Workshop on African Natural Language Processing, 176–191. https://doi.org/10.18653/v1/2025.africanlp-1.27

Ojo, J., Ogueji, K., Stenetorp, P., & Ifeoluwa, D. (2023). How good are large language models on African languages? [Unpublished manuscript or conference paper].

Olojo, S., Zakrzewski, J., Smart, A., Van Liemt, E., Miceli, M., Ebinama, A., & Amugongo, L. M. (2025). Lost in machine translation: The sociocultural implications of language technologies in Nigeria. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, 1(1). https://doi.org/10.1145/3715275.3732153

Post, M. (2018). A call for clarity in reporting BLEU scores. Proceedings of the Third Conference on Machine Translation: Research Papers, 186–191.

Raja, R., & Vats, A. (2025). Parallel corpora for machine translation in low-resource Indic languages: A comprehensive review. Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages, 129–143. https://doi.org/10.18653/v1/2025.loresmt-1.12

Sharma, S., Diwakar, M., Singh, P., Singh, V., Kadry, S., & Kim, J. (2023). Machine translation systems based on classical-statistical-deep-learning approaches. Electronics, 12(7), 1–29. https://doi.org/10.3390/electronics12071716

Siu, S. C. (2023). ChatGPT and GPT-4 for professional translators: Exploring the potential of large language models in translation. SSRN Electronic Journal, 1–36. https://doi.org/10.2139/ssrn.4448091

Skenduli, M. P., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages. ACM Computing Surveys, 55(11), 1–37.

Tafa, T. O., Hashim, S. Z. M., Othman, M. S., Alhussian, H., Nasser, M., Abdulkadir, S. J., Huspi, S. H., Adeyemo, S. O., & Bena, Y. A. (2025). Machine translation performance for low-resource languages: A systematic literature review. IEEE Access, 13, 72486–72505. https://doi.org/10.1109/ACCESS.2025.3562918

Thai, K., Karpinska, M., Krishna, K., Ray, W., Inghilleri, M., Wieting, J., & Iyyer, M. (2022). Exploring document-level literary machine translation with parallel paragraphs from world literature. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9882–9902. https://doi.org/10.18653/v1/2022.emnlp-main.672

Zahorodna, O., Saienko, V., Tolchieva, H., Tymoshchuk, N., Kulinich, T., & Shvets, N. (2022). Developing communicative professional competence in future economic specialists in the conditions of postmodernism. Postmodern Openings, 13(2), 77–96.

chart of the data creation process

Downloads

Published

06-10-2026

How to Cite

Ibrahim Abdullahi, R., Obunadike, G. N., Abdullahi, A., & Irhebhude, M. E. (2026). Resource Sufficient Data Generation Technique for Machine Translation Models Using Named Entity Transliteration and Cross- Lingual Word Embeddings. FUDMA Journal of Sciences, 10(20), 30-36. https://doi.org/10.33003/fjs-2026-1020-6074