An LLM-Driven Framework for Addressing Code-Switching and Orthographic Variance in Nigerian Language Data Collection

Authors

  • Roseline Oghogho Osaseri Dr
  • Agharase Rosemary Usiobaifo
  • Kaizer John Onyekachukwu

DOI:

https://doi.org/10.33003/fjs-2026-1012-5226

Keywords:

Code-Switching, Large Language Model, Orthographic Variance, Natural Language Processing, Data Preprocessing, Low-Resource Languages, Nigerian Languages

Abstract

Natural Language Processing (NLP) systems for low-resource languages continue to face significant challenges due to the poor quality and structural inconsistency of available textual data. In the Nigerian linguistic context, informal digital communication is characterized by frequent code-switching between English, Nigerian Pidgin, and indigenous languages, as well as high levels of orthographic variance. These characteristics reduce the effectiveness of conventional preprocessing pipelines, which are typically designed for monolingual and standardized text. Existing approaches based on rule-based filtering, statistical methods, or conventional language identification models often fail to accurately interpret multilingual and inconsistent language patterns.

This paper presents a reasoning-driven preprocessing framework for improving Nigerian language data collection for NLP systems. The proposed approach leverages Large Language Models (LLMs) to perform context-aware language identification, relevance filtering, and orthographic normalization. Unlike traditional preprocessing methods that rely primarily on lexical or statistical matching, the framework treats language identification and normalization as contextual reasoning tasks. The system is designed to preserve semantic meaning while reducing spelling variability and dataset noise.

To evaluate the framework, textual data was collected from informal online sources, including blogs, forums, and social media platforms containing multilingual Nigerian language content. Experimental results demonstrate that the proposed approach achieved an 85% reduction in noisy data while improving the handling of code-switched and orthographically inconsistent text. The findings highlight the potential of LLM-based preprocessing for enhancing data quality in low-resource NLP environments and provide a scalable foundation for developing more robust and inclusive language technologies for Nigerian languages.

Author Biography

  • Agharase Rosemary Usiobaifo

    Department of Computer Science, Faculty of Computing, University of Benin.

References

Adelani, D. I., Abbott, J., Neubig, G., et al. (2023). AfroLM: A self-active learning framework for low-resource African languages. Transactions of the Association for Computational Linguistics, 11, 1301-1317. https://aclanthology.org/2023.tacl-1.75/

Adebara, I. O., & Abdul-Mageed, M. (2022). Towards advancing natural language processing for African languages. arXiv preprint. https://arxiv.org/abs/2208.08200

Alhafni, B. (2025). Orthographic normalization for low-resource and dialectal languages. Journal of Natural Language Engineering, 31(2), 145-162.

Chakravarthi, B. R., Rani, P., Arcan, M., & McCrae, J. P. (2021). A survey of orthographic information in machine translation. SN Computer Science, 2(5), 1-19. https://doi.org/10.1007/s42979-021-00723-4

Das, R., Ranjan, S., Pathak, S., & Jyothi, P. (2023). Improving pretraining techniques for code-switched NLP. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), 1043-1058. https://aclanthology.org/2023.acl-long.66/

Diwan, A., Vaideeswaran, R., Shah, S., & Singh, A. (2021). Multilingual and code-switching automatic speech recognition for low-resource languages. arXiv preprint. https://arxiv.org/abs/2104.00235

Nazir, M. K., et al. (2025). Multilingual transformer models for code-mixed data processing. IEEE Access, 13, 145678-145690. https://ieeexplore.ieee.org/document/10835765/

Nguyen, L., Bryant, C., & Mayeux, O. (2023). Machine translation for low-resource code-switching. Findings of the Association for Computational Linguistics (ACL 2023), 893-905. https://aclanthology.org/2023.findings-acl.893/

Pakray, P., & Gelbukh, A. (2025). Natural language processing applications for low-resource languages. Natural Language Engineering. Cambridge University Press. https://www.cambridge.org/core/journals/natural-language-engineering

Sabty, C. (2024). Computational approaches to code-switching in natural language processing. arXiv preprint. https://arxiv.org/abs/2410.13318

Zhong, Y., Li, H., & Chen, X. (2024). Context-aware language modeling for multilingual low-resource NLP systems. Proceedings of the International Conference on Computational Linguistics, 455-468.

Architecture of the proposed LLM-driven preprocessing framework for multilingual Nigerian language data

Downloads

Published

24-07-2026

How to Cite

Osaseri, R. O., Usiobaifo, A. R., & Onyekachukwu, K. J. (2026). An LLM-Driven Framework for Addressing Code-Switching and Orthographic Variance in Nigerian Language Data Collection. FUDMA JOURNAL OF SCIENCES, 10(12), 53-58. https://doi.org/10.33003/fjs-2026-1012-5226

Most read articles by the same author(s)