An LLM-Driven Framework for Addressing Code-Switching and Orthographic Variance in Nigerian Language Data Collection
DOI:
https://doi.org/10.33003/fjs-2026-1012-5226Keywords:
Code-Switching, Large Language Model, Orthographic Variance, Natural Language Processing, Data Preprocessing, Low-Resource Languages, Nigerian LanguagesAbstract
Natural Language Processing (NLP) systems for low-resource languages continue to face significant challenges due to the poor quality and structural inconsistency of available textual data. In the Nigerian linguistic context, informal digital communication is characterized by frequent code-switching between English, Nigerian Pidgin, and indigenous languages, as well as high levels of orthographic variance. These characteristics reduce the effectiveness of conventional preprocessing pipelines, which are typically designed for monolingual and standardized text. Existing approaches based on rule-based filtering, statistical methods, or conventional language identification models often fail to accurately interpret multilingual and inconsistent language patterns.
This paper presents a reasoning-driven preprocessing framework for improving Nigerian language data collection for NLP systems. The proposed approach leverages Large Language Models (LLMs) to perform context-aware language identification, relevance filtering, and orthographic normalization. Unlike traditional preprocessing methods that rely primarily on lexical or statistical matching, the framework treats language identification and normalization as contextual reasoning tasks. The system is designed to preserve semantic meaning while reducing spelling variability and dataset noise.
To evaluate the framework, textual data was collected from informal online sources, including blogs, forums, and social media platforms containing multilingual Nigerian language content. Experimental results demonstrate that the proposed approach achieved an 85% reduction in noisy data while improving the handling of code-switched and orthographically inconsistent text. The findings highlight the potential of LLM-based preprocessing for enhancing data quality in low-resource NLP environments and provide a scalable foundation for developing more robust and inclusive language technologies for Nigerian languages.
References
Adelani, D. I., Abbott, J., Neubig, G., et al. (2023). AfroLM: A self-active learning framework for low-resource African languages. Transactions of the Association for Computational Linguistics, 11, 1301-1317. https://aclanthology.org/2023.tacl-1.75/
Adebara, I. O., & Abdul-Mageed, M. (2022). Towards advancing natural language processing for African languages. arXiv preprint. https://arxiv.org/abs/2208.08200
Alhafni, B. (2025). Orthographic normalization for low-resource and dialectal languages. Journal of Natural Language Engineering, 31(2), 145-162.
Chakravarthi, B. R., Rani, P., Arcan, M., & McCrae, J. P. (2021). A survey of orthographic information in machine translation. SN Computer Science, 2(5), 1-19. https://doi.org/10.1007/s42979-021-00723-4
Das, R., Ranjan, S., Pathak, S., & Jyothi, P. (2023). Improving pretraining techniques for code-switched NLP. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), 1043-1058. https://aclanthology.org/2023.acl-long.66/
Diwan, A., Vaideeswaran, R., Shah, S., & Singh, A. (2021). Multilingual and code-switching automatic speech recognition for low-resource languages. arXiv preprint. https://arxiv.org/abs/2104.00235
Nazir, M. K., et al. (2025). Multilingual transformer models for code-mixed data processing. IEEE Access, 13, 145678-145690. https://ieeexplore.ieee.org/document/10835765/
Nguyen, L., Bryant, C., & Mayeux, O. (2023). Machine translation for low-resource code-switching. Findings of the Association for Computational Linguistics (ACL 2023), 893-905. https://aclanthology.org/2023.findings-acl.893/
Pakray, P., & Gelbukh, A. (2025). Natural language processing applications for low-resource languages. Natural Language Engineering. Cambridge University Press. https://www.cambridge.org/core/journals/natural-language-engineering
Sabty, C. (2024). Computational approaches to code-switching in natural language processing. arXiv preprint. https://arxiv.org/abs/2410.13318
Zhong, Y., Li, H., & Chen, X. (2024). Context-aware language modeling for multilingual low-resource NLP systems. Proceedings of the International Conference on Computational Linguistics, 455-468.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Roseline Oghogho Osaseri, Agharase Rosemary Usiobaifo, Kaizer John Onyekachukwu

This work is licensed under a Creative Commons Attribution 4.0 International License.