Ethical Safeguards in Autonomous Data Collection: Protecting Cultural Context and Privacy in Indigenous Language NLP
DOI:
https://doi.org/10.33003/fjs-2026-1015-5225Keywords:
Autonomous Data Collection, Ethical AI, Indigenous Data Governance, Large Language Models (LLMs), Multi-Agent SystemsAbstract
The increasing dependence of Natural Language Processing (NLP) systems on large-scale textual datasets has intensified the use of automated data collection methods such as web scraping, crawling, and autonomous extraction pipelines. While these approaches improve scalability and support the development of Large Language Models (LLMs), they also raise important ethical concerns, particularly in indigenous and low-resource language contexts. Existing data collection systems primarily prioritize data availability and technical efficiency, often overlooking issues related to privacy, cultural sensitivity, consent, and community-centered governance. Consequently, automated pipelines may collect personally identifiable information (PII), culturally sensitive narratives, and indigenous knowledge without adequate contextual understanding or ethical safeguards.
This paper proposes an ethical-by-design framework for autonomous indigenous language data collection using an LLM-based multi-agent architecture. The framework integrates ethical safeguards directly into system operations through context-aware filtering, hybrid PII detection, cultural sensitivity classification, and auditability mechanisms.
A Design Science Research methodology was adopted to develop and evaluate the framework using scenario-based validation involving privacy-sensitive and culturally sensitive textual data. The evaluation demonstrates that the proposed approach supports transparent and accountable data collection while reducing ethical risks associated with privacy violations and inappropriate extraction of culturally protected information.
This study contributes to ongoing research in responsible AI, indigenous data governance, and ethical NLP by demonstrating how ethical principles can be operationalized within autonomous data collection infrastructures. The proposed framework provides a foundation for the development of transparent, culturally aware, and privacy-conscious NLP data collection systems for indigenous and low-resource language technologies.
References
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.3445922
Birhane, A. (2020). Algorithmic colonization of Africa. SCRIPTed, 17(2), 389-409. https://script-ed.org/article/algorithmic-colonization-of-africa/
Carroll, S. R., Garba, I., Figueroa-Rodríguez, O. L., Holbrook, J., Lovett, R., Materechera, S., Parsons, M., Raseroka, K., Rodriguez-Lonebear, D., Rowe, R., Sara, R., Walker, J. D., Anderson, J., & Hudson, M. (2020). The CARE principles for Indigenous data governance. Data Science Journal, 19(1), 43. https://doi.org/10.5334/dsj-2020-043
Couldry, N., & Mejias, U. A. (2019). The costs of connection: How data is colonizing human life and appropriating it for capitalism. Stanford University Press.
Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., Luetge, C., Madelin, R., Pagallo, U., Rossi, F., Schafer, B., Valcke, P., & Vayena, E. (2018). AI4People-An ethical framework for a good AI society. Minds and Machines, 28(4), 689-707. https://doi.org/10.1007/s11023-018-9482-5
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723
Hutchinson, B. (2024). Modeling the sacred: Considerations for NLP on religious and culturally sensitive text. Findings of the Association for Computational Linguistics: NAACL 2024. https://aclanthology.org/2024.findings-naacl.65/
Kukutai, T., & Taylor, J. (2016). Indigenous data sovereignty: Toward an agenda. ANU Press. https://doi.org/10.22459/CAEPR38.11.2016
Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2). https://doi.org/10.1177/2053951716679679
Ricaurte, P. (2022). Ethics for the majority world: AI and the question of violence at scale. Media, Culture & Society, 44(4), 582-599. https://doi.org/10.1177/01634437221099612
Whittaker, M., Crawford, K., Dobbe, R., Fried, G., Kaziunas, E., Mathur, V., West, S. M., Richardson, R., Schultz, J., & Schwartz, O. (2018). AI Now Report 2018. AI Now Institute. https://ainowinstitute.org/reports.html
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Roseline Oghogho Osaseri, Agharase Rosemary Usiobaifo, Kaizer John Onyekachukwu

This work is licensed under a Creative Commons Attribution 4.0 International License.