Formalizing Feedback as Reusable Knowledge in Data Preparation Systems: A Persistent Feedback Artefact Model

Authors

  • Sa’adatu Abdulkadir Kaduna State University image/svg+xml
  • Philip Odion
  • Darius Chinyio
  • Isah Saidu
  • Muhammad Ahmad

DOI:

https://doi.org/10.33003/fjs-2026-1012-5272

Keywords:

Data Preparation, Feedback Persistence, Feedback Reuse, Feedback Artefacts, Knowledge Representation, Human-in-the-Loop

Abstract

Feedback-driven data preparation systems use user corrections to improve data quality and transformation accuracy. However, existing approaches typically treat feedback as a transient interaction that is consumed during execution and discarded after task completion. This limits the reuse of corrective knowledge across datasets, users, and preparation cycles. This study proposes a Persistent Feedback Artefact Model (PeFAM) as a conceptual and formal representation for treating feedback as a reusable computational knowledge object. The model represents feedback as a structured artefact comprising the original value, corrected value, contextual attributes, metadata, timestamp information, and provenance. A context-aware similarity mechanism is defined to guide how feedback artefacts may be retrieved and assessed for reuse across preparation tasks. The model was assessed through a worked analytical scenario that examined representational completeness, traceability, contextual matching, and threshold-based eligibility for human review. The study contributes a formal representation of feedback knowledge, a similarity-based reuse mechanism, and a conceptual arrangement of functions supporting persistent feedback management in data preparation environments. The analytical demonstration indicates that the model can support representation, traceability, contextual comparison, and threshold-based eligibility within the defined scenario, while empirical evaluation of efficiency, accuracy, usability, and scalability remains necessary.

References

Abdulkadir, S., Odion, P. O., Chinyio, D. T., Saidu, I. R., & Ahmad, M. A. (2026). Feedback-driven automation in data preparation: A systematic literature review. Science World Journal (SWJ), 21(1), 290-303. https://doi.org/10.4314/swj.v21i1.39

Azeroual, O. (2020). Data wrangling in database systems: Purging of dirty data. Data, 5(2), Article 50. https://doi.org/10.3390/data5020050

Fernandes, A. A. A., Koehler, M., Konstantinou, N., Pankin, P., Paton, N. W., & Sakellariou, R. (2023). Data preparation: A technological perspective and review. SN Computer Science, 4(4), Article 425. https://doi.org/10.1007/s42979-023-01828-8

Garba, M., Usman, M., & Saidu, M. (2025). Enhancing employee attrition prediction: The impact of data preprocessing on machine learning model performance. FUDMA Journal of Sciences, 9(1), 205–210. https://doi.org/10.33003/fjs-2025-0901-3030

Herschel, M., Diestelkämper, R., & Lahmar, H. B. (2017). A survey on provenance: What for? What form? What from? The VLDB Journal, 26(6), 881–906. https://doi.org/10.1007/s00778-017-0486-1

Konstantinou, N., Abel, E., Bellomarini, L., Bogatu, A., Civili, C., Irfanie, E., Koehler, M., Mazilu, L., Sallinger, E., Fernandes, A. A. A., Gottlob, G., Keane, J. A., & Paton, N. W. (2019). VADA: An architecture for end user informed data preparation. Journal of Big Data, 6(74), 1-32. https://doi.org/10.1186/s40537-019-0237-9

Konstantinou, N., & Paton, N. W. (2020). Feedback driven improvement of data preparation pipelines. Information Systems, 92, Article 101480. https://doi.org/10.1016/j.is.2019.101480

Liu, L., Hasegawa, S., Sampat, S. K., Xenochristou, M., Chen, W., Kato, T., Kakibuchi, T., & Asai, T. (2024). AutoDW: Automatic data wrangling leveraging large language models. In ASE ’24: Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (pp. 2041–2052). ACM. https://doi.org/10.1145/3691620.3695267

Narayan, A., Chami, I., Orr, L., & Ré, C. (2022). Can Foundation Models Wrangle Your Data? Proceedings of the VLDB Endowment, 16(4), 738-746. https://doi.org/10.14778/3574245.357425

Paton, N. (2019). Automating data preparation: Can we? Should we? Must we? Research Explorer, the University of Manchester. https://research.manchester.ac.uk/en/publications/automating-data-preparation-can-we-should-we-must-we/

Rezig, E. K., Ouzzani, M., Elmagarmid, A. K., Aref, W. G., & Stonebraker, M. (2019). Towards an End-to-End Human-Centric data cleaning framework. In HILDA ’19: Proceedings of the Workshop on Human-In-The-Loop Data Analytics (pp. 1–7). ACM. https://doi.org/10.1145/3328519.3329133

Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. M. (2021). “Everyone wants to do the model work, not the data work”: Data cascades in High-Stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (pp. 1–15). ACM. https://doi.org/10.1145/3411764.3445518

Shome, A., Cruz, L., Spinellis, D., & Arie, V. D. (2024). Understanding feedback mechanisms in Machine Learning JuPyter Notebooks. arXiv, 1-37. https://doi.org/10.48550/arxiv.2408.00153

Singh, J., Cobbe, J., & Norval, C. (2019). Decision Provenance: Harnessing Data Flow for Accountable Systems. IEEE Access, 7, 6562-6574. https://doi.org/10.1109/ACCESS.2018.2887201

Somasundaram, P. (2022). Streamlining Data Wrangling Processes through Automation and Tooling. International Journal of Science and Research (IJSR), 11(2), 1323–1325. https://doi.org/10.21275/sr24418101747

Wirth, R., & Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining (pp. 29–39).

Wu, X., Xiao, L., Sun, Y., Zhang, J., Ma, T., & He, L. (2023). A survey of human-in-the-loop for machine learning. Future Generation Computer Systems, 135, 364–381. https://doi.org/10.1016/j.future.2022.05.014

Yin, Y., Wang, B., Zuo, H., & Childs, P. (2024). Effects of different Human-In-The-Loop approaches on Human-AI Co-Design: A comparison between Human-Learning HITL approach and Machine-Learning HITL Approach. In Proceedings of the ASME 2024 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference (Vol. 2A, Article V02AT02A050). ASME. https://doi.org/10.1115/detc2024-143295

Lifecycle of a feedback artefact from creation to reviewed reuse and outcome recording

Downloads

Published

31-07-2026

How to Cite

Abdulkadir, S., Odion, P., Chinyio, D., Saidu, I., & Ahmad, M. (2026). Formalizing Feedback as Reusable Knowledge in Data Preparation Systems: A Persistent Feedback Artefact Model. FUDMA JOURNAL OF SCIENCES, 10(12), 242-251. https://doi.org/10.33003/fjs-2026-1012-5272

Most read articles by the same author(s)