AI-Orchestrated Data Quality Control and End-to-End Lineage Management
DOI:
https://doi.org/10.5281/zenodo.22111485Keywords:
Cognitive Data Engineering, AI-Governed Data Pipelines, Data Quality Metrics and Evaluation, Automated Anomaly Detection and Correction, Metadata Standards, Internet-Scale Data Ecosystems, Data Lineage and Provenance, Compliance-Aware Data Architecture, Cost-Aware Pipeline Orchestration, AI Governance Frameworks, End-to-End Data Pipeline Management, Risk Identification and Quantification, Policy-Driven Data Engineering, Decision Rights in AI Systems, Data Provenance Acceleration, Intelligent Metadata Management, Regulated Data Pipelines, Autonomous Data Quality Management, Enterprise AI Governance, Scalable Data Orchestration.Abstract
Artificial intelligence (AI) is establishing itself as the next generation of technology. However, the data required to train such expandable, self-learning, powerful systems have so far been collected and pre-processed in a traditional manner. Moreover, because of the lack of governance in many such AI initiatives, these processes remain uncontrolled and often produce low-quality results. Both issues urgently need solution.
Three core principles of cognitive data engineering have been developed that are enabled by the recent expansion of AI technology. First, quality metrics and evaluation, as well as anomaly detection and correction mechanisms, have been formalized to provide a comprehensive AI-governed data-quality framework. Next, a set of metadata standards that define describe and affect Internet-scale data ecosystems is proposed. Their implementation provides a sophisticated method of capturing data lineage and provenance information and using that data for compliance and efficient data query acceleration. Third, the data pipeline architecture is designed to produce the definition and execution of complex data pipelines and orchestration in a cost-aware manner. These contributions enable an AI governance framework for complete data pipelines. Such a framework defines roles, policies, and decision rights to identify risks in the use of data pipelines, assess those risks quantitatively, and provide mitigation guidelines.
References
1. Ambe, R. (2024). Classification and quantification of timestamp data quality issues and its impact on data quality outcome. Data Intelligence, 6(3), 812–833.
2. Côté, P.-O., Nikanjam, A., Ahmed, N., Humeniuk, D., & Khomh, F. (2024). Data cleaning and machine learning: A systematic literature review. Automated Software Engineering, 31, Article 48.
3. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).
4. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.
5. Jakubik, J., Vössing, M., Kühl, N., Walk, J., & Satzger, G. (2024). Data-centric artificial intelligence. Business & Information Systems Engineering, 66, 507–515.
6. Inala, R. (2023). AI-powered investment decision support systems: Building smart data products with embedded governance controls. Journal for ReAttach Therapy and Developmental Diversities, 6(10), 2251-2266.
7. Latendresse, J., Abedu, S., Abdellatif, A., & Shihab, E. (2024). An exploratory study on machine learning model management. ACM Transactions on Software Engineering and Methodology, 34(1), Article 16.
8. Northcutt, C. G., Jiang, L., & Chuang, I. L. (2021). Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research, 70, 1373–1411.
9. Kolla, S. K., & Mangalampalli, B. M. (2024). Edge-Based Deep Learning Systems for Point-of-Care Diagnostic Intelligence. Journal of Neonatal Surgery, 13(1), 2387-2399.
10. Priestley, M., O’Donnell, F., & Simperl, E. (2023). A survey of data quality requirements that matter in ML development pipelines. Journal of Data and Information Quality, 15(2), Article 11.
11. Pushkarna, M., Zaldivar, A., & Kjartansson, O. (2022). Data cards: Purposeful and transparent dataset documentation for responsible AI. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1776–1826.
12. Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. M. (2021). “Everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–15.
13. Inala, R., & Somu, B. (2024). Agentic ai in retail banking: Redefining customer service and financial decision-making. Journal of Artificial Intelligence and Big Data Disciplines, 1(1), 1-19.
14. Shankar, S., Li, H., Asawa, P., Hulsebos, M., Lin, Y., Zamfirescu-Pereira, J. D., Chase, H., Fu-Hinthorn, W., Parameswaran, A. G., & Wu, E. (2024). spade: Synthesizing data quality assertions for large language model pipelines. Proceedings of the VLDB Endowment, 17(12), 4173–4186.
15. Wittner, M., et al. (2024). Toward a common standard for data and specimen provenance in life sciences. Learning Health Systems.
16. Kolla, S. K., & Reddy, V. A. R. (2024). Evaluating Cloud-Native vs. Hybrid Architectures for Health Benefit Administration Systems. International Journal of Medical Toxicology and Legal Medicine, 27(5), 1042-1053.
17. Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., & Hu, X. (2023). Data-centric AI: Perspectives and challenges. AI Magazine, 44(4), 452–464.
18. Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., Jiang, Z., Zhong, S., & Hu, X. (2023). Data-centric artificial intelligence: A survey. ACM Computing Surveys.
19. Gottimukkala, V. R. R. (2024). Federated Learning Approaches for Fraud Detection in International Payment Systems. https://www. jisem-journal. com/download/118_JISEM. pdf.
20. Thoutam, P. (2024). Automated data preparation through deep learning: A novel framework for intelligent data cleansing and standardization. International Journal of Scientific Research in Computer Science, Engineering and Information Technology.
21. Schneider, J., et al. (2023). Data-centric artificial intelligence: A systematic literature review. Data Science and Management, 6(3), 144–157.
22. Kim, H., et al. (2022). Data quality management and assessment for machine learning applications. Journal of Data and Information Quality.
23. Aramburu, M. J., et al. (2023). Data quality enhancement for machine learning: Assessing completeness, feature accuracy, and label accuracy. Data Mining and Knowledge Discovery.
24. Whang, S. E., Roh, Y., Song, H., & Lee, J.-G. (2023). Data collection and quality challenges in machine learning: A data-centric AI perspective. IEEE Data Engineering Bulletin, 46(4), 5–17.
25. Kolla, T. (2024). Graph Neural Networks for HCC Risk Adjustment and Interoperability. International Journal of Science, Research and Technology, 7(6), 13244-13255.
26. Sambasivan, N., et al. (2021). Understanding data work in machine learning systems: Data quality, documentation, and organizational practices. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2).
27. Suresh, H., & Guttag, J. V. (2021). A framework for understanding sources of harm throughout the machine learning life cycle. Proceedings of the 2021 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, 1–9.
28. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., & Gebru, T. (2021). Model cards for model reporting. Communications of the ACM, 64(12), 56–67.
29. Holland, S., Hosny, A., Newman, S., Joseph, J., & Chmielinski, K. (2021). The dataset nutrition label: A framework to drive higher data quality standards. IEEE Data Engineering Bulletin, 44(1), 5–14.
30. Hynes, N., Sculley, D., & Taylor, J. (2021). Data provenance and reproducibility in machine learning pipelines. Proceedings of the VLDB Endowment, 14(12), 3051–3064.
31. Schelter, S., Biessmann, F., Januschowski, T., Salinas, D., Seufert, S., & Szarvas, G. (2022). On challenges in machine learning model monitoring and data quality management. ACM Journal of Data and Information Quality, 14(2), 1–27.
32. Breck, E., Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2021). Data validation for machine learning. Proceedings of the VLDB Endowment, 14(12), 2861–2874.
33. Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2021). Data lifecycle management for machine learning systems. Proceedings of the VLDB Endowment, 14(12), 2875–2887.
34. Hulsebos, M., Demiralp, Ç., & others. (2023). Data quality and data lineage challenges in machine learning pipelines. IEEE Transactions on Knowledge and Data Engineering, 35(9), 8750–8764.
35. Van der Aalst, W. M. P. (2022). Process mining and data quality: Event data, provenance, and process intelligence. ACM Transactions on Management Information Systems, 13(4), 1–27.
36. Nargesian, F., Zhu, E., Miller, R. J., Pu, K. Q., & Arocena, P. C. (2021). Data lake management: Challenges and opportunities. Proceedings of the VLDB Endowment, 14(12), 2901–2914.
37. Papenbrock, T., & Naumann, F. (2022). Data profiling and data quality management for scalable data integration. ACM Journal of Data and Information Quality, 14(3), 1–26.
Additional Files
Published
Data Availability Statement
None
Issue
Section
License
Copyright (c) 2024 Katarzyna Nowak (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This work is licensed under a Creative Commons
Attribution 4.0 International License (CC BY 4.0).
You are free to:
- Share: copy and redistribute the material
- Adapt: remix, transform, and build upon the material
for any purpose, even commercially.
Under the following terms:
Attribution — You must give appropriate credit to
the original author(s) and source.
Full license text: https://creativecommons.org/licenses/by/4.0/