Engineering Autonomous Data Ecosystems with AI and MLOps
DOI:
https://doi.org/10.5281/zenodo.22111443Keywords:
Autonomous Data Platforms, Intelligent Data Pipelines, Data Engineering Convergence, MLOps Integration, AI Service Platforms, Autonomous Analytics, Predictive Modeling Pipelines, Automated Model Deployment, Continuous Model Monitoring, Cloud-Native Data Platforms, Product-Oriented Data Architecture, Data Management Principles, Architectural Patterns for Data Platforms, Self-Managing Data Pipelines, AutoML and Forecasting Systems, Enterprise AI Enablement, CRM Analytics Automation, Risk Analytics Platforms, End-to-End ML Lifecycle, Autonomous Learning Systems.Abstract
An Autonomous Data Platform integrates data engineering, MLOps, and AI services into a single platform. The convergence is important because many business problems that require data analysis, predictive modeling, and monitoring can be realized as autonomous data pipelines. Organizations are struggling to establish best practices and standards for autonomous data platforms. The research considers the data management principles, conceptual models, and architectural patterns of data platforms from a product perspective. Emphasis is placed on the convergence with MLOps and cloud engineering. The use cases and evaluations of autonomous data platforms enabled by the convergence are examined.
With the proliferation of data and new generation artificial intelligence (AI) technologies, organizations are exploring new roles, processes, and technology products to groom the data and build data models for predictive analytics and forecasting. The autonomy of data pipelines is becoming popular as organizations increasingly require learners and predictors to be created, deployed, and monitored automatically. The concept of an Autonomous Data Platform describes a converged product combining data engineering, MLOps, and AI services within an organization. Autonomous Data Platforms can be realized and realized as intelligent data pipelines that groom data and support organizations in various business functions such as customer relationship management and risk analytics systems.
References
1. Tomarchio, O., Calcaterra, D., Di Modica, G., & Mazzaglia, P. (2021). TORCH: A TOSCA-based orchestrator of multi-cloud containerised applications. Journal of Grid Computing, 19, 5.
2. Nawrocki, P., Grzywacz, M., & Śnieżyński, B. (2021). Adaptive resource planning for cloud-based services using machine learning. Journal of Parallel and Distributed Computing, 152, 88–97.
3. Kolla, S. H. (2022). Strategic Information Integration Models for Cross-Functional Service Optimization in Large-Scale Enterprises. International Journal of Emerging Trends in Engineering and Management Research, 7(3), 11811.
4. Hernández, F., & Pérez, G. (2021). DISCERNER: Dynamic selection of resource manager in hyper-scale cloud-computing data centres. Future Generation Computer Systems, 116, 190–199.
5. Yeh, T., & Yu, S. (2022). Realizing dynamic resource orchestration on cloud systems in the cloud-to-edge continuum. Journal of Parallel and Distributed Computing, 160, 100–109.
6. Inala, R. (2023). Revolutionizing Customer Master Data in Insurance Technology Platforms: An AI and MDM Architecture Perspective. International Journal of Finance (IJFIN)-ABDC Journal Quality List, 36(6), 579-606.
7. Lin, J., Xie, D., Huang, J., Liao, Z., & Ye, L. (2022). A multi-dimensional extensible cloud-native service stack for enterprises. Journal of Cloud Computing, 11, 83.
8. Brogi, A., Carrasco, J., Durán, F., Pimentel, E., & Soldani, J. (2022). Self-healing trans-cloud applications. Computing, 104, 809–833.
9. ization.
10. a, S. K., & Reddy, V. A. R. (2023). Deep Learning Architectures For Multimodal Medical Data
11. Chen, J., Chen, P., Niu, X., Wu, Z., Xiong, L., & Shi, C. (2022). Task offloading in hybrid-decision-based multi-cloud computing network: A cooperative multi-agent deep reinforcement learning. Journal of Cloud Computing, 11, 90.
12. Dogani, J., Khunjush, F., & Seydali, M. (2022). K-AGRUED: A container autoscaling technique for cloud-based web applications in Kubernetes using attention-based GRU encoder-decoder. Journal of Grid Computing, 20, 40.
13. Mangalampalli, B. M. Generative AI Applications In Healthcare Data Mart Design And Optim
14. Carrión, C. (2022). Kubernetes scheduling: Taxonomy, ongoing issues and challenges. ACM Computing Surveys, 55(7), 138.
15. Injadat, M., Moubayed, A., Nassif, A. B., & Shami, A. (2022). Machine learning towards intelligent systems for cloud resource management: A review. Journal of Network and Computer Applications, 204, 103405.
16. Singh, B., Kaur, R., Woodside, M., & Chinneck, J. W. (2023). Low-power multi-cloud deployment of large distributed service applications with response-time constraints. Journal of Cloud Computing, 12, 1.
17. Amistapuram, K. (2023). Privacy-Preserving Machine Learning Models for Sensitive Customer Data in Insurance Systems. Educational Administration: Theory and Practice, 29(4), 5950-5958.
18. Ullah, A., Kiss, T., Kovács, J., Tusa, F., Deslauriers, J., Dagdeviren, H., Arjun, R., & others. (2023). Orchestration in the Cloud-to-Things compute continuum: Taxonomy, survey and future directions. Journal of Cloud Computing, 12, 135.
19. Senjab, K., Abbas, S., Ahmed, N., & Khan, A. U. R. (2023). A survey of Kubernetes scheduling algorithms. Journal of Cloud Computing, 12, 87.
20. Carlini, E., Coppola, M., Dazzi, P., Ferrucci, L., Kavalionak, H., Korontanis, I., Mordacchini, M., & Tserpes, K. (2023). SmartORC: Smart orchestration of resources in the compute continuum. Frontiers in High Performance Computing, 1, 1164915.
21. Yandamuri, U. S. (2023). An Intelligent Analytics Framework Combining Big Data and Machine Learning for Business Forecasting. Zenodo.
22. Tusa, F., & Clayman, S. (2023). End-to-end slices to orchestrate resources and services in the cloud-to-edge continuum. Future Generation Computer Systems, 141, 473–488.
23. Alyas, T., et al. (2023). Optimizing resource allocation framework for multi-cloud environment. Computers, Materials & Continua, 75(2), 4119–4136.
24. Shan, C., Wu, C., Xia, Y., Guo, Z., Liu, D., & Zhang, J. (2023). Adaptive resource allocation for workflow containerization on Kubernetes. Journal of Systems Engineering and Electronics, 34.
25. Inala, R. (2023). AI-powered investment decision support systems: Building smart data products with embedded governance controls. Journal for ReAttach Therapy and Developmental Diversities, 6(10), 2251-2266.
26. Purahong, B., Sithiyopasakul, J., Sithiyopasakul, P., Lasakul, A., & Benjangkaprasert, C. (2023). Automated resource management system based upon container orchestration tools comparison. Journal of Advances in Information Technology, 14(3), 501–509.
27. Senjab, K., Abbas, S., Ahmed, N., & Khan, A. U. R. (2023). Container orchestration and scheduling for cloud-native applications: A systematic review. Journal of Cloud Computing, 12.
28. Ullah, A., Kiss, T., Kovács, J., Tusa, F., Deslauriers, J., Dagdeviren, H., & Arjun, R. (2023). Resource management and application orchestration across the cloud-to-edge continuum. Journal of Cloud Computing, 12.
29. Inala, R. Designing Scalable Technology Architectures for Customer Data in Group Insurance and Investment Platforms.
30. Malla, S., & Christensen, K. (2021). Kubernetes scheduling and resource management for cloud-native applications. IEEE International Conference on Cloud Computing.
31. Shi, Z., Jiang, C., Jiang, L., & Liu, X. (2021). HPKS: High performance Kubernetes scheduling for dynamic blockchain workloads in cloud computing. IEEE International Conference on Cloud Computing, 456–466.
32. Wang, Y., Zhang, X., Zhang, J., & others. (2021). Container scheduling techniques: A survey and assessment. Journal of King Saud University—Computer and Information Sciences, 34.
33. Chen, L., Liu, Y., Zhang, Y., & Wang, H. (2021). Machine learning-based resource scheduling for cloud computing environments. Future Generation Computer Systems, 118.
34. Nawrocki, P., Grzywacz, M., & Śnieżyński, B. (2021). Machine learning approaches for adaptive resource provisioning in cloud computing. Journal of Parallel and Distributed Computing, 152.
35. KollIntegration. South Eastern European Journal of Public Health, 248–260.
36. Fu, Y., Machlovi, N., Mao, Y., Wang, J., & Cheng, L. (2022). Performance evaluation of resource management schemes for cloud native platforms with computing containers. IEEE International Performance, Computing, and Communications Conference, 414–415.
37. Di Nitto, E., & Vladušič, D. (2022). Orchestrating heterogeneous applications: Motivation and state of the art. In E. Di Nitto, E. Gorroñogoitia Cruz, I. Kumara, D. Radolović, K. Tokmakov, & Z. Vasileiou (Eds.), Deployment and operation of complex software in heterogeneous execution environments. Springer.
38. Di Nitto, E., Vladušič, D., & others. (2022). The SODALITE runtime environment. In Deployment and operation of complex software in heterogeneous execution environments (pp. 67–92). Springer.
39. Chen, J., Chen, P., Niu, X., Wu, Z., Xiong, L., & Shi, C. (2022). Cooperative deep reinforcement learning for resource-aware computation offloading in multi-cloud environments. Journal of Cloud Computing, 11.
40. Ahmed, A., Srirama, S. N., & others. (2023). Intelligent resource orchestration and management in distributed cloud-native environments. Future Generation Computer Systems, 141.
Additional Files
Published
Data Availability Statement
none
Issue
Section
License
Copyright (c) 2023 Hiroshi Tanaka (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This work is licensed under a Creative Commons
Attribution 4.0 International License (CC BY 4.0).
You are free to:
- Share: copy and redistribute the material
- Adapt: remix, transform, and build upon the material
for any purpose, even commercially.
Under the following terms:
Attribution — You must give appropriate credit to
the original author(s) and source.
Full license text: https://creativecommons.org/licenses/by/4.0/