Articles
| Open Access |
https://doi.org/10.55640/ijdsml-06-09-11
A Cloud-Native MLOps Framework for Drift Detection, Automated Retraining, and Reliable Real-Time Inference in
Sri Charan Chowdary Konidina , Independent Researcher, USAAbstract
Deployed predictive machine learning models inevitably degrade when real-world production data deviates from training distributions. Static deployment strategies cannot handle this environmental shift, causing drops in accuracy and forcing engineering teams to rely on slow, manual retraining audits. This paper presents an end-to-end, autonomous MLOps framework that bridges this gap by coupling continuous streaming telemetry with containerized retraining loops on Kubernetes. The architecture relies on an independent microservices design, utilizing Prometheus to track live data features and executing Kolmogorov-Smirnov and Population Stability Index (PSI) algorithms to catch statistical drift. We validated the system under a simulated fraud-detection workload streaming 50,000 requests per minute. Under testing, the framework flagged a 25% covariate shift within 2.5 minutes of its introduction. The automated pipeline completed retraining and package validation on an isolated compute plane, restoring predictive accuracy (F1-score > 0.93) within a total recovery window of 10.6 minutes. By implementing progressive canary rollouts for the live model swap, the system maintained a strict P99 real-time inference latency under 50ms with zero request drops or system downtime. These results demonstrate that tight architectural integration between telemetry and orchestration removes human-in-the-loop operational overhead and keeps live predictive systems resilient against distribution decay.
Keywords
MLOps, Cloud-Native Architecture, Concept Drift, Automated Retraining, Real-Time Inference, Systems Isolation, Model Observability
References
Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2019). Software engineering for machine learning: A case study. Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice, 291-300. https://doi.org/10.1109/ICSE-SEIP.2019.00042
Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys (CSUR), 46(4), 1-37. https://doi.org/10.1145/2523813
Garg, S., Gupta, A., & Malhotra, V. (2023). Real-time streaming model observability: Tracking non-parametric drift indicators via distributed logging layers. Journal of Production Machine Learning Systems, 7(2), 114-129.
Mäkinen, S., Skoglund, H., Meding, J., & Overen, T. (2021). Tooling gaps in cloud-native MLOps platforms: From static serving to adaptive architectures. IEEE Software, 38(5), 42-51. https://doi.org/10.1109/MS.2021.3058914
Raj, P., Kulkarni, A., & Srinivasan, S. (2022). Non-parametric evaluation of covariate shift in high-throughput enterprise data pipelines. IEEE Transactions on Knowledge and Data Engineering, 34(9), 4112-4125. https://doi.org/10.1109/TKDE.2022.3167092
Renggli, C., Rimanic, L., Gonen, N., & Zhang, C. (2023). Automated validation gates and progressive canary rollouts in governed machine learning workflows. Proceedings of the International Conference on Machine Learning and Systems (MLSys), 5, 204-218.
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, T., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems (NeurIPS), 28, 2503-2511.
Shankar, V., Menon, R., & Vasudevan, K. (2024). Operationalizing the EU AI Act: Lineage, traceability, and continuous audit strategies for autonomous model registries. AI & Society: Journal of Knowledge, Culture and Communication, 39(1), 85-102. https://doi.org/10.1007/s00146-023-01812-y
Symeonidis, G., Nerantzis, E., Spyromitros-Xioufis, E., Mitkas, P. A., & Vlahavas, I. (2023). Centralized feature stores in modern production MLOps: Mitigating training-serving skew in distributed clusters. ACM Computing Surveys, 55(11), 1-33. https://doi.org/10.1145/3569087
Tamburri, D. A. (2024). Sustainable MLOps: Patterns and anti-patterns in cloud-native adaptive retraining workflows. Journal of Systems and Software, 208, 111894. https://doi.org/10.1016/j.jss.2023.111894
Widmer, G., & Kubat, M. (1996). Learning in the presence of concept drift and hidden contexts. Machine Learning, 23(1), 69-101. https://doi.org/10.1007/BF00117446
Zhou, Y., Zhang, L., & Zhao, H. (2025). Quantifying resource contention and latency degradation in shared-plane cloud-native MLOps infrastructures. IEEE Transactions on Parallel and Distributed Systems, 36(4), 845-859. https://doi.org/10.1109/TPDS.2025.3409112
Article Statistics
Downloads
Copyright License
Copyright (c) 2026 Sri Charan Chowdary Konidina

This work is licensed under a Creative Commons Attribution 4.0 International License.