Weiyao Ma
Robert H. Smith School of Business, University of Maryland, College Park, MD 20742, USA.
*Corresponding author: Weiyao Ma
Abstract
In the context of the continuous expansion of data scale and highly digitized business systems, data infrastructure has gradually become an important technical foundation for supporting data analysis, intelligent decision-making, and real-time services. Traditional data processing chains that rely on scripts often exhibit loose structures, complex task dependencies, and insufficient fault recovery capabilities. Once a data processing node fails, data delays, result deviations, and service disruptions may spread along the processing chain. Automated data pipelines provide a new engineering approach for reconfiguring data processing systems. The core lies in leveraging a unified task orchestration system to organize data collection, transformation, and loading processes, while relying on metadata governance, system monitoring, and quality verification mechanisms to form a stable operating framework. The data infrastructure constructed through automated pipelines has a clearer data link structure and stronger operational resilience. The observability and fault recovery capabilities of the data processing process are enhanced, thereby forming a stable data supply environment and providing reliable support for the operation of complex data platforms.
References
[1] Xiaer X, Jialong C, Bangyi Z, et al. Research on safety resilience evaluation model of data center physical infrastructure: an ANP-based approach. Buildings. 2022;12(11):1911.
[2] Luke M. Injecting failure: data center infrastructures and the imaginaries of resilience. Inf Soc. 2020;36(3):167-176.
[3] Kayikci Y. Stream processing data decision model for higher environmental performance and resilience in sustainable logistics infrastructure. J Enterp Inf Manag. 2020;34(1):140-167.
[4] Zhang Y, Wang N, Zeng Q, et al. Automating data preparation pipeline efficiently via Monte Carlo tree search. Inf Sci. 2026;724: 122730.
[5] Yang H, Qingfeng L, Kaini Q, et al. PhiPipe: a multi-modal MRI data processing pipeline with test-retest reliability and predicative validity assessments. Hum Brain Mapp. 2022;44(5):2062-2084.
[6] Lyu W, Uranüs LA, Vadrot BA. From data rationales to data infrastructure: implications for the BBNJ clearing-house mechanism. Mar Policy. 2026;188:107079.
[7] Coran G, Costantini M, Gelsumini S, et al. Scalable and interoperable data management in the spoke 3 big data infrastructure. Astron Comput. 2026;55:101065.
[8] Porta V, Connolly FMA, Rod HN, et al. A qualitative comparison of data infrastructures for COVID-19 health-related data: lessons for the European Health Data Space. Policy Stud. 2026;47(2):338-358.
[9] Pankararu JC, Toneu TI, Odonne G, et al. A global biodiversity use data infrastructure acknowledging indigenous and local knowledge. NPJ Biodivers. 2026;5(1):7.
Copyright
© 2026 by the author(s).
This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial-NoDerivatives (CC BY-NC-ND) license, which permits non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited and is not modified or adapted.
https://creativecommons.org/licenses/by-nc-nd/4.0/
How to cite this paper
Reliable Data Infrastructure Supported by Automated Data Pipelines
How to cite this paper: Weiyao Ma. (2026). Reliable Data Infrastructure Supported by Automated Data Pipelines. Engineering Advances, 6(2), 126-130.
DOI: http://dx.doi.org/10.26855/ea.2026.06.009