magazinelogo

Advances in Computer and Communication

ISSN Online: 2767-2875 CODEN: ACCDC3
Frequency: quarterly Email: acc@hillpublisher.com
Total View: 1274020 Downloads: 257845 Citations: 204 (From Dimensions)
ArticleOpen Access http://dx.doi.org/10.26855/acc.2026.09.010

Research on Parallel Computing and Runtime Scheduling Mechanisms for Large Language Models in High-concurrency Scenarios

Yiling Yuan

Information Networking Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA.

*Corresponding author: Yiling Yuan

Published: July 24, 2026

Abstract

High-concurrency situations demand stringent constraints on the parallel computing efficiency of large language models. Typically, complex load distributions and imperfect parallel scheduling architectures result in resource wastage and increased inference latency during model deployment. This paper comprehensively investigates the operational principles of large language model inference under a high volume of concurrent requests and examines the adaptability limitations of major parallel computing schemes in high-speed service environments. To address these challenges, a multi-dimensional hybrid parallel computing structure is designed to overcome the limitations of single parallel computing approaches, and an adaptive granularity optimization method at the model inference level is proposed to accommodate varying task complexities. Additionally, a layered resource allocation system is established to manage heterogeneous resources hierarchically. Concurrently, a dynamic load-sensitive scheduling decision model and a multi-level request priority sorting mechanism are developed to enable real-time resource scheduling and dynamic migration support. The results demonstrate a significant reduction in performance bottlenecks associated with parallel computing for large models under high concurrency, thereby enhancing the overall throughput and stability of industrial model applications.

Keyword

Large language model; high concurrency; parallel computing; dynamic scheduling

References

[1] Koevorden VJ, Aben N, Struben V, et al. Validating large language model–assisted data extraction from clinical notes. ESMO Real World Data Digit Oncol. 2026;12:100718.

[2] Jones G, Williams J, Berg T, et al. Generative large language models for predictive maintenance planning. Comput Ind Eng. 2026;218:112095.

[3] Li T, Ge X, Wang Z, et al. Multimodal large language model (MLLM) benchmark for intelligent construction in underground engineering. Autom Constr. 2026;188:106997.

[4] Sulaiman HM, Muda N. Automated triaging of hospital complaints using large language model-assisted content analysis and machine learning. Eng Appl Artif Intell. 2026;178(Pt 2):114942.

[5] Nosrati K, Tepljakov A, Belikov J, et al. When control meets large language models: from words to dynamics. Eng Appl Artif Intell. 2026;178(Pt 2):115119.

[6] Elhosary E, Moselhi O. Knowledge-augmented large language models to support automated HAZOP report generation. Process Saf Environ Prot. 2026;213:108997.

[7] Zhou G, Chen C, Qiu P, et al. CCJA: context-coherent jailbreak attack for aligned large language models. Mach Learn. 2026; 115(6):132.

[8] Lieslehto J, Tiihonen J, Lähteenvuo M, et al. Large language model approach to uncover reasoning patterns in forensic psychiatric assessment. Sci Rep. 2026.
DOI:10.1038/s41598-026-53275-z.

Copyright

© 2026 by the author(s).
This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial-NoDerivatives (CC BY-NC-ND) license, which permits non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited and is not modified or adapted.
https://creativecommons.org/licenses/by-nc-nd/4.0/

How to cite this paper

Research on Parallel Computing and Runtime Scheduling Mechanisms for Large Language Models in High-concurrency Scenarios

How to cite this paper: Yiling Yuan. (2026) Research on Parallel Computing and Runtime Scheduling Mechanisms for Large Language Models in High-concurrency Scenarios. Advances in Computer and Communication7(3), 163-166.

DOI: http://dx.doi.org/10.26855/acc.2026.09.010