Computing Resource-Aware Operation Optimization Strategy for MPI Jobs in Cloud-Native Environment

Authors

  • Wenxiao Wang State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023, China Author
  • Zibo Gao State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023, China Author
  • Guoding Ji State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023, China Author
  • Chentian Yong SINOPEC Geophysical Research Institute, Nanjing 211103, China Author
  • Bo Li SINOPEC Geophysical Research Institute, Nanjing 211103, China Author

DOI:

https://doi.org/10.64509/jicn.23.137

Keywords:

Cloud-Native, MPI, Resource-Aware Optimization, NUMA Affinity, Dynamic Parallelism

Abstract

Large-scale high-performance computing workloads in petroleum geophysical exploration commonly use the Message Passing Interface (MPI) as their parallel programming model. However, MPI does not provide native mechanisms for resource management, which complicates the efficient execution of multiple MPI jobs in shared-resource environments. As MPI workloads migrate to cloud-native platforms, their high degrees of parallelism and communication-intensive behavior may lead to scheduling delays and contention for shared resources, resulting in prolonged execution time and inefficient resource utilization. This paper identifies two major limitations of existing cloud-native MPI deployments: cluster-unaware process-count selection and insufficient exploitation of Non-Uniform Memory Access (NUMA) locality. To address these limitations, we propose a resource-aware optimization framework that dynamically selects the MPI process count and performs node- and NUMA-aware process placement. Experimental results show that the proposed parallelism-selection method reduces task-sequence execution time by at least 18% and improves the evaluated resource-utilization metrics by more than 30%. The topology-aware placement method further reduces execution time by at least 10% compared with the evaluated affinity baselines under the tested shared-resource cluster configurations.

Downloads

Download data is not yet available.

References

[1] Baysal, E., Kosloff, D.D., Sherwood, J.W.C.: Reverse Time Migration. Geophysics 48(11), 1514-1524 (1983). https://doi.org/10.1190/1.1441434 DOI: https://doi.org/10.1190/1.1441434

[2] Zhou, H.-W., Hu, H., Zou, Z., Wo, Y., Youn, O.: Reverse Time Migration: A Prospect of Seismic Imaging Methodology. Earth-Science Reviews 179, 207-227 (2018). https://doi.org/10.1016/j.earscirev.2018.02.008 DOI: https://doi.org/10.1016/j.earscirev.2018.02.008

[3] Araya-Polo, M., Cabezas, J., Hanzich, M., Pericas, M., Rubio, F., Gelado, I., Shafiq, M., Morancho, E., Navarro, N., Ayguade, E., et al.: Assessing Accelerator-Based HPC Reverse Time Migration. IEEE Transactions on Parallel and Distributed Systems 22(1), 147-162 (2011). https://doi.org/10.1109/TPDS.2010.144 DOI: https://doi.org/10.1109/TPDS.2010.144

[4] Gabriel, E., Fagg, G.E., Bosilca, G., Angskun, T., Dongarra, J.J., Squyres, J.M., Sahay, V., Kambadur, P., Barrett, B., Lumsdaine, A., et al.: Open MPI: Goals, Concept, and Design of a Next Generation MPI Implementation. In Recent Advances in Parallel Virtual Machine and Message Passing Interface: the 11th European PVM/MPI Users' Group Meeting (Euro PVM/MPI 2004), pp. 97-104 (2004). https://doi.org/10.1007/978-3-540-30218-6_19 DOI: https://doi.org/10.1007/978-3-540-30218-6_19

[5] https://www.docker.com/ Accessed 2026-07-22

[6] https://kubernetes.io/ Accessed 2026-07-22

[7] https://spark.apache.org/ Accessed 2026-07-22

[8] Huang, X., Gu, R., Huang, Y.: Towards Efficient Serverless MapReduce Computing on Cloud-Native Platforms. Big Data Mining and Analytics 8(3), 575-591 (2025). https://doi.org/10.26599/BDMA.2024.9020084 DOI: https://doi.org/10.26599/BDMA.2024.9020084

[9] Zhang, X., Wang, X., Shang, J., et al.: Efficient Serverless Stream Processing Based on High-Level Programming and Parallelism AutoTuning. In the 2024 IEEE International Conference on High Performance Computing and Communications (HPCC 2024), pp. 474-481 (2024). https://doi.org/10.1109/HPCC64274.2024.00070 DOI: https://doi.org/10.1109/HPCC64274.2024.00070

[10] Kubeflow: Kubeflow MPI Operator. https://github.com/kubeflow/mpi-operator Accessed 2026-07-22

[11] Kubernetes: Kubernetes Documentation: Assigning Pods to Nodes. https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/ Accessed 2026-07-22

[12] Kubernetes: Kubernetes Documentation: Control Topology Management Policies on a Node. https://kubernetes.io/docs/tasks/administer-cluster/topology-manager/ Accessed 2026-07-22

[13] Mujkanovic, N., Durillo, J.J., Hammer, N., Müller, T.: Survey of Adaptive Containerization Architectures for HPC. In Proceedings of the SC '23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, pp. 165-176 (2023). https://doi.org/10.1145/3624062.3624588 DOI: https://doi.org/10.1145/3624062.3624588

[14] Beltre, A.M., Saha, P., Govindaraju, M., Younge, A., Grant, R.E.: Enabling HPC Workloads on Cloud Infrastructure Using Kubernetes Container Orchestration Mechanisms. In 2019 IEEE/ACM International Workshop on Containers and New Orchestration Paradigms for Isolated Environments in HPC (CANOPIE-HPC), pp. 11-20 (2019). https://doi.org/10.1109/CANOPIE-HPC49598.2019.00007 DOI: https://doi.org/10.1109/CANOPIE-HPC49598.2019.00007

[15] Hursey, J.: Design Considerations for Building and Running Containerized MPI Applications. In the 2020 2nd International Workshop on Containers and New Orchestration Paradigms for Isolated Environments in HPC (CANOPIE-HPC 2020), pp. 35-44 (2020). https://doi.org/10.1109/CANOPIEHPC51917.2020.00010 DOI: https://doi.org/10.1109/CANOPIEHPC51917.2020.00010

[16] Zhou, N.: Containerization and Orchestration on HPC Systems. In Sustained Simulation Performance 2019 and 2020: the Joint Workshop on Sustained Simulation Performance (PJWSSP 2019/2020), pp. 133-147 (2021). https://doi.org/10.1007/978-3-030-68049-7_10 DOI: https://doi.org/10.1007/978-3-030-68049-7_10

[17] Zhang, J., Lu, X., Panda, D.K.: High Performance MPI Library for Container-Based HPC Cloud on InfiniBand Clusters. In 2016 45th International Conference on Parallel Processing (ICPP), pp. 268-277 (2016). https://doi.org/10.1109/ICPP.2016.38 DOI: https://doi.org/10.1109/ICPP.2016.38

[18] Liu, P., Guitart, J.: Performance Comparison of Multi-Container Deployment Schemes for HPC Workloads: An Empirical Study. The Journal of Supercomputing 77(6), 6273-6312 (2021). https://doi.org/10.1007/s11227-020-03518-1 DOI: https://doi.org/10.1007/s11227-020-03518-1

[19] Denis, A., Jaeger, J., Jeannot, E., Perache, M., Taboada, H.: Study on Progress Threads Placement and Dedicated Cores for Overlapping MPI Nonblocking Collectives on Manycore Processor. The International Journal of High Performance Computing Applications 33(6), 1240-1254 (2019). https://doi.org/10.1177/1094342019860184 DOI: https://doi.org/10.1177/1094342019860184

[20] Gallardo, E., Vienne, J., Fialho, L., Teller, P., Browne, J.: MPI Advisor: A Minimal Overhead Tool for MPI Library Performance Tuning. In Proceedings of the 22nd European MPI Users' Group Meeting, pp. 1-10 (2015). https://doi.org/10.1145/2802658.2802667 DOI: https://doi.org/10.1145/2802658.2802667

[21] Ashwini, J.P., Sanjay, H.A., Nayana, M.C., Shastry, K.A., K, M.M.M.: Framework for Performance Enhancement of MPI-Based Application on Cloud. Scalable Computing: Practice and Experience 24(1), 55-67 (2023). https://doi.org/10.12694/scpe.v24i1.2086 DOI: https://doi.org/10.12694/scpe.v24i1.2086

[22] Feng, Y., Yuan, J., Liu, J., Lu, Y., Wu, H.: Cross-Domain Collaborative Federated Intelligence for Wireless Computing Power Networks. Journal of Intelligent Computing and Networking 2(1), 22-34 (2026). https://doi.org/10.64509/jicn.21.75 DOI: https://doi.org/10.64509/jicn.21.75

[23] Chen, W., Deng, D., Xiang, C., Liu, Z., Xiao, B.: When One-Shot Federated Learning Meets Diffusion Models at the Edge: Technological Advances and Applications. Journal of Intelligent Computing and Networking 2(1), 35-54 (2026). https://doi.org/10.64509/jicn.21.64 DOI: https://doi.org/10.64509/jicn.21.64

[24] Lameter, C.: NUMA (Non-Uniform Memory Access): An Overview. Queue 11(7), 40-51 (2013). https://doi.org/10.1145/2508834.2513149 DOI: https://doi.org/10.1145/2508834.2513149

[25] Springer, R., Lowenthal, D.K., Rountree, B., Freeh, V.W.: Minimizing Execution Time in MPI Programs on an Energy-Constrained, Power-Scalable Cluster. In the 2006 Eleventh ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 2006), pp. 230-238 (2006). https://doi.org/10.1145/1122971.1123006 DOI: https://doi.org/10.1145/1122971.1123006

[26] Kumar, M., Kaur, G.: Containerized MPI Application on InfiniBand based HPC: An Empirical Study. In 2022 3rd International Conference for Emerging Technology (INCET), pp. 1-6 (2022). https://doi.org/10.1109/INCET54531.2022.9824366 DOI: https://doi.org/10.1109/INCET54531.2022.9824366

[27] Souza Filho, P., Bulcão, A., Panetta, J., Lough, M., Monnet, B.: Efficient Execution of MPI Containers. In the 2023 EAGE Seventh High Performance Computing Workshop (HPCW 2023), vol. 2023, pp. 1-5 (2023). https://doi.org/10.3997/2214-4609.2023630030 DOI: https://doi.org/10.3997/2214-4609.2023630030

[28] NASA Advanced Supercomputing Division: NAS Parallel Benchmarks. https://www.nas.nasa.gov/software/npb.html Accessed 2026-07-22

[29] Intel: Intel Performance Counter Monitor (Intel PCM). https://github.com/intel/pcm Accessed 2026-07-22

JICN137

Downloads

Published

2026-08-06

Issue

Section

Articles

How to Cite

Wang, W., Gao, Z., Ji, G., Yong, C., & Li, B. (2026). Computing Resource-Aware Operation Optimization Strategy for MPI Jobs in Cloud-Native Environment. Journal of Intelligent Computing and Networking, 2(3), 14-29. https://doi.org/10.64509/jicn.23.137

Similar Articles

1-10 of 18

You may also start an advanced similarity search for this article.