二维水流模型的多GPU并行计算研究

Multi-GPU parallel computation for a two-dimensional flow model

  • 摘要: 针对二维水流模型计算效率慢的问题,本文采用多GPU并行计算技术提升模型运算性能。基于CUDA并行计算平台,构建融合区域分解与交界连通的并行策略,将整体计算域剖分为若干计算子区域,保证各子区域网格数量与计算负载均衡,由独立GPU负责单个子区域的水动力求解;在相邻子区域的上下游交界位置建立跨子域数据交换机制,在每一时间步完成水位、流量等水力信息的同步传递,保证了全计算域流场的连续一致性。选取长江干流澄通段作为工程算例开展二维水动力模拟,设置2万、11万、43万、113万共4种不同规模的计算网格,分析单GPU与多GPU方案的并行加速性能。以CPU串行程序耗时作为基准,在113万网格规模下,单GPU并行最大加速比可达约191倍,4块GPU协同并行的最大加速比约310倍。研究表明,本文提出的区域分解与交界连通并行策略,能够稳定实现二维水流模型的多GPU并行计算,显著缩短模型迭代耗时,大幅提升计算效率。该并行方案在流域大尺度水动力模拟、河道洪水演进预报以及精细化水情分析等场景将具备良好的工程应用价值。

     

    Abstract: This study investigates multi-GPU parallel computation for a two-dimensional water flow model aims to address the limited computational efficiency of conventional serial explicit schemes for solving two-dimensional shallow water equations based on the finite-volume method. The governing two-dimensional shallow water equations are discretized by a finite-volume explicit scheme, which calculates numerical fluxes at cell interfaces and updates hydraulic variables subject to the CFL stability criterion. Traditional serial numerical simulation suffers severe computational bottlenecks for high-resolution and large-scale hydrodynamic grids. The explicit scheme requires iterative calculations for every grid cell at each time step, which substantially increases computation time and limits the application of the model in refined hydrodynamic simulation and long-term water evolution prediction. To overcome the efficiency limit of serial computation, multi-GPU parallel computing is adopted in this study to improve the overall computational performance of the hydrodynamic model. The heterogeneous computing platform uses one Intel Xeon Gold 6150 CPU (2.7 GHz base frequency) and four NVIDIA Tesla P100 GPU accelerators. Based on NVIDIA’s CUDA parallel computing platform, this study proposes a parallel framework combining domain decomposition and inter-subdomain boundary communication. The full computational domain of the two-dimensional flow model is partitioned into several independent subdomains with balanced grid numbers and computational workloads. This strategy achieves load balancing across computing devices and prevents idle computing resources caused by uneven grid distribution. Each subdomain is assigned to an individual GPU for independent computation. Ghost-cell halo regions are reserved at the boundaries of neighboring subdomains, and non-blocking data exchange protocols are defined for these interfaces. At the end of each time step, hydraulic variables including water level and discharge are exchanged across the halo regions to maintain flow continuity and the physical consistency of the global flow field, enabling high-precision simulation of flow evolution across the entire domain. Within each time step, GPU kernels compute numerical fluxes and update flow states for interior cells, while asynchronous CUDA memory copies transfer boundary data between the host and device. CUDA multi-stream technology overlaps computation and communication to reduce the latency of CPU-GPU data transmission. The two-dimensional hydrodynamic simulation of the Chengtong reach of the Yangtze River is used as the engineering validation case. Field hydrological observations, including time-series water level and discharge data at typical cross-sections, are used to verify the reliability and numerical stability of the multi-GPU parallel model. Good agreement between simulated and measured hydrographs confirms that the parallel implementation does not introduce extra numerical dissipation and preserves the mass conservation property of the original finite-volume solver. Parallel acceleration performance is quantitatively evaluated for four mesh sizes (20 000, 110 000, 430 000 and 1 130 000 cells) to explore the efficiency characteristics of single-GPU and multi-GPU parallel computing. The wall-clock runtime of the original CPU serial solver is taken as the benchmark for acceleration ratio calculation. Table 1 lists model computation time, the fraction of GPU computation and CPU-GPU communication in total runtime, as well as parallel acceleration ratios under these four grid resolutions. Experimental results show that for the finest mesh with 1.13 million cells, single-GPU parallel computation reaches an acceleration ratio of approximately 191 times compared with serial computation, and four-GPU collaborative parallel computation further lifts the acceleration ratio to around 310 times. Statistical results reveal that pure GPU computation occupies 51%–63% of total parallel runtime, while CPU-GPU data communication accounts for the remaining 37%–49%. As grid size increases, the relative proportion of communication overhead gradually decreases and the acceleration ratio keeps rising. This indicates that larger computational domains can better leverage GPU arithmetic throughput and amortize fixed communication costs. The results demonstrate that the proposed parallel strategy based on domain decomposition and boundary connectivity can realize efficient multi-GPU parallel computation for two-dimensional water flow models. This parallel optimization reduces iteration time substantially while maintaining good numerical accuracy and simulation stability. It provides efficient and reliable technical support for large-scale basin hydrodynamic simulation, high-precision river flood evolution forecasting and refined water regime analysis, and has broad application prospects and important engineering value in hydrological numerical simulation.

     

/

返回文章
返回