Image Details

Choose export citation format:

STORM: RDMA-based Monte Carlo Transport Scheme for Distributed-memory Particle Simulations

  • Authors: Maor Mizrachi, Barak Raveh, Elad Steinberg

Maor Mizrachi et al 2026 The Astrophysical Journal Supplement Series 286 .

  • Provider: AAS Journals

Caption: Figure 10.

Average per-rank wall-clock time breakdown for a representative Monte Carlo step of the cylindrical Hohlraum IMC benchmark on 40 nodes (4480 ranks), comparing the OFI (RDMA) and P2P backends. All values are averages across ranks. Computation: Physics—the physics->step() call (face-intersection geometry, opacity lookup, scattering, and energy deposition). Communication: MPI Progress—per-iteration overhead of MPI progress calls, cell-move bookkeeping, particle-list management, and send-buffer queuing within the inner particle loop; Send/Recv—send and receive calls, buffer management, and flushing of aggregated particle buffers to the network; Termination—probing and advancing the tree-based distributed termination counter; Realloc.—asynchronous buffer reallocation progress (one-sided backends only). Busy waiting: accumulated per-iteration overhead of timing instrumentation and loop control across millions of main-loop iterations. Both backends perform the same 9.31 × 109 particle steps; the physics time is identical. The 1.41× speedup originates from reduced MPI progress overhead and lower communication costs.

Other Images in This Article
Copyright and Terms & Conditions

Additional terms of reuse