You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
<li><b>Add the module to <spanclass="tt">docs/module_categories.json</span></b> so it appears in this page</li>
430
430
</ol>
431
431
<p>Follow the pattern of existing modules like <spanclass="tt">m_body_forces</span> (simple) or <spanclass="tt">m_viscous</span> (more involved) as a template.</p>
<p>💡 <b>Tip:</b> If you encounter a validation error, check the relevant section above or review <ahref="https://github.com/MFlowCode/MFC/blob/master/toolchain/mfc/case_validator.py"><spanclass="tt">case_validator.py</span></a> for complete validation logic.</p>
<p>This page covers how to achieve maximum performance with MFC, including optimization techniques and benchmark results across various hardware platforms.</p>
<p>The single most impactful optimization is <b>case optimization</b>, which can provide <b>up to 10x speedup</b> for both CPU and GPU runs.</p>
169
169
<p>Case optimization works by hard-coding your simulation parameters at compile time, enabling aggressive compiler optimizations (loop unrolling, constant propagation, dead code elimination).</p>
<p>The following table outlines observed performance as nanoseconds per grid point (ns/gp) per equation (eq) per right-hand side (rhs) evaluation (lower is better), also known as the grind time. We solve an example 3D, inviscid, 5-equation model problem with two advected species (8 PDEs) and 8M grid points (158-cubed uniform grid). The numerics are WENO5 finite volume reconstruction and HLLC approximate Riemann solver. This case is located in <spanclass="tt">examples/3D_performance_test</span>. You can run it via <spanclass="tt">./mfc.sh run -n <num_processors> -j $(nproc) ./examples/3D_performance_test/case.py -t pre_process simulation --case-optimization</span> for CPU cases right after building MFC, which will build an optimized version of the code for this case then execute it. For benchmarking GPU devices, you will likely want to use <spanclass="tt">-n <num_gpus></span> where <spanclass="tt"><num_gpus></span> should likely be <spanclass="tt">1</span>. If the above does not work on your machine, see the rest of this documentation for other ways to use the <spanclass="tt">./mfc.sh run</span> command.</p>
233
233
<p>Results are for MFC v4.9.3 (July 2024 release), though numbers have not changed meaningfully since then. Similar performance is also seen for other problem configurations, such as the Euler equations (4 PDEs). All results are for the compiler that gave the best performance. Note:</p><ul>
<p><b>All grind times are in nanoseconds (ns) per grid point (gp) per equation (eq) per right-hand side (rhs) evaluation, so X ns/gp/eq/rhs. Lower is better.</b></p>
<p>Strong scaling results are obtained by keeping the problem size constant and increasing the number of processes so that work per process decreases.</p>
<p>The base case utilizes 8 GPUs with one MPI process per GPU for these tests. The performance is analyzed at two problem sizes: 16M and 64M grid points. The "base case" uses 2M and 8M grid points per process.</p>
<p>Instead of building MFC from scratch, you can use containers to quickly access a pre-built version of MFC and its dependencies. In brief, you can run the latest MFC container: </p><divclass="fragment"><divclass="line">docker run -it --rm --entrypoint bash sbryngelson/mfc:latest-cpu</div>
287
287
</div><!-- fragment --><p> Please refer to the <aclass="el" href="docker.html" title="Containers">Docker</a> document for more information.</p>
<p>MFC has example cases in the <spanclass="tt">examples</span> folder. You can run such a case interactively using 2 tasks by typing:</p>
295
295
<divclass="fragment"><divclass="line">./mfc.sh run examples/2D_shockbubble/case.py -n 2</div>
296
296
</div><!-- fragment --><p>Please refer to the <aclass="el" href="running.html" title="Running">Running</a> document for more information on <spanclass="tt">case.py</span> files and how to run them.</p>
<p>MFC is <b>unit-agnostic</b>: the solver performs no internal unit conversions. Whatever units you provide for initial conditions, boundary conditions, and material properties, the same units appear in the output.</p>
300
300
<p>The only requirement is <b>consistency</b> — all inputs must use the same unit system. Note that some parameters use <b>transformed stored forms</b> rather than standard physical values (e.g., <spanclass="tt">gamma</span> expects \(1/(\gamma-1)\), not \(\gamma\) itself). See <aclass="el" href="equations.html#sec-stored-forms" title="Stored Parameter Conventions">Stored Parameter Conventions</a> for details.</p>
<divclass="line">./mfc.sh viz examples/2D_shockbubble/ --var pres --step all --mp4</div>
312
312
</div><!-- fragment --><p>Output images and videos are saved to the <spanclass="tt">viz/</span> subdirectory of the case. For more options, see <aclass="el" href="visualization.html" title="Flow visualization">Flow Visualization</a> or run <spanclass="tt">./mfc.sh viz -h</span>.</p>
0 commit comments