Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1013 commits
Select commit Hold shift + click to select a range
4799a4c
Ledger (69): goal v2 -- gated increments are pushed the session they …
sbryngelson Sep 5, 2026
13a18f0
Merge branch 'task10/batchcount' into up/mega
sbryngelson Sep 5, 2026
704582d
Walk the rebuild box loop over this rank's participants, not every bo…
sbryngelson Sep 4, 2026
b188e75
Walk the rebuild's old-block loops over the stashes this rank holds
sbryngelson Sep 4, 2026
560f21c
Check seam topology from this rank's owned blocks, not all pairs
sbryngelson Sep 4, 2026
e2fc388
Print per-rank seconds for the regrid sub-phases in the phase-rank table
sbryngelson Sep 4, 2026
1a4344d
Drop the participant role array: the consumers' own predicates alread…
sbryngelson Sep 4, 2026
83e484a
Ledger (70): load balance replayed offline at zero cost -- do not imp…
sbryngelson Sep 5, 2026
082f65f
Merge branch 'task9/rebuild' into up/mega
sbryngelson Sep 5, 2026
5f2e184
Ledger (71): Task 9 merged -- the regrid rebuild's O(P) rows fall fro…
sbryngelson Sep 5, 2026
3ed5bae
Add the amr_device_pack case flag (default F): requires amr, excludes…
sbryngelson Sep 5, 2026
2f650c3
F1/F2 coarse-patch gather: one fused pack/unpack kernel per family pe…
sbryngelson Sep 5, 2026
f64fea7
Ledger (72): the controlled ladder and the MI250X A/B, reviewed -- ba…
sbryngelson Sep 5, 2026
5e0f5ea
Ledger (73): the steady AMR excess is 1.44 s/step (2.1x target) with …
sbryngelson Sep 5, 2026
80225d4
Ledger (74): the step is only ~26% MPI wait (a lower bound), and the …
sbryngelson Sep 5, 2026
d238236
Merge branch 'up/mega' into task10/fusedpack
sbryngelson Sep 5, 2026
b126ee8
Ledger (75, 76): the fused gather packs are worth 0.14 s/step with 12…
sbryngelson Sep 5, 2026
23800cf
AMR: correct the amr_device_pack description (sends fuse per family, …
sbryngelson Sep 5, 2026
a11b4fe
Merge task10/fusedpack: fused F1/F2 gather packs behind the default-o…
sbryngelson Sep 5, 2026
3e208d3
Ledger (75): record the gates the fused-pack merge passed, and the 3 …
sbryngelson Sep 5, 2026
2d3381c
Ledger (76): the 8-GPU/node first doubling is 1.30x -- the 21 percent…
sbryngelson Sep 5, 2026
a4618f6
Ledger (77): my own hypothesis falsified -- the allocator setting rec…
sbryngelson Sep 5, 2026
e36a680
Ledger (75): all 70 AMR goldens pass on the merged tree -- the 3 chem…
sbryngelson Sep 5, 2026
4b534eb
Ledger (78): this node's intra-node MPI wait degraded 4.4x during the…
sbryngelson Sep 5, 2026
cede444
Ledger (78): name the canary script and record that its first design …
sbryngelson Sep 5, 2026
7babe17
AMR: delete amr_rg_gather and its 35 unreachable sites -- a flag noth…
sbryngelson Sep 5, 2026
97eedbb
Ledger (80): per-block cost is ~13 ms/block/step across six phases, p…
sbryngelson Sep 5, 2026
04fb2f0
Ledger (81): negative, pre-registered, falsifier fired -- pooling the…
sbryngelson Sep 5, 2026
9002f3c
AMR instrument: five bracket-free host-time rows (h:slot/shell/own/un…
sbryngelson Sep 5, 2026
8dc669a
AMD OpenMP lane: per-file opt-in defaultmap(present:allocatable) (MFC…
sbryngelson Sep 5, 2026
00caa29
Ledger (82): the per-block AMR cost is amdflang's per-launch re-map o…
sbryngelson Sep 5, 2026
256355c
Ledger (82) correction: unallocated module arrays abort under present…
sbryngelson Sep 5, 2026
a235b5a
AMD OpenMP lane: m_amr_registers.fpp opts in to defaultmap(present:al…
sbryngelson Sep 5, 2026
22b4fba
m_amr_registers opt-in header: fypp comments only (the formatter had …
sbryngelson Sep 5, 2026
11f4a77
Ledger (83): m_amr_registers opted in -- at most 2% at cap 32, nothin…
sbryngelson Sep 5, 2026
b5b1782
AMR: one device-resident slab table + one GPU_UPDATE per launch for t…
sbryngelson Sep 5, 2026
daaa80c
Ledger (84): one device slab table + one update per launch, pre-regis…
sbryngelson Sep 5, 2026
55c735d
AMR: amr_batched_gather (default F) -- pool the gathered coarse patch…
sbryngelson Sep 5, 2026
5ee8e1d
AMR: amr_batched_gather -- wire the F2 parent-fill wave to the pooled…
sbryngelson Sep 5, 2026
f920995
AMR: amr_batched_gather -- keep the per-member tables host-only so co…
sbryngelson Sep 5, 2026
f1510c7
AMR: amr_batched_gather -- the pooled unpack copies in only this wave…
sbryngelson Sep 5, 2026
ab091b8
amr_batched_gather rebase: restore the four preprocessor directives t…
sbryngelson Sep 6, 2026
6ddd8f1
Ledger (85): ledger 81 re-tested under the clause -- the pooled gathe…
sbryngelson Sep 6, 2026
f04a2e4
Ledger (86): scorecard item 2 re-measured on the pushed tip, two code…
sbryngelson Sep 6, 2026
9395593
Ledger (87): negative, pre-registered -- a per-block cost weight at K…
sbryngelson Sep 6, 2026
b3f12b0
Ledger (88): the per-block cost driver named from 66k batch records -…
sbryngelson Sep 6, 2026
20cdd96
AMR instrument: per-batch swap/rhs/restore/rk timing with member ids …
sbryngelson Sep 6, 2026
5e816d6
AMR batch instrument: open the per-rank log once (a logical flag; new…
sbryngelson Sep 6, 2026
36071e2
AMR batched advance: amr_bat_pad (default 0) lets a smaller block joi…
sbryngelson Sep 6, 2026
f29f8d3
amr_bat_pad: the batched capture reads member extents from amr_bat_me…
sbryngelson Sep 6, 2026
cea4202
amr_bat_pad: declare amr_bat_mext in m_global_parameters beside the o…
sbryngelson Sep 6, 2026
6bfa859
Ledger (89): the padded batch A/B -- batches -58%, summed rhs -16%, s…
sbryngelson Sep 6, 2026
04d1c6f
Ledger (90): amr_bat_pad at cap 96 (negative: wait up on every rank, …
sbryngelson Sep 6, 2026
6f6febb
Ledger (91): amr_lb_block_cost K=2 on top of padded batching -- null …
sbryngelson Sep 6, 2026
516399a
Ledger (92): the per-batch fixed cost named from a kernel+copy trace …
sbryngelson Sep 6, 2026
8c81242
amdflang: opt m_rhs and m_weno into defaultmap(present:allocatable) (…
sbryngelson Sep 6, 2026
09f7a17
Ledger (93): m_rhs and m_weno present:allocatable opt-in -- -0.09/-0.…
sbryngelson Sep 6, 2026
4e44eb8
AMR instruments: f_amr_wtime() wraps MPI_Wtime under MFC_MPI (the ser…
sbryngelson Sep 6, 2026
fa972ef
post_process: the AMR overlay no longer stores block-local mixture fi…
sbryngelson Sep 7, 2026
a10128b
test harness: post-process cases keep parallel_io = F on a no-MPI bui…
sbryngelson Sep 7, 2026
ea54255
Grid minima over the filled range: dx/dy/dz are allocated to the _all…
sbryngelson Sep 6, 2026
50b4e47
Ledger (94): the PR's CI read in full -- every Frontier lane's heap c…
sbryngelson Sep 7, 2026
b0b2919
amdflang: opt m_riemann_solver_hllc into defaultmap(present:allocatab…
sbryngelson Sep 6, 2026
5c68785
Ledger (95): m_riemann_solver_hllc present:allocatable opt-in behind …
sbryngelson Sep 7, 2026
d02ca91
Serial I/O: write bc_type.dat and bc_buffers.dat into every simulatio…
sbryngelson Sep 7, 2026
05e0a9d
post_process: the save-index gap skip under cfl_dt now probes the ser…
sbryngelson Sep 7, 2026
45f5312
post_process: write the rectilinear-grid coordinates with the Silo da…
sbryngelson Sep 7, 2026
5717670
post_process: the Lagrangian-bubble and immersed-body point meshes an…
sbryngelson Sep 7, 2026
e934894
Ledger (96): the PR's CI closed out -- the three remaining failure cl…
sbryngelson Sep 7, 2026
5f3ccfa
Merge upstream/master into up/mega: the Phoenix single-job benchmark …
sbryngelson Sep 7, 2026
8644c8b
AMR batched stage: skip the restore-side device push of the grid stat…
sbryngelson Sep 7, 2026
4c519b8
Ledger (97): restore-side grid-state device push skipped between cons…
sbryngelson Sep 8, 2026
67b5480
Ledger (98): two pre-registered negatives close the per-launch copy c…
sbryngelson Sep 8, 2026
b52479a
Ledger (99): cap 96 explained and re-measured -- amr_bat_pad is a sma…
sbryngelson Sep 8, 2026
f5f5152
Ledger (99) correction: the validator already prohibits batched advan…
sbryngelson Sep 8, 2026
e8cecbd
AMR: the batched fine advance turns on by default where the case admi…
sbryngelson Sep 8, 2026
fbd72e6
AMR batching default: amr_device_pack does not ride along (its cap-32…
sbryngelson Sep 8, 2026
79c108f
Ledger (100): the batched fine advance on by default where the case a…
sbryngelson Sep 8, 2026
16751fd
test harness: 27 AMR goldens get amr_max_grid_size pinned at the valu…
sbryngelson Sep 8, 2026
84dbdd0
Ledger (101): 27 AMR goldens pin amr_max_grid_size at the derived val…
sbryngelson Sep 8, 2026
a984dab
AMR batched advance: apply the static-body immersed-boundary correcti…
sbryngelson Sep 8, 2026
efced0f
AMR batched IB correction: hold amr_bat_n at 1 while the members are …
sbryngelson Sep 8, 2026
90defe8
Ledger (103): the batched advance skipped the post-RK hooks -- the st…
sbryngelson Sep 8, 2026
43cfccc
AMR fine RHS: zero body cells by the block's OWN fine markers (ib_mar…
sbryngelson Sep 8, 2026
f849a13
AMR golden: static IBM circle -> dynamic regrid -> batched pair (two …
sbryngelson Sep 8, 2026
95d5f08
AMR IB goldens regenerated for the fine-marker RHS zeroing (6 cases: …
sbryngelson Sep 8, 2026
04c5d82
Ledger (104): the fine RHS zeroed body cells by the coarse marker pat…
sbryngelson Sep 8, 2026
99e62cf
Ledger (102): item 4 scoped and measured -- the np=8 exchange is alre…
sbryngelson Sep 8, 2026
d397293
Toolchain default: amr_device_pack rides with the batching default wh…
sbryngelson Sep 8, 2026
2dc132c
Ledger (106): amr_device_pack A/B at caps 32 and 96 closes GOAL v3 it…
sbryngelson Sep 8, 2026
8a059f9
AMR regrid: the global box union is no longer truncated to amr_max_bl…
sbryngelson Sep 8, 2026
5986d9b
Ledger (105): the 2-node rung found a correctness cliff -- the global…
sbryngelson Sep 8, 2026
18e16b7
Ledger (107): scorecard item 2 re-measured on the shipped defaults, t…
sbryngelson Sep 8, 2026
bbf39c8
Ledger 107 same-session correction: the 'unbracketed sixth' was an ac…
sbryngelson Sep 8, 2026
b9b0039
AMR migration: the wire buffers are device-resident and, under rdma_m…
sbryngelson Sep 8, 2026
796f2fa
Ledger (108): migration off the host (GOAL v4 item 1) -- the regrid's…
sbryngelson Sep 8, 2026
62bd10e
Ledger (110): halo width is not where MFC's AMR excess sits -- the S0…
sbryngelson Sep 8, 2026
7e0d3ac
AMR regrid hysteresis (amr_snap, default 0 = off): a new box within a…
sbryngelson Sep 8, 2026
343470c
Ledger (109): regrid hysteresis (amr_snap, default off) -- a new box …
sbryngelson Sep 8, 2026
74bec70
Toolchain default: amr_snap = min(2, amr_buf - 2) rides with the batc…
sbryngelson Sep 8, 2026
2c23b6c
Ledger (111): scorecard item 2 with the regrid hysteresis on -- two c…
sbryngelson Sep 8, 2026
601e25a
Merge upstream master d2d8cac2 into up/mega: state-dependent equation…
sbryngelson Sep 8, 2026
4ac310f
Ledger (112): upstream master d2d8cac2 merged (state-dependent equati…
sbryngelson Sep 9, 2026
a79369e
Ledger (113): item 4's first read closed -- the two-node doubling cos…
sbryngelson Sep 9, 2026
343d57c
Ledger (115): the rep-to-rep climb is not node state and not intrinsi…
sbryngelson Sep 9, 2026
bde8fb1
Ledger (114): the two-node rung with the regrid hysteresis on -- 1.26…
sbryngelson Sep 9, 2026
08ae3e2
Ledger (117): the clean 2x statement -- two codes, one node, one wind…
sbryngelson Sep 9, 2026
2d84e7c
Ledger (118): the rung with all shipped defaults -- np8 5.14 -> 3.13 …
sbryngelson Sep 9, 2026
9cb243a
Docs: mark every 4.96 % noise-floor citation as superseded by the mea…
sbryngelson Sep 9, 2026
6c9c2d5
Ledger (120): the floor of the differenced protocol -- five back-to-b…
sbryngelson Sep 9, 2026
4f71ca4
AMR migration: bound the device-resident wire pools (2 GiB); above th…
sbryngelson Sep 9, 2026
503244c
Ledger (125) + AMR migration: bound the device-resident wire pools at…
sbryngelson Sep 9, 2026
7e958f1
AMR: hoist the coarse cons halo before the coarse RHS and convert ove…
sbryngelson Sep 9, 2026
c529bc5
Ledger (122) + AMR: the coarse cons halo runs before the coarse RHS a…
sbryngelson Sep 9, 2026
f20dbeb
AMR fold: retire the last comments naming the deleted freg wave
sbryngelson Sep 9, 2026
f4d8b7a
Ledger (123) + AMR fold: the level>=2 freg faces ride the restrict-pa…
sbryngelson Sep 9, 2026
3e79085
AMR rebuild: walk the boxes owner-interleaved per level (round-robin …
sbryngelson Sep 9, 2026
b4ae46b
Ledger (126) + AMR rebuild: walk the boxes owner-interleaved per leve…
sbryngelson Sep 9, 2026
85bb4a1
AMR seam wave: post at the top of the stage (sends read stage-entry i…
sbryngelson Sep 9, 2026
0e6805e
Ledger (129) + AMR seam wave: post at the top of the stage with priva…
sbryngelson Sep 9, 2026
1ddedac
Ledger (119): the np16 rebuild-free window probe -- the two-node wait…
sbryngelson Sep 9, 2026
7cb0e3f
Ledger 127: ownership stickiness at rebuild does not reduce migration…
sbryngelson Sep 9, 2026
a35a8ad
Ledger 130: clean-node np8 reads of the rendezvous cuts and the rebui…
sbryngelson Sep 9, 2026
f93fd33
Ledger 132: scorecard item 2 re-baselined on the fixed pin, one node,…
sbryngelson Sep 10, 2026
40173b2
Ledger 133: five np16 rungs on a healthy node pair (doublings 1.23-1.…
sbryngelson Sep 10, 2026
9607583
Ledger 134: the np32/np48 NaN of ledger 133 was a wrong pre_process b…
sbryngelson Sep 11, 2026
e93c7ce
Ledger 135: the largest-first batch order halves the fine-advance ran…
sbryngelson Sep 11, 2026
5146bc6
EOS: bake the state-dependent-EOS flag at build time and keep the HLL…
sbryngelson Sep 10, 2026
2cad101
Ledger 131: the master merge cost the fine RHS +37 % through the stat…
sbryngelson Sep 11, 2026
d94b349
Ledger 136: one profiler for both codes -- MFC spends a smaller share…
sbryngelson Sep 11, 2026
c6f41b7
amr: split the reflux-faces wave into post and drain and run the L0 c…
sbryngelson Sep 11, 2026
9cc07b0
Ledger 137: post the reflux-faces wave before the L0 coarse RHS and d…
sbryngelson Sep 11, 2026
11f63a0
Revert the reflux-wave post/drain split: against its true parent it i…
sbryngelson Sep 11, 2026
6169032
Ledger 138: retract ledger 137 -- its control pin and its uniform ter…
sbryngelson Sep 11, 2026
9345b39
Mark ledger 137 retracted in its own header (ledger 138 supersedes it)
sbryngelson Sep 11, 2026
875d2cb
docs: table the per-step rendezvous count before and after the GOAL v…
sbryngelson Sep 11, 2026
2ec578b
Ledger 139: the first pin-matched two-code reads of the current code …
sbryngelson Sep 11, 2026
39e0225
Ledger 139: apply the six reviewer corrections lost from the landed t…
sbryngelson Sep 11, 2026
c0fbe7b
Ledger 140: growing both uniform arms to the AMR arm's 200-step windo…
sbryngelson Sep 11, 2026
4b8ec25
Ledger 141: GOAL v8 item 2 is not built -- a nowait target region enc…
sbryngelson Sep 11, 2026
4ad4ca6
Ledger 128: the weak-scaling curve on the rung deck, np8 to np32 -- 1…
sbryngelson Sep 12, 2026
b2be495
Correct the restr phase-split comment: rs:wave is neither deleted nor…
sbryngelson Sep 12, 2026
cb8ded1
Move only the six face planes the level>=2 reflux apply touches, not …
sbryngelson Sep 12, 2026
ade997c
Ledger 151: statement 2 re-read case-optimized is 0.627 and 1.71x AMR…
sbryngelson Sep 12, 2026
e6ea130
Ledger 152: 56 percent of statement 2's excess is MPI wait with a 251…
sbryngelson Sep 12, 2026
9f7c41e
Merge upstream master into up/mega: take upstream's continuum-damage …
sbryngelson Sep 13, 2026
3d82e94
Merge upstream master again: adopt the AMD_NUM_SPECIES_MAX species bo…
sbryngelson Sep 13, 2026
36e80f2
Give eos_state_dependent a Fypp default in the CMake build (fixes doc…
sbryngelson Sep 13, 2026
9121920
Ledger 150: five code-review leads taken to verdicts, with three meas…
sbryngelson Sep 13, 2026
25a3c31
Document that a generated-case Fypp variable also needs a CMake default
sbryngelson Sep 13, 2026
ee642c5
Quote file:line in the plan doc so Doxygen stops autolinking it (fixe…
sbryngelson Sep 13, 2026
fdd49c5
Regenerate the AMR cont_damage golden for upstream's undamaged-modulu…
sbryngelson Sep 13, 2026
aa6a589
Size HLLC star states by AMD_SYS_SIZE_MAX, the half of #1852 the merg…
sbryngelson Sep 13, 2026
001a945
Ledger 153: the master merge changed exactly one test, the AMR cont_d…
sbryngelson Sep 13, 2026
13788bc
Reword file:line in the plan doc; Doxygen autolinks it even inside a …
sbryngelson Sep 13, 2026
a544996
Measure phase nesting at run time so the budget residual sums only to…
sbryngelson Sep 13, 2026
4e7da65
Ledger 154: the step budget measures its own nesting, and the measure…
sbryngelson Sep 13, 2026
b994bc0
Ledger 153: the full suite confirms the merge changed only the one re…
sbryngelson Sep 13, 2026
c1eea07
Ledger 155: Phase 2 priced before it was built; the per-rank skew is …
sbryngelson Sep 13, 2026
561a9da
Ledger 156: the amr_bat_pad probe; merging shapes saves ~3.3 ms per l…
sbryngelson Sep 13, 2026
f4aea3a
Ledger 153: record the merge's conflict resolutions
sbryngelson Sep 13, 2026
639325c
Add the amr_equal_tiles parameter: validator rules, docs, default off
sbryngelson Sep 13, 2026
fffab63
Add s_amr_equal_tile_extent: equal-tile extent within 2-cell-floored …
sbryngelson Sep 13, 2026
25dab1e
Add amr_equal_tiles call sites, [amr-tile] report, and a multi-level …
sbryngelson Sep 13, 2026
888e31b
Golden for the amr_equal_tiles sibling of the pinned-cap multi-level …
sbryngelson Sep 13, 2026
9705ab0
Name the [amr-tile] counters for what they count (all_dims, some_dims)
sbryngelson Sep 13, 2026
cfbec7f
Ledger 157: equal tiles stops at its eligibility gate; no level-2 box…
sbryngelson Sep 13, 2026
c9a67f7
Ledger 158: the wait floor is imbalance, not transfer; statement 2 re…
sbryngelson Sep 13, 2026
d1b379a
Ledger 158: replace the withdrawn rendezvous design with the priced t…
sbryngelson Sep 13, 2026
2141f7f
Add amr_lb_beta: time-feedback block weights in the regrid partition …
sbryngelson Sep 13, 2026
8afa115
Golden for the amr_lb_beta sibling of the pinned-cap multi-level deck
sbryngelson Sep 13, 2026
e9e4c0b
Ledger 159: time-feedback block weights equalise time by making it la…
sbryngelson Sep 13, 2026
25663cd
WENO5: reconstruct one cell in a device routine so its six work array…
sbryngelson Sep 14, 2026
1c7e1df
Move the HLLC per-face body into device routines (137 -> 23 kernel ar…
sbryngelson Sep 14, 2026
7926d48
GPU_INLINE_CALL: force-inline the per-cell device routines on amdflan…
sbryngelson Sep 14, 2026
b68af5a
Move the conservative-to-primitive kernel body into a device routine …
sbryngelson Sep 14, 2026
453ba96
HLLC face routine takes explicit-shape arrays; WENO cell results go t…
sbryngelson Sep 14, 2026
205ec78
Kernel work arrays go BLOCK-local instead of private: hllc 46.9 -> 5.…
sbryngelson Sep 14, 2026
4315bc8
WENO5: the block with the work arrays wraps the seq i-loop, not each …
sbryngelson Sep 14, 2026
fc09a52
Merge upstream master (aa4e4585) into up/mega
sbryngelson Sep 14, 2026
88819ac
Ledger 160: the descriptor tax -- BLOCK-local work arrays take 42 of …
sbryngelson Sep 14, 2026
6b5c734
Ledger 161: the rendezvous timeline -- ~14 syncs/step, two carry the …
sbryngelson Sep 14, 2026
76eb142
AMR fine-window wire pools live on the device under rdma_mpi; MPI sen…
sbryngelson Sep 14, 2026
74fffb9
Ledger 162: device-resident wire pools -- wall -73 ms/step (5/5 predi…
sbryngelson Sep 14, 2026
f357bb0
Ledger 163: batched-advance launches are no longer the wall (leaders …
sbryngelson Sep 14, 2026
cb428ca
Phase report: every phase per rank
sbryngelson Sep 14, 2026
8f2e4a5
Fine-block reconstruction: clip the transverse window to the block in…
sbryngelson Sep 14, 2026
651f752
Ledger 164: fine-block reconstruction window clipped to the interior …
sbryngelson Sep 15, 2026
5bceeda
Put the MFC_GPU preprocessor directives in s_convert_conservative_to_…
sbryngelson Sep 15, 2026
121585d
Ledger 165: statement 2 re-read node-matched A-B-A-A on k004-005 -- 1…
sbryngelson Sep 15, 2026
952f396
Delete the AMR instruments the campaign rejected: amr_equal_tiles, am…
sbryngelson Sep 15, 2026
917cc39
amr_status: record the tier-1 cleanup and what remains a product deci…
sbryngelson Sep 15, 2026
634cd5f
Clean up the AMR record: campaign notebooks move to misc/amr_ledger w…
sbryngelson Sep 15, 2026
d8f5509
case.md: cross-page anchors as Doxygen refs (fixes the docs link check)
sbryngelson Sep 15, 2026
d298d05
Kernel work arrays back to private= (drop the BLOCK-in-kernel form: N…
sbryngelson Sep 15, 2026
fc05bed
Migration pack/unpack kernels: drop the present= clause (map(present,…
sbryngelson Sep 15, 2026
eaf2aee
Batched fine advance admits MHD, relativity, hypoelasticity, continuu…
sbryngelson Sep 15, 2026
90b81be
Batched fine advance admits chemistry and IGR: seed the whole tempera…
sbryngelson Sep 15, 2026
021b983
Batched fine advance no longer needs a pinned amr_max_grid_size: the …
sbryngelson Sep 15, 2026
26c527e
Lock-step goldens for the 6-equation and moving-body AMR cases, gener…
sbryngelson Sep 15, 2026
56a5bbf
Batched fine advance admits the 6-equation model and moving bodies: t…
sbryngelson Sep 15, 2026
c20da8b
Remove the unused locals, one orphan module variable, one dead functi…
sbryngelson Sep 15, 2026
2a7a88d
Drop the campaign's scaling counters: gather-byte, clustering-tree an…
sbryngelson Sep 15, 2026
ff77999
Split m_amr into a chain of eight service modules plus the driver (pu…
sbryngelson Sep 16, 2026
adf5d06
Source lint rejects the per-executable preprocessor guards (MFC_PRE_P…
sbryngelson Sep 16, 2026
1831336
Non-Newtonian IBM golden: absolute tolerance 1e-8 (its zero field rea…
sbryngelson Sep 16, 2026
88a9857
Ledger 166: subcycling priced against the batched default (1.20x slow…
sbryngelson Sep 16, 2026
ac58fd7
Flush stdout in s_mpi_abort before every abort: a prohibit's reason w…
sbryngelson Sep 16, 2026
6ac213f
Remove AMR subcycling: every level advances in lock-step at the case …
sbryngelson Sep 16, 2026
8879319
Retire the per-block fine advance: the batched advance is the only pa…
sbryngelson Sep 16, 2026
6591f58
Ledger README: the per-block advance is retired to the amr-per-block …
sbryngelson Sep 16, 2026
773a902
Ledger 167: the per-block advance leaves the branch (amr-per-block); …
sbryngelson Sep 16, 2026
245b386
Fix the CI red on 773a9028: rewrap the m_amr_registers Fypp header no…
sbryngelson Sep 16, 2026
99e8fa3
Share the AMR restart file layout between the simulation and post_pro…
sbryngelson Sep 16, 2026
c00284c
One exchange engine for the AMR waves (m_amr_wave): plan sides, pools…
sbryngelson Sep 16, 2026
cb9c71d
Plain Fortran for the single-instance Fypp templates in m_amr_exchang…
sbryngelson Sep 16, 2026
3bb14b0
Instrumentation: one phase-budget module, twelve top-level phases
sbryngelson Sep 16, 2026
a40e134
Split the two largest AMR routines
sbryngelson Sep 17, 2026
279502b
Delete the dead batched cons->prim conversion; trim the AMR comments
sbryngelson Sep 17, 2026
c3d76e7
Merge upstream master (IB prescribed kinematics and force history) in…
sbryngelson Sep 17, 2026
c4a07e8
Move the RK-stage AMR hooks out of m_time_steppers into m_amr_stage
sbryngelson Sep 17, 2026
66d200b
One gather mechanism: init and the regrid rebuild fill coarse patches…
sbryngelson Sep 17, 2026
9dea90e
Merge remote-tracking branch 'upstream/master' into up/mega
sbryngelson Sep 17, 2026
5526ac1
Stash migration on the wave engine
sbryngelson Sep 17, 2026
162ffd1
Split the grid swap, the L0 tile init and the restart reader
sbryngelson Sep 17, 2026
6fb2535
Merge upstream master (#1887 num_gps device refresh, already present …
sbryngelson Sep 18, 2026
77c0bbe
Cleanup: one dimension mask, one clusterer module, one of each geomet…
sbryngelson Sep 18, 2026
ba5f1e3
Registers: one capture kernel for creg and freg, one zeroing routine,…
sbryngelson Sep 18, 2026
c7371e8
L0 tiles: one box router for seed/scatter/reflux/restrict, one face-B…
sbryngelson Sep 18, 2026
0761c8a
Exchange/transfer: slab lists as (3,6) boxes, one restrict kernel bod…
sbryngelson Sep 18, 2026
8de0c12
Frame/store/state/L0: Fypp folds over the three coordinate axes, one …
sbryngelson Sep 18, 2026
3dff89b
Vectorize the per-dimension guards in regrid, distribution, transfer,…
sbryngelson Sep 18, 2026
2934c82
One IGR sigma seed loop for batched and single blocks; trim the store…
sbryngelson Sep 18, 2026
3a0dab5
Docs: describe the reflux exchange as the wave it now is
sbryngelson Sep 18, 2026
36063ed
One stable key sort for the SFC cut, the seam-pair build and the clus…
sbryngelson Sep 18, 2026
f1ab695
One receive drain and one unpack kernel for the level-1 and parent fi…
sbryngelson Sep 18, 2026
45ee6f2
s_amr_copy_fine_fields takes the block slot and reads its own bounds
sbryngelson Sep 18, 2026
034a3ca
Fix the coarse-ghost exchange test reading uninitialized locals
sbryngelson Sep 18, 2026
9cdd36a
Guard the five unchecked combinations review found, and make the AMR …
sbryngelson Sep 18, 2026
4eac13a
Make the three empty goldens real, and fail a golden case that produc…
sbryngelson Sep 18, 2026
4947327
Push restored AMR state to the device where the host copy is authorit…
sbryngelson Sep 18, 2026
3fe45d7
Apply the post-RK sources on fine blocks too, or reject them
sbryngelson Sep 18, 2026
677d8c9
Guard the jac_old seed kernel instead of cycling past it (nvfortran)
sbryngelson Sep 18, 2026
1654204
Carry a level>=2 block's own solution across a regrid, as level 1 alr…
sbryngelson Sep 18, 2026
f7342cf
Read the serial AMR overlay by the file's own record flag, not by geo…
sbryngelson Sep 18, 2026
ded685c
Restart a coexist case: read the L0 tile records instead of rejecting…
sbryngelson Sep 19, 2026
fb7dd61
Second review round: a case-optimized build could not compile, and th…
sbryngelson Sep 19, 2026
4769306
Test species diffusion across a rank seam, and correct two stale clai…
sbryngelson Sep 19, 2026
97ab30e
Cut the AI tics out of the user-facing AMR doc
sbryngelson Sep 19, 2026
d915d2c
Tighten the AMR doc: shorter sentences, fewer of them
sbryngelson Sep 19, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .lychee.toml
Original file line number Diff line number Diff line change
Expand Up @@ -33,5 +33,6 @@ exclude = [
"https://code\\.visualstudio\\.com/?$", # Root page returns 403 to automated requests
"https://stackoverflow\\.com", # Returns 403 to automated requests
"https://marketplace\\.visualstudio\\.com", # Returns 503 to automated requests
"_8md\\.html$", # Doxygen auto-links backticked *.md filenames in prose to per-file pages it never generates for markdown inputs; the real md_*.html page links are still checked
"https://web\\.eng\\.ucsd\\.edu", # San Diego mechanism page has an untrusted SSL cert
]
6 changes: 6 additions & 0 deletions .typos.toml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@ extend-ignore-identifiers-re = [
AttributeIDSupressMenu = "AttributeIDSupressMenu"

[default.extend-words]
# Cray CCE spells it this way in the lib-4425 runtime error; quoted verbatim in the AMR ledger so the
# message stays greppable against what the machine actually prints.
Unitialized = "Unitialized"
INOUT = "INOUT"
WRONLY = "WRONLY"
nd = "nd"
Expand All @@ -22,6 +25,9 @@ TKE = "TKE"
HSA = "HSA"
infp = "infp"
Sur = "Sur"
thi = "thi" # AMR clustering local: tagged-box hi index (tlo/thi)
alo = "alo" # AMR clustering local: accepted-box lo array (alo/ahi)
thr = "thr" # AMR clustering local: min-separation merge threshold
equil = "equil" # abbreviation for "equilibrium" (flamelet chemistry)
chioces = "chioces" # typo for "choices" - tests constraint key validation
reqires = "reqires" # typo for "requires" - tests dependency key validation
Expand Down
1 change: 1 addition & 0 deletions cmake/Fypp.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -119,6 +119,7 @@ macro(HANDLE_SOURCES target useCommon)
-D MFC_COMPILER="${CMAKE_Fortran_COMPILER_ID}"
-D MFC_CASE_OPTIMIZATION=False
-D chemistry=False
-D eos_state_dependent=False
--line-numbering
--no-folding
--line-length=999
Expand Down
6 changes: 6 additions & 0 deletions cmake/GPU.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -90,9 +90,15 @@ elseif (CMAKE_Fortran_COMPILER_ID STREQUAL "Cray")
add_link_options("SHELL:-hkeepfiles")

if (CMAKE_BUILD_TYPE STREQUAL "Debug")
# -h bounds: array-bounds and pointer checking, the Cray equivalent of gfortran's
# -fcheck=bounds,pointer / Intel's -check bounds / NVHPC's -Mbounds, all of which the
# debug branches above already set. Cray was the ONLY compiler whose debug build had no
# bounds checking, so an out-of-bounds write showed up here only as a later, unrelated
# allocation failing with an uninitialised descriptor.
add_compile_options(
"SHELL:-h acc_model=auto_async_none"
"SHELL: -h acc_model=no_fast_addr"
"SHELL: -h bounds"
"SHELL: -K trap=fp" "SHELL: -g" "SHELL: -O0"
)
add_link_options("SHELL: -K trap=fp" "SHELL: -g" "SHELL: -O0")
Expand Down
297 changes: 297 additions & 0 deletions docs/documentation/amr.md

Large diffs are not rendered by default.

505 changes: 505 additions & 0 deletions docs/documentation/amr_implementation.md

Large diffs are not rendered by default.

188 changes: 188 additions & 0 deletions docs/documentation/case.md

Large diffs are not rendered by default.

11 changes: 7 additions & 4 deletions docs/documentation/contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,10 +107,12 @@ reconfigure.

**Fypp and per-target stubs.** Fypp resolves `#:include` at parse time, so every `.fpp`
file sees exactly the include path for the target being compiled. `src/common/` is
compiled once per executable with the `MFC_<TARGET>` preprocessor define
(`MFC_PRE_PROCESS`, `MFC_SIMULATION`, or `MFC_POST_PROCESS`) — this is intentional.
It is what lets common modules include per-target generated files and gate
simulation-only code with `#ifdef MFC_SIMULATION` without duplication.
compiled once per executable, and that include path is how common modules pick up the
per-target generated files. Do not gate code on the executable: the per-executable
preprocessor guards (`MFC_PRE_PROCESS`, `MFC_SIMULATION`, `MFC_POST_PROCESS`) were
removed from the sources and the source lint rejects them. Stage-varying behaviour is
passed in as an argument or an initialization policy, and stage-only code lives in that
stage's directory.

For which of the 15 files is manual vs. generated and what each contains, see the
"How to Add a New Simulation Parameter" section below.
Expand Down Expand Up @@ -222,6 +224,7 @@ Both human reviewers and AI code reviewers reference this section.
- **Runtime checks go where they run.** Shared constraints belong in `src/common/m_checker_common.fpp`, simulation-only ones in `src/simulation/m_checker.fpp`, and pre- and post-process ones in their own `m_checker.fpp`. Those two `s_check_inputs` are currently empty; that is still the correct home for their checks, not `m_checker_common`.
- **Analytic initial conditions are compiled into the binary** and their expressions are AST-validated at case load, so syntax errors and unknown variables surface immediately and by name. Each IC variable maps to an `eqn_idx` expression in `QPVF_IDX_VARS` (`toolchain/mfc/case.py`); adding a patch-settable conserved variable means updating that map and the Fortran `eqn_idx` builder together, because a mismatch is a silent wrong index.
- **Under `--case-optimization` the baked-in constants are dropped from the namelist**, so changing one requires a rebuild rather than a case-file edit.
- **A Fypp variable that the generated `case.fpp` sets also needs a default in `cmake/Fypp.cmake`.** `./mfc.sh` writes a per-case `case.fpp` whose `#:set` lines define those variables, so every `./mfc.sh build` sees them. The checked-in fallback `src/common/include/case.fpp` defines none of them, and that is what the documentation build and any bare CMake build use — the `-D <name>=False` list in `cmake/Fypp.cmake` is the only thing keeping those builds alive. A new `#:set` without a matching `-D` therefore preprocesses fine everywhere a developer looks and fails only in the documentation lane, with a bare `name ... is not defined` from the file that reads it.

### Compiler Portability

Expand Down
13 changes: 13 additions & 0 deletions docs/documentation/gpuParallelization.md
Original file line number Diff line number Diff line change
Expand Up @@ -926,3 +926,16 @@ answer is wrong, or one backend diverges from all the others.


<div style='text-align:center; font-size:0.75rem; color:#888; padding:16px 0 0;'>Page last updated: 2026-02-04</div>

## Silent-failure traps (AMD flang)

- NEVER put a `GPU_PARALLEL_LOOP` inside a Fortran `block` construct: amdflang compiles it clean but silently
DROPS the region from the device image — the first launch dies with `HSA_STATUS_ERROR_INVALID_SYMBOL_NAME`
naming an `__omp_offloading_*` symbol. Hoist the kernel into its own module subroutine.
- **Never `GPU_UPDATE` a NON-CONTIGUOUS array section.** ``GPU_UPDATE(device='[q%%sf(a:b, c:d, e:f)]')`` on a
sub-box emits correct OpenMP, but AMD flang copies it as `size(section)` CONTIGUOUS elements starting at the
first: only the leading run lands where it is named and the rest overwrites neighbouring cells with stale data
— no error, no warning. A leading section (`arr(1:n)`, or a fixed trailing index like `freg(d)%%lo(:,:,:,k)`)
IS contiguous and safe; anything that strides is not. To move a sub-box, pack/unpack it with a device kernel
(`s_l0_pack_unpack_block`, `s_amr_restrict_device_wire`) — that is why those exist. Measured: 10 of 60 covered
cells delivered in the AMR cross-rank restrict, mass off 1.4e-5 per regrid.
1 change: 1 addition & 0 deletions docs/documentation/readme.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ Welcome to the Multi-component Flow Code (MFC) documentation.

- @ref architecture "Code Architecture" - How the source code is organized, data flow, and module map
- @ref expectedPerformance "Performance" - Optimization and benchmarks
- @ref amr "Adaptive Mesh Refinement" - Block-structured AMR: algorithm, physics support, and parameters
- @ref gpuParallelization "GPU Parallelization" - GPU macro API (developer reference)
- @ref docker "Containers" - Docker usage
- @ref troubleshooting "Troubleshooting" - Debugging and common issues
Expand Down
225 changes: 124 additions & 101 deletions docs/module_categories.json
Original file line number Diff line number Diff line change
@@ -1,103 +1,126 @@
[
{
"category": "Solver Core",
"modules": [
"m_rhs",
"m_time_steppers",
"m_weno",
"m_riemann_solvers",
"m_riemann_state",
"m_riemann_solver_hlld",
"m_riemann_solver_hll",
"m_riemann_solver_lf",
"m_riemann_solver_hllc",
"m_riemann_solver_hypo_hlld",
"m_muscl",
"m_variables_conversion",
"m_thinc"
]
},
{
"category": "Physics Models",
"modules": [
"m_viscous",
"m_hb_function",
"m_surface_tension",
"m_reactive_burn",
"m_bubbles",
"m_bubbles_EE",
"m_bubbles_EL",
"m_bubbles_EL_kernels",
"m_qbmm",
"m_hypoelastic",
"m_phase_change",
"m_chemistry",
"m_acoustic_src",
"m_body_forces",
"m_pressure_relaxation",
"m_collisions"
]
},
{
"category": "Boundary Conditions",
"modules": [
"m_cbc",
"m_compute_cbc",
"m_boundary_common",
"m_boundary_primitives",
"m_boundary_io",
"m_ibm",
"m_particle_cloud",
"m_igr",
"m_ib_patches",
"m_compute_levelset"
]
},
{
"category": "I/O and Startup",
"modules": [
"m_start_up",
"m_data_output",
"m_data_input",
"m_delay_file_access"
]
},
{
"category": "Infrastructure",
"modules": [
"m_derived_types",
"m_global_parameters",
"m_global_parameters_common",
"m_mpi_common",
"m_mpi_proxy",
"m_constants",
"m_precision_select",
"m_helper",
"m_helper_basic",
"m_compile_specific",
"m_fftw",
"m_nvtx",
"m_model",
"m_finite_differences",
"m_checker",
"m_checker_common",
"m_sim_helpers",
"m_derived_variables",
"m_patch_geometries"
]
},
{
"category": "Pre-Process",
"modules": [
"m_grid",
"m_icpp_patches",
"m_initial_condition",
"m_assign_variables",
"m_check_patches",
"m_check_ib_patches",
"m_perturbation",
"m_simplex_noise",
"m_boundary_conditions"
]
}
{
"category": "Solver Core",
"modules": [
"m_rhs",
"m_time_steppers",
"m_weno",
"m_riemann_solvers",
"m_riemann_state",
"m_riemann_solver_hlld",
"m_riemann_solver_hll",
"m_riemann_solver_lf",
"m_riemann_solver_hllc",
"m_riemann_solver_hypo_hlld",
"m_muscl",
"m_variables_conversion",
"m_thinc",
"m_active_box"
]
},
{
"category": "Physics Models",
"modules": [
"m_viscous",
"m_hb_function",
"m_surface_tension",
"m_reactive_burn",
"m_bubbles",
"m_bubbles_EE",
"m_bubbles_EL",
"m_bubbles_EL_kernels",
"m_qbmm",
"m_hypoelastic",
"m_phase_change",
"m_chemistry",
"m_acoustic_src",
"m_body_forces",
"m_pressure_relaxation",
"m_collisions"
]
},
{
"category": "Boundary Conditions",
"modules": [
"m_cbc",
"m_compute_cbc",
"m_boundary_common",
"m_boundary_primitives",
"m_boundary_io",
"m_ibm",
"m_particle_cloud",
"m_igr",
"m_ib_patches",
"m_compute_levelset"
]
},
{
"category": "I/O and Startup",
"modules": [
"m_start_up",
"m_data_output",
"m_data_input",
"m_delay_file_access",
"m_load_weight",
"m_load_balance",
"m_sfc_partition"
]
},
{
"category": "Infrastructure",
"modules": [
"m_derived_types",
"m_global_parameters",
"m_global_parameters_common",
"m_mpi_common",
"m_mpi_proxy",
"m_constants",
"m_precision_select",
"m_helper",
"m_helper_basic",
"m_compile_specific",
"m_fftw",
"m_nvtx",
"m_model",
"m_finite_differences",
"m_checker",
"m_checker_common",
"m_sim_helpers",
"m_derived_variables",
"m_patch_geometries",
"m_box",
"m_amr",
"m_amr_state",
"m_amr_distribution",
"m_amr_wave",
"m_amr_store",
"m_amr_exchange",
"m_amr_frame",
"m_amr_transfer",
"m_amr_advance",
"m_amr_stage",
"m_amr_l0",
"m_amr_cluster",
"m_amr_regrid",
"m_amr_restart",
"m_amr_restart_io",
"m_amr_registers",
"m_amr_xchg_audit",
"m_phase_timing"
]
},
{
"category": "Pre-Process",
"modules": [
"m_grid",
"m_icpp_patches",
"m_initial_condition",
"m_assign_variables",
"m_check_patches",
"m_check_ib_patches",
"m_perturbation",
"m_simplex_noise",
"m_boundary_conditions"
]
}
]
Loading
Loading