Skip to content

Reduce memory usage of sample-level result accumulators - #109

Merged
akrivi merged 11 commits into
mainfrom
al/fix_threads_samples_memory
Sep 2, 2026
Merged

akrivi merged 11 commits into
mainfrom
al/fix_threads_samples_memory

Conversation

@akrivi

@akrivi akrivi commented May 20, 2026 •

Copy link
Copy Markdown
Collaborator

Previously, threaded execution created worker-local sample accumulators sized for the full Monte Carlo sample count. For sample-based result specs like ShortfallSamples this caused memory usage to scale as O(regions x timesteps x samples x threads). After the CVAR implementation added per-sample shortfall totals, Shortfall was affected as well.

This PR changes sample-based result accumulation so each threaded worker's recorder stores only that worker's assigned sample range. The partitions are then copied into their corresponding positions during finalization to produce the complete full-size sample result. Non-sample statistics continue to be merged as before.

Example for 3 threaded workers:

Before, each worker's sample accumulator allocated the full sample dimension:

Worker 1 accumulator: [total number of samples] 

Worker 2 accumulator: [total number of samples]

Worker 3 accumulator: [total number of samples]

Now, samples are split into ranges:

Worker 1 accumulator: [first 1/3 of total number of samples] 

Worker 2 accumulator: [second 1/3 of total number of samples]

Worker 3 accumulator: [last 1/3 of total number of samples]

Benchmarks

System: Guam 2028, 13 regions, 8760 timestamps, hourly resolution
Simulation: Run on HPC, using standard nodes (104 cores, 250 GB)
Result: ShortfallSamples()

1000 MC Samples

threads Elapsed time(s) - main Elapsed time(s) - PR Max RSS(GB) - main Max RSS(GB) - PR
1 23.61 23.75 1.57 1.57
2 13.04 12.96 2.37 2.38
4 8.28 8.16 4.09 2.38
8 6.68 5.24 7.47 2.41
16 6.84 3.87 14.31 2.43
24 7.39 3.48 21.14 2.46
32 9.59 3.28 27.99 2.49
48 11.16 3.06 41.69 2.62
64 14.53 2.97 55.28 2.61
80 15.04 2.96 69.01 2.75
96 21.48 3.06 82.61 2.79

10000 MC Samples

threads Elapsed time(s) - main Elapsed time(s) - PR Max RSS(GB) - main Max RSS(GB) - PR
1 210.73 208.58 9.21 9.20
2 109.29 108.45 17.67 17.65
4 61.39 56.88 34.64 17.64
8 46.96 33.21 68.58 17.66
16 43.08 18.88 136.53 17.69
24 51.95 13.68 204.37 17.77

@codecov-commenter

codecov-commenter commented May 20, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.87500% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 84.57%. Comparing base (734c206) to head (0151207).

Files with missing lines Patch % Lines
PRASCore.jl/src/Results/Shortfall.jl 81.81% 2 Missing ⚠️
PRASCore.jl/src/Simulations/Simulations.jl 97.14% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #109      +/-   ##
==========================================
+ Coverage   84.52%   84.57%   +0.04%     
==========================================
  Files          45       45              
  Lines        2592     2658      +66     
==========================================
+ Hits         2191     2248      +57     
- Misses        401      410       +9     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@akrivi
akrivi force-pushed the al/fix_threads_samples_memory branch from f6313c9 to b4b5ba3 Compare July 5, 2026 21:26
@akrivi
akrivi force-pushed the al/fix_threads_samples_memory branch from b113d5e to 8adce48 Compare August 3, 2026 17:47
@akrivi akrivi changed the title fix: High memory usage from sample-level result accumulators Reduce memory usage of sample-level result accumulators Aug 3, 2026
@akrivi
akrivi requested a review from scdhulipala August 5, 2026 19:11
Comment thread PRASCore.jl/src/Results/Results.jl Outdated
Comment thread PRASCore.jl/src/Results/Results.jl Outdated
@akrivi
akrivi requested a review from scdhulipala August 17, 2026 17:02

sampledata(acc::LineAvailabilityAccumulator) = acc.available

accumulatortype(::LineAvailability) = LineAvailabilityAccumulator

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is accumulatortype only defined for LineAvailability? Was this not defined to begin with?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

accumulatortype(::LineAvailability) was already defined before this PR. Each Resultspec defines its corresponding accumulatortype method in its own source file. The only new method this PR added here is sampledata(::LineAvailabilityAccumulator)

Comment thread PRASCore.jl/src/Results/Results.jl Outdated
Comment thread PRASCore.jl/src/Results/Results.jl Outdated
import ..Results
import ..Results: ResultSpec, ResultAccumulator,
accumulator, resultchannel, finalize
accumulator, finalize, resultchannel, usesamplepartitions

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

usesamplepartitions is internal API, right?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s correct and it is not exported from Results. Simulations.jl imports it only to determine whether each worker accumulator should be allocated for its local sample partition or for the complete sample count.

else
assess(system, method, sampleseeds, results, resultspecs...)
if method.threaded && threads == 1
@warn "It looks like you haven't configured JULIA_NUM_THREADS before you started the julia repl. \n If you want to use multi-threading, stop the execution and start your julia repl using : \n julia --project --threads auto"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we also change this warning to reflect the doc string?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I guess what I was saying here is: threaded::Bool=true`: Enable threaded Monte Carlo simulation is included in the docs. Should we just rephrase this warning to say the same? To enable threaded Monte Carlo ... I will resolve this for now.

Comment thread PRASCore.jl/src/Simulations/Simulations.jl
Comment thread PRASCore.jl/test/Simulations/runtests.jl
Comment thread PRASCore.jl/test/Simulations/runtests.jl
@akrivi
akrivi force-pushed the al/fix_threads_samples_memory branch from 97dc530 to 0151207 Compare September 1, 2026 01:08
@akrivi
akrivi requested a review from scdhulipala September 1, 2026 01:24
@akrivi
akrivi merged commit 0deb435 into main Sep 2, 2026
33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants