Skip to content

Add compressed binary accessibility caches (.agz) - #250

Open
Alexander-Mitrofanov wants to merge 5 commits into
masterfrom
feat/245-accessibility-binary
Open

Alexander-Mitrofanov wants to merge 5 commits into
masterfrom
feat/245-accessibility-binary

Conversation

@Alexander-Mitrofanov

@Alexander-Mitrofanov Alexander-Mitrofanov commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Repeated target screens spend time formatting and parsing large accessibility matrices. This adds gzip-compressed .agz caches containing versioned Boost binary archives of exact internal ED values, including the extra interval length needed for dangling-end probabilities.

Use --out=tAcc:target.agz (or any existing query/target Acc/Pu output), then reuse it with --tAcc=E --tAccFile=target.agz. The extension selects the binary reader in either E or P mode; both use the stored ED values unchanged. Existing text and .gz behavior is preserved.

Unconstrained AccessibilityVrna and AccessibilityFromStream pass their stored ED rows directly to Boost through non-owning row views, without copying rows or calling getED() per cell. Other accessibility implementations and constrained data retain generic export to preserve their accessibility semantics. Loading fills retained matrix rows in place and validates discarded tails when reading a narrower band. The version 1 archive layout is unchanged.

The reader checks the format, sequence, dimensions, energy values, trailing data and gzip integrity. Smaller requested interaction lengths are supported. Boost.Serialization is checked during configure and linked into the CLI and pkg-config consumers. README, CLI help and ChangeLog document usage and compatibility.

Validation on Linux x86-64 with GCC 16.2.0, Boost 1.85.0 and ViennaRNA 2.7.2:

  • Release with native std::mdspan and debug with bundled Kokkos mdspan: make tests -j2 passes all 73,630 assertions in 46 API cases, plus both CLI suites. The debug build was built from the source distribution.
  • Exact round trips cover ViennaRNA, base-pair, disabled, reversed and stream-loaded accessibilities, constraints/infinity, requested lengths, and malformed/truncated archives. Direct export tests verify zero getED() calls and byte-identical generic/direct payloads; a constrained short-sequence regression preserves masking. Invalid discarded-band values are rejected.
  • CLI checks verify identical predictions, P/E input, query/target paths, uppercase extensions, multiple sequences, compressed text reuse, and controlled errors for truncated trailers and invalid CRCs.
  • make install, independent compilation of affected installed public headers, and a linked installed-library consumer/benchmark. An archive produced by the original version 1 implementation reloads and re-exports byte-for-byte identically.
  • make dist includes the benchmark source, documentation and per-trial CSV data; git diff --check passes.

Compression was measured after the direct matrix optimization, using the same 100,000-base ED matrix for each format. Median of three trials for AccessibilityVrna:

Dataset Format Write (s) Read (s) Bytes
Synthetic RNA raw binary 0.105 0.022 40,479,873
Synthetic RNA binary + gzip 2.796 0.248 16,620,982
E. coli NC_000913.3, bases 1–100,000 raw binary 0.108 0.024 40,479,873
E. coli NC_000913.3, bases 1–100,000 binary + gzip 3.019 0.271 16,850,736

Gzip is retained: it costs roughly 2.7–2.9 seconds per write and 0.23–0.25 seconds per read in this run, but saves 58–59% of storage. Raw archives are about 2.4 times larger. The benchmark also compares direct/generic export and AccessibilityFromStream; all 48 reloads matched every ED cell exactly. Gzip dominates compressed write time, while the raw timings show the benefit of direct export.

Reproduction commands, hardware, full tables and CSV data are included. Timings include stream close but exclude folding/verification, use cached filesystem reads, and do not include fsync. They measure local I/O, not end-to-end genome-screen speedups.

Compatibility: archives use Boost's native binary format, requiring a compatible architecture/archive version. Cached temperature/folding/model settings are not stored or checked; reuse with matching settings. macOS/Apple Clang has not been tested locally.

Fixes #245.

@Alexander-Mitrofanov

Copy link
Copy Markdown
Collaborator Author

@martin-raden The implementation for #245 is ready for review. This PR adds compressed .agz accessibility caches through the existing CLI options. Local release and debug tests pass; benchmark results and compatibility details are included in the PR description.

Comment thread src/IntaRNA/general.cpp

// gzipped input file stream
if (boost::iends_with(in, ".gz")) {
if (boost::iends_with(in, ".gz") || boost::iends_with(in, ".agz")) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

check if useful for this kind of data

@martin-raden

Copy link
Copy Markdown
Member

@Alexander-Mitrofanov

revise the code of this PR given the following requests:

  • consider my recent commits to this branch
  • (step 1) extend the serialization interface and AccessibilityArchive class in a way, such that the default AccessibilityVrna class can export its stored EdMatrix directly without copying the data (which is currently done, if I understand it correctly). The same holds for AccessibilityFromStream in case the read in values are to be stored again. Other Accessibility subclasses like AccessibilityBasePair and AccessibilityDisabled can use the current generic save strategy, which accesses the data via the subclass' interface.
  • (step 2) after the optimizations from above are implemented: evaluate and check whether zip compression (a) results in significantly smaller output file sizes and (b) how much time the additional compression costs. if none-compressed output is not much larger but can be done significantly faster (especially reading), remove zip compression. otherwise maintain it.

@Alexander-Mitrofanov

Copy link
Copy Markdown
Collaborator Author

@martin-raden Addressed both requested steps in 2e2c5f0, on top of your commits:

  • AccessibilityVrna and AccessibilityFromStream now serialize their stored ED row views directly, without row copies or per-cell getED() calls. Other producers keep generic export. Constrained data also keeps the generic path to preserve masking. Reads fill retained rows directly. The version 1 archive bytes are unchanged.
  • I kept gzip after measuring both synthetic RNA and E. coli (100,000 bases each): raw archives are 40.5 MB versus 16.6–16.9 MB compressed, a 58–59% saving. With direct export, gzip adds about 2.7–2.9 s per write and 0.23–0.25 s per read on this machine. Raw is faster but about 2.4× larger, so the size saving is substantial.

Full measurements, direct/generic comparisons, raw trial data and reproduction commands are included. All 48 benchmark reloads matched exactly. Release/native and debug/Kokkos suites pass (73,630 assertions in 46 API cases, plus both CLI suites), along with installed-header and original-archive compatibility checks.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

IntaRNA-specific binary format for probability data

2 participants