Skip to content

Carry aligned reads as CRAM, with --aligned_format cram|bam - #207

Draft
ljwharbers wants to merge 13 commits into
devfrom
experiment/cram-io
Draft

ljwharbers wants to merge 13 commits into
devfrom
experiment/cram-io

Conversation

@ljwharbers

@ljwharbers ljwharbers commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Carry the aligned and haplotagged reads as CRAM instead of BAM, behind a new --aligned_format cram|bam (default cram). On long reads this makes the aligned files a third (ONT) to two thirds (Revio) smaller, with identical records and identical calls. Full write-up with all measurements (shared on request): https://claude.ai/artifact/9gUVWBRFuLtv5kSFBgp3aR

Draft: not ready for review yet (see "Before review" below).

Why these CRAM settings

Every tool downstream of alignment was re-run on CRAM with its own command and container, with the reference not reachable through the CRAM header. Two settings turned out to be required, because the failures without them are silent:

  • CRAM 3.0, not 3.1. Clair3 and ClairS-TO's Verdict bundle an htslib that cannot decode 3.1 slices. Clair3 then writes an empty VCF and exits 0; Verdict crashes. On 3.0 both match BAM.
  • Reference embedded (embed_ref=1). ClairS's internal samtools mpileup, Severus and SAVANA's read counting open the reads without the FASTA. ClairS then writes empty VCFs and exits 0. It only appears to work on VSC because the UR: path in the header, which points into the minimap2 work dir, happens to be mounted inside the containers. An embedded reference adds 1-7 % to the file and makes each CRAM readable anywhere.

Changes

  • MINIMAP2_ALIGN (nf-core, patched): bam_format = 'cram' makes samtools sort write -O cram --reference, with a cram emit and a .crai index.
  • SAMTOOLS_MERGE gets the FASTA. The merge/index joins take whichever of bam/cram and bai/crai the format produced, so merged replicates no longer depend on .out.bam.
  • LONGPHASE_HAPLOTAG (nf-core, patched): with ext.cram, longphase writes into a FIFO that samtools encodes, because longphase's own --cram cannot embed the reference. No temporary BAM reaches the disk. PHASING_HAPLOTYPING joins on bai or crai.
  • CRAMINO_POST gets --reference. ASCAT keeps no FASTA: the module's image has no ref.fasta argument, so passing one fails ascat.prepareHTS.
  • --aligned_format (schema, nextflow.config, conf/modules.config closures). bam keeps today's files and .bai indexes.
  • These apply to both formats:
    • Duplicate RG tags fixed: every aligned record carried two RG tags, the basecaller's (via samtools fastq -T '*' and -y) and the pipeline's -R read group. CRAM keeps one per record and kept the basecaller's, which has no @RG line. samtools reset -x RG now drops the uBAM's copy.
    • CLAIRSTO cleanup: deletes tmp_<sample>/phasing_output/phased_bam_output when done, the full haplotagged copy of the tumour reads (157-162 GiB for one ONT sample).
  • Docs: docs/usage.md, docs/output.md, CHANGELOG.md.

Validation

  • Per-tool probe (test profile, CRAM 3.0 + embedded reference): identical to BAM for Clair3, ClairS, ClairS-TO incl. Verdict, DeepVariant/DeepSomatic, longphase phase/modcall/haplotag, Severus, SAVANA, mosdepth, samtools stats/flagstat/idxstats and cramino.
  • Test profile, BAM vs CRAM:
    • Clair callers: 124 files identical.
    • Deep callers: 119 files identical.
    • Replicate samplesheet (test_sheet_2, sample4 merged), --aligned_format bam vs cram: 185 files identical, merged reads included.
    • The only differences are Severus read_ids.csv, which also differs between two BAM runs.
  • Full sample (BL1_ont, ONT fiber-seq, tumour-only, CHM13):
    • Identical to BAM: ClairS-TO somatic and germline calls, Severus and SAVANA SVs, ASCAT purity/ploidy, Wakhan solutions, modkit combined pileup.
    • Outdir 167 → 108 GiB. Work dir 707 → about 440 GiB with the ClairS-TO cleanup.
    • Runtime within ±5 % on the large steps.
    • Phase-set IDs differ for 0.28 % of germline variants. The cause is longphase modcall, which is not reproducible: run twice on the same BAM it gave 100 and 102 sites. It is not caused by CRAM.
  • nf-test, apptainer on Mindwell, after the dev merge: all five pipeline tests pass. deep_only runs --aligned_format bam; the other four run the CRAM default. That includes clair_only, the only test with replicates, so sample4's two replicates go through the CRAM merge. nft-bam reads the merged CRAM without a FASTA (embedded reference), and its reads MD5 equals the one the BAM merge gave (88c8d3cf…, 7272 reads), so the snapshot entry is unchanged. deep_only is also the test DeepVariant/DeepSomatic slow down most on the thin CRAM test data (79 → 11 min); on BAM it reproduces dev's snapshot md5s exactly. Snapshots were regenerated; the only drift is bamfiles/ (.bam/.bai ↔ .cram/.crai) and two files whose header names the input file (samtools .stats, Severus breakpoints_double.csv). CRAM/CRAI md5s are not snapshotted, because multi-threaded samtools lays the containers out differently on every run.

Before review

  • On the thin chr19 test data, DeepVariant/DeepSomatic make_examples are 13-14× slower on CRAM. On real data there is no slowdown: 88 s BAM vs 87 s CRAM on chr20:10-15 Mb, same 111 224 examples. deep_only runs on BAM, and consensus/union still run both deep callers on CRAM.
  • Coral #205 (CoRAL) adds readers of the aligned reads that have only seen BAM. Whichever of Coral #205/Carry aligned reads as CRAM, with --aligned_format cram|bam #207 merges second has to check CoRAL on CRAM.
  • -profile conda needs samtools in the patched longphase/haplotag environment for the FIFO; it has been added, but conda is not in the CI matrix.

PR checklist

  • This comment contains a description of changes (with reason).
  • If you've fixed a bug or added code that should be tested, add tests!
  • Make sure your code lints (nf-core pipelines lint).
  • Ensure the test suite passes (apptainer on Mindwell, after the dev merge).
  • Check for unexpected warnings in debug mode.
  • Usage Documentation in docs/usage.md is updated.
  • Output Documentation in docs/output.md is updated.
  • CHANGELOG.md is updated.

🤖 Generated with Claude Code

ljwharbers and others added 10 commits September 23, 2026 14:30
MINIMAP2_ALIGN (patched) writes a reference-compressed CRAM + .crai when
bam_format is 'cram'; SAMTOOLS_MERGE gets the reference and the joins
follow the cram/crai outputs; LONGPHASE_HAPLOTAG writes CRAM (--cram) and
PHASING_HAPLOTYPING joins on bai or crai. ASCAT and CRAMINO_POST now get
the FASTA so they can decode CRAM. CLAIRSTO drops its intermediate
haplotagged copy of the tumor reads (as large as the input) once done.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
samtools fastq -T '*' carried the basecaller's RG tag through minimap2 -y,
next to the @rg added by -R, so every record had two RG tags. BAM keeps both
(readers see the first), but CRAM keeps one read group per record: the last,
the basecaller's, which has no @rg line in the aligned header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Probe on the test-profile tasks (CRAM input, reference not reachable
through the header, no EBI fallback):
- Clair3 and ClairS-TO's Verdict alleleCounter bundle an htslib that cannot
  decode CRAM 3.1 slices; Clair3 then calls nothing and still exits 0.
  CRAM 3.0 decodes fine.
- ClairS (internal mpileup), Severus and SAVANA's read counting open the
  reads without the reference; ClairS then writes empty VCFs and exits 0.
  embed_ref=1 makes every CRAM decodable without one, at ~1-3% size.

longphase haplotag's own --cram cannot embed the reference, so with
ext.cram it writes into a FIFO that samtools encodes (ext.args2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ASCAT in the module's image has no ref.fasta argument, so passing the
FASTA fails ascat.prepareHTS outright; the CRAMs embed their reference, so
alleleCounter decodes them without one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0.5.1 writes tmp_<sample_name>/, so the cleanup missed the 162 GiB
haplotagged copy of BL1_ont's tumour reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MINIMAP2_ALIGN, SAMTOOLS_MERGE and LONGPHASE_HAPLOTAG write CRAM 3.0 with
the reference embedded only when aligned_format is cram; bam keeps BAM and
.bai throughout. The merge/index joins take whichever output the format
produced. The uBAM RG-tag fix applies to both formats.

Tests: clair_only (the replicate samplesheet) runs --aligned_format bam and
keeps its merged-reads MD5 check; the other pipeline tests expect CRAM.
Snapshots not yet regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All five pipeline tests pass on apptainer (default, consensus, deep_only,
union on CRAM; clair_only on --aligned_format bam, snapshot unchanged).
The only drift is bamfiles/ (.bam/.bai -> .cram/.crai) and two files whose
header names the input file: samtools stats' command line and Severus'
breakpoints_double.csv column names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

nf-core pipelines lint overall result: Passed ✅ ⚠️

Posted for pipeline commit d5b9d7b

+| ✅ 215 tests passed       |+
#| ❔  21 tests were ignored |#
#| ❔   1 tests had warnings |#
!| ❗  40 tests had warnings |!
Details

❗ Test warnings:

  • nextflow_config - Config manifest.version should end in dev: 1.1.0
  • pipeline_todos - TODO string in nextflow.config: Specify your pipeline's command line flags
  • pipeline_todos - TODO string in nextflow.config: Update the field with the details of the contributors to your pipeline. New with Nextflow version 24.10.0
  • pipeline_todos - TODO string in README.md: Include a figure that guides the user through the major workflow steps. Many nf-core
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in lint_log.txt: Named file extensions MUST be emitted for ALL output channels
  • pipeline_todos - TODO string in lint_log.txt: List additional required output channels/values here
  • pipeline_todos - TODO string in meta.yml: #Add a description of the module and list keywords
  • pipeline_todos - TODO string in methods_description_template.yml: #Update the HTML below to your preferred methods description, e.g. add publication citation for this pipeline
  • pipeline_todos - TODO string in nextflow.config: Specify any additional parameters here
  • pipeline_todos - TODO string in base.config: Check the defaults for all processes
  • pipeline_todos - TODO string in base.config: Customise requirements for specific processes.
  • pipeline_todos - TODO string in CONTRIBUTING.md: Add any pipeline specific contribution guidelines here, such as coding styles, procedures, checklists etc.
  • schema_description - Ungrouped param in schema: skip_modkit
  • schema_description - No description provided in schema for parameter: generate_gvcf
  • schema_description - No description provided in schema for parameter: autocorrelation
  • schema_description - No description provided in schema for parameter: vep_custom
  • schema_description - No description provided in schema for parameter: vep_custom_tbi
  • schema_description - No description provided in schema for parameter: severus_minsupport
  • schema_description - No description provided in schema for parameter: wakhan_chroms
  • local_component_structure - prepare_vep_plugins.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - deepsomatic.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - small_variant_consensus.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - prepare_reference_files.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - prepare_annotation.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - prepare_signatures.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure
  • local_component_structure - phasing_haplotyping.nf in subworkflows/local should be moved to a SUBWORKFLOW_NAME/main.nf structure

❔ Tests ignored:

  • files_exist - File is ignored: CODE_OF_CONDUCT.md
  • files_exist - File is ignored: assets/nf-core-lrsomatic_logo_light.png
  • files_exist - File is ignored: docs/images/nf-core-lrsomatic_logo_light.png
  • files_exist - File is ignored: docs/images/nf-core-lrsomatic_logo_dark.png
  • files_exist - File is ignored: .github/ISSUE_TEMPLATE/config.yml
  • files_exist - File is ignored: .github/workflows/awstest.yml
  • files_exist - File is ignored: .github/workflows/awsfulltest.yml
  • files_exist - File is ignored: .github/CONTRIBUTING.md
  • nextflow_config - Config variable ignored: manifest.name
  • nextflow_config - Config variable ignored: manifest.homePage
  • files_unchanged - File ignored due to lint config: CODE_OF_CONDUCT.md
  • files_unchanged - File ignored due to lint config: .github/ISSUE_TEMPLATE/bug_report.yml
  • files_unchanged - File ignored due to lint config: .github/PULL_REQUEST_TEMPLATE.md
  • files_unchanged - File ignored due to lint config: .github/workflows/branch.yml
  • files_unchanged - File ignored due to lint config: .github/workflows/linting.yml
  • files_unchanged - File ignored due to lint config: assets/email_template.txt
  • files_unchanged - File ignored due to lint config: assets/nf-core-lrsomatic_logo_light.png
  • files_unchanged - File ignored due to lint config: docs/images/nf-core-lrsomatic_logo_light.png
  • files_unchanged - File ignored due to lint config: docs/images/nf-core-lrsomatic_logo_dark.png
  • files_unchanged - File ignored due to lint config: docs/README.md
  • actions_awstest - 'awstest.yml' workflow not found: /home/runner/work/lrsomatic/lrsomatic/.github/workflows/awstest.yml

❔ Tests fixed:

✅ Tests passed:

Run details

  • nf-core/tools version 4.1.0
  • Run at 2026-09-28 18:14:54

ljwharbers and others added 3 commits September 28, 2026 15:30
samtools view --threads lays CRAM containers out differently on every
run, so the published CRAM and CRAI bytes never repeat: PR #207 CI got a
different set of md5s on each shard and attempt while flagstat, idxstats
and stats (the read content) matched. Re-encoding one BAM three times
with the module's image confirms it: with 4 threads the .crai changes
between identical runs, with 1 thread only the path-derived file ID and
@pg differ.

Ignore */bamfiles/*.cram{,.crai} in tests/.nftignore (the names stay in
stable_name) and drop the 10 md5 entries from the four CRAM snapshots.
Verified on Mindwell: default (twice), consensus, deep_only and union
pass without --update-snapshot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The CHANGELOG said ASCAT now gets the FASTA, which 77bcf54 reverted, and
quoted the CRAM 3.1 size savings and a 1-3 % embed_ref cost. The shipped
3.0 + embed_ref files are 33 % (ONT), 52 % (older PacBio) and 69 % (Revio)
smaller, and the embedded reference costs 1-7 %. It now also says that a
resumed run realigns, since -x RG changes MINIMAP2_ALIGN in both formats.

The patched LONGPHASE_HAPLOTAG pipes longphase's output through
samtools view, which its container has but its conda environment did not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
clair_only is the only pipeline test with replicates, but it ran
--aligned_format bam, so the default CRAM merge (SAMTOOLS_MERGE with the
FASTA and version=3.0/embed_ref=1) had no test. It now runs the default.
nft-bam's htsjdk decodes the embedded-reference CRAM without a FASTA, and
sample4's merged reads MD5 is the one the BAM merge gave (88c8d3cf...,
7272 reads), so that snapshot entry is unchanged.

deep_only takes --aligned_format bam instead. Its snapshot md5s are now
identical to dev's, and the test drops from 79 to 11 min, since
DeepVariant/DeepSomatic make_examples are slow on CRAM only on the thin
chr19 test data. consensus and union still run both deep callers on CRAM.

Snapshots regenerated with apptainer on Mindwell; the drift is bamfiles/,
samtools .stats and Severus breakpoints_double.csv, whose headers name
the input file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant