diff --git a/CHANGELOG.md b/CHANGELOG.md index 3f6e2a92..51bfabd1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -20,6 +20,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### `Changed` +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - In tumour-only mode with both `deepvariant` and `deepsomatic` selected, DeepVariant's germline calls are now kept only where DeepSomatic's verdict is `GERMLINE` or `PON` (`INFO/DS_VERDICT`); with `deepvariant` alone they are unfiltered and may include somatic variants (@robert-a-forsyth). +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - The phased somatic VCF is now split from the merged germline+somatic VCF by an `INFO/SOMATIC` provenance flag (new `VCFTAG` module) instead of by position, so germline records at somatic coordinates no longer leak into it; ClairS-TO germline records keep their original `FILTER` in `INFO/ORIG_FILTER` (@robert-a-forsyth). +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - Added `--smallvar_filter_pass` (default `true`), which restricts each small variant caller's output to `PASS` records before the caller consensus, phasing, VEP and the report; the published per-caller VCFs are unchanged (@robert-a-forsyth). - [#201](https://github.com/IntGenomicsLab/lrsomatic/pull/201) - `CLAIRSTO` and `CLAIRSTO_VERDICT_TAG` now pull the fork image from Docker Hub: `oras://docker.io/ljwharbers/clairs-to-sif:0.5.1-verdict-chm13-c0687e8-flat` under Singularity/Apptainer and `docker.io/ljwharbers/clairs-to:0.5.1-verdict-chm13-c0687e8-flat` otherwise, instead of `ghcr.io/ljwharbers/clairs-to`. The `-cpu` SIF on ghcr failed with `PROTOCOL_ERROR` on slow links: ghcr redirects every blob download to an Azure URL that expires at the next 5-minute mark and resets a stream still open then, and Apptainer resumes neither an `oras://` nor a `docker://` download. Docker Hub's download URLs are valid for 50 minutes and only checked when the request starts. `-flat` is the same software copied into an empty image in a few layers (3.3 GB instead of 7 GB); the software and its outputs are unchanged. `docs/usage.md` describes `pullTimeout`, Docker Hub's anonymous pull limit, pre-pulling, and how to recover the remaining `oras://ghcr.io` SIFs resumably (@ljwharbers). - [#199](https://github.com/IntGenomicsLab/lrsomatic/pull/199) - `CLAIRSTO` and `CLAIRSTO_VERDICT_TAG` now run the `-cpu` rebuild of the fork image (`0.5.1-verdict-chm13-c0687e8-cpu`), which swaps PyTorch's CUDA build for the CPU build of the same version. The software is otherwise unchanged, but the Apptainer SIF drops from 6.53 GB to 3.46 GB. The old image could not be pulled on a normal VSC link: Apptainer fetches an `oras://` SIF as a single unresumable stream, and the signed blob URL ghcr redirects to expires on a 15-minute wall-clock boundary, so 6.53 GB needed 7.3 MB/s sustained and was otherwise cut mid-transfer with `PROTOCOL_ERROR` (@ljwharbers). - [#197](https://github.com/IntGenomicsLab/lrsomatic/pull/197) - `CLAIRSTO` now runs `ghcr.io/ljwharbers/clairs-to:0.5.1-verdict-chm13-c0687e8` (ClairS-TO 0.5.1) instead of `docker.io/hkubal/clairs-to:v0.4.2`: a fork that lets Verdict read its CNA resources from `--cna_resource_dir`, fixes four places where Verdict's Python port of ASCAT departed from R, and disables Verdict with a warning when its resources cannot be read. **GRCh38 results move as well as CHM13 ones.** Revert to the upstream image once HKU-BAL/ClairS-TO carries these changes. The module also selects the SIF under `-profile apptainer` and sets explicit output prefixes (@ljwharbers). @@ -41,6 +44,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### `Fixed` +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - `*_var_combine = 'all'` now keeps both callers' private calls; the prioritized caller's own private calls were previously dropped. Invalid `combine_method`/`prioritize_caller` values now raise an error (@robert-a-forsyth). +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - The caller consensus now splits multi-allelic records before intersecting, so they match across callers, and rejoins them before phasing (@robert-a-forsyth). +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - `LRSOMATICREPORT` now renders each sample as soon as its own inputs are ready instead of waiting for the whole batch, including samples without ASCAT's optional raw segments output (@robert-a-forsyth). +- [#200](https://github.com/IntGenomicsLab/lrsomatic/pull/200) - Bumped `WAKHAN` from 0.4.3 to 0.4.4, fixing a deterministic `StatisticsError` crash in BAF binning that aborted the run (@robert-a-forsyth). - [#196](https://github.com/IntGenomicsLab/lrsomatic/pull/196) - `LRSOMATICREPORT` now points `XDG_CACHE_HOME` at the task directory alongside `HOME` and `TMPDIR`. Singularity/Apptainer inherit the host environment, so on sites that set it outside the bind-mounted work tree the render died with `Read-only file system (os error 30): mkdir '<...>/.cache/quarto'` (@AmberVerhasselt, @ljwharbers). - [#193](https://github.com/IntGenomicsLab/lrsomatic/pull/193) - `--vep_eve https://evemodel.org/api/proteins/bulk/download/` was rejected at launch because the "needs preparing" check keyed on a `.zip` suffix; it now checks whether the value is already a prepared bgzipped file (@AmberVerhasselt). - [#188](https://github.com/IntGenomicsLab/lrsomatic/pull/188) - `MODKIT_PILEUP` now runs a patched modkit 0.6.4 ([ljwharbers/modkit@pacbio-conflict-fix](https://github.com/ljwharbers/modkit/tree/pacbio-conflict-fix)): `ghcr.io/ljwharbers/modkit:0.6.4-pacbiofix-6e0afa2` under Docker and `oras://ghcr.io/ljwharbers/modkit-sif:0.6.4-pacbiofix-6e0afa2` under Singularity/Apptainer. Stock modkit 0.4.3-0.6.4 dropped 32-65 % of reads from recent PacBio HiFi BAMs and returned empty `--cpg` pileups ([nanoporetech/modkit#612](https://github.com/nanoporetech/modkit/issues/612), fix proposed in [nanoporetech/modkit#720](https://github.com/nanoporetech/modkit/pull/720)), and ignored `--phased`/`--modified-bases` for PacBio BAMs with 6mA calls. The image is `linux/amd64` only and Conda is not supported (use `--skip_modkit` there); return to the biocontainer once a release includes the fix (@ljwharbers). diff --git a/conf/modules.config b/conf/modules.config index 8b20e4fd..c87d5dff 100644 --- a/conf/modules.config +++ b/conf/modules.config @@ -121,13 +121,26 @@ process { withName: '.*:BCFTOOLS_NORM' { ext.prefix = { "${meta.id}.${meta.caller}_norm" } + // Split multi-allelics so BCFTOOLS_ISEC, which matches exact REF/ALT, can intersect them per ALT. + // BCFTOOLS_NORM_REJOIN rejoins them before phasing. ext.args = { - "-Oz" + "-m -any -Oz" } publishDir = [ enabled: false ] } + // Rejoin split multi-allelics; LongPhase keys variants by position and cannot take two records at one POS. + withName: '.*:GERMLINE_CONSENSUS:BCFTOOLS_NORM_REJOIN' { + ext.prefix = { "${meta.id}_germline_sorted" } + ext.args = { '-m +any --output-type z --write-index=tbi' } + publishDir = [ enabled: false ] + } + withName: '.*:SOMATIC_CONSENSUS:BCFTOOLS_NORM_REJOIN' { + ext.prefix = { "${meta.id}_somatic_sorted" } + ext.args = { '-m +any --output-type z --write-index=tbi' } + publishDir = [ enabled: false ] + } withName: '.*:BCFTOOLS_ISEC' { ext.prefix = { "${meta.id}_isec" } @@ -138,22 +151,41 @@ process { enabled: false ] } - withName: '.*STANDARDIZE_AF' { - ext.prefix = { "${meta.id}.${meta.caller}_standardized" } - ext.args = { - meta.rename_to == 'VAF' - ? "--rename-annots <(printf 'FORMAT/AF\\tFORMAT/VAF\\n') -Oz -W=tbi" - : "--rename-annots <(printf 'FORMAT/VAF\\tFORMAT/AF\\n') -Oz -W=tbi" + withName: '.*:BCFTOOLS_ANNOTATE' { + ext.prefix = { "${meta.id}.${meta.caller}" } + // Stamp INFO/CALLER; in 'all' mode (meta.rename_to set) also unify the FORMAT allele frequency key. + ext.args = { + def rename = meta.rename_to == 'VAF' + ? "--rename-annots <(printf 'FORMAT/AF\\tFORMAT/VAF\\n') " + : meta.rename_to == 'AF' + ? "--rename-annots <(printf 'FORMAT/VAF\\tFORMAT/AF\\n') " + : "" + rename + '''-h <(echo '##INFO=') \ + -c CHROM,POS,REF,ALT,INFO/CALLER \ + -Oz \ + -W=tbi''' } publishDir = [ enabled: false ] } - withName: '.*:BCFTOOLS_ANNOTATE' { - ext.prefix = { "${meta.id}.${meta.caller}" } + // GERMLINE VERDICT TRANSFER (tumor-only): DeepSomatic's verdict filters DeepVariant's germline calls. + withName: '.*:DS_VERDICT_QUERY' { + ext.prefix = { "${meta.id}.ds_verdict" } + // Exclude PASS/RefCall rather than include GERMLINE/PON: bcftools fails on an undeclared FILTER. ext.args = { - '''-h <(echo '##INFO=') \ - -c CHROM,POS,REF,ALT,INFO/CALLER \ + "-e 'FILTER=\"PASS\" || FILTER=\"RefCall\"' -f '%CHROM\t%POS\t%REF\t%ALT\t%FILTER\n'" + } + publishDir = [ + enabled: false + ] + } + + withName: '.*:DS_VERDICT_ANNOTATE' { + ext.prefix = { "${meta.id}.deepvariant_verdict" } + ext.args = { + '''-h <(echo '##INFO=') \ + -c CHROM,POS,REF,ALT,INFO/DS_VERDICT \ -Oz \ -W=tbi''' } @@ -161,6 +193,18 @@ process { enabled: false ] } + + withName: '.*:DS_GERMLINE_SELECT' { + ext.prefix = { "${meta.id}.deepvariant_germline" } + // Keep only DeepSomatic-adjudicated germline sites; all others lack DS_VERDICT or fail this test. + ext.args = { + "-i 'INFO/DS_VERDICT=\"GERMLINE\" || INFO/DS_VERDICT=\"PON\"' --output-type z --write-index=tbi" + } + publishDir = [ + enabled: false + ] + } + withName: '.*:BCFTOOLS_QUERY' { ext.args = { "-f '%CHROM\t%POS\t%REF\t%ALT\t${meta.caller}\n'" @@ -401,8 +445,22 @@ process { enabled: false ] } + // Provenance flags stamped before the phasing merge, so the somatic arm is selected by origin. + withName: '.*:TAG_SOMATIC' { + ext.prefix = { "${meta.id}_somatic_tagged" } + publishDir = [ + enabled: false + ] + } + withName: '.*:TAG_GERMLINE' { + ext.prefix = { "${meta.id}_germline_tagged" } + publishDir = [ + enabled: false + ] + } withName: '.*:PHASING_HAPLOTYPING:BCFTOOLS_VIEW' { ext.prefix = { "somatic_smallvariants" } + ext.args = { "-i 'INFO/SOMATIC=1'" } publishDir = [ path: { "${params.outdir}/${meta.id}/variants/phased" }, mode: params.publish_dir_mode, @@ -496,26 +554,26 @@ process { ] } withName: '.*:GERMLINE_CONSENSUS:BCFTOOLS_SORT' { - ext.prefix = { "${meta.id}_germline_sorted" } + ext.prefix = { "${meta.id}_germline_split_sorted" } ext.args = {'-Oz -W=tbi'} publishDir = [ enabled: false ] } withName: '.*:SOMATIC_CONSENSUS:BCFTOOLS_SORT' { - ext.prefix = { "${meta.id}_somatic_sorted" } + ext.prefix = { "${meta.id}_somatic_split_sorted" } ext.args = {'-Oz -W=tbi'} publishDir = [ enabled: false ] } withName: '.*:GERMLINE_CONSENSUS:BCFTOOLS_SORT_CONSENSUS' { - ext.prefix = { "${meta.id}_germline_sorted" } + ext.prefix = { "${meta.id}_germline_split_sorted" } ext.args = { '-Oz -W=tbi' } publishDir = [ enabled: false ] } withName: '.*:SOMATIC_CONSENSUS:BCFTOOLS_SORT_CONSENSUS' { - ext.prefix = { "${meta.id}_somatic_sorted" } + ext.prefix = { "${meta.id}_somatic_split_sorted" } ext.args = { '-Oz -W=tbi' } publishDir = [ enabled: false ] } @@ -575,7 +633,7 @@ process { } withName: '.*:CLAIR3' { - ext.args = { "--sample_name=${meta.id}" } + // --sample_name is passed by the module itself, from ext.prefix (which defaults to meta.id). publishDir = [ path: { "${params.outdir}/${meta.id}/variants/clair3" }, mode: params.publish_dir_mode, @@ -703,6 +761,20 @@ process { ] } + // PASS-only per-caller copies for downstream steps; ClairS-TO is covered by VCFSPLIT. + // --write-index=tbi: the module's index output is optional and the join needs it. + withName: '.*:(CLAIR3|CLAIRS|DEEPVARIANT|DEEPSOMATIC)_PASS_FILTER' { + ext.args = '--apply-filters PASS --output-type z --write-index=tbi' + publishDir = [ + enabled: false + ] + } + + withName: '.*:CLAIR3_PASS_FILTER' { ext.prefix = { "${meta.id}_clair3_pass" } } + withName: '.*:CLAIRS_PASS_FILTER' { ext.prefix = { "${meta.id}_clairs_pass" } } + withName: '.*:DEEPVARIANT_PASS_FILTER' { ext.prefix = { "${meta.id}_deepvariant_pass" } } + withName: '.*:DEEPSOMATIC_PASS_FILTER' { ext.prefix = { "${meta.id}_deepsomatic_pass" } } + withName : '.*:SIGPROFILER_MATRIXGENERATOR' { ext.args = { params.sigprofiler_matrix_args ?: '' } publishDir = [ diff --git a/docs/usage.md b/docs/usage.md index c9fe4d26..427a12a8 100644 --- a/docs/usage.md +++ b/docs/usage.md @@ -135,6 +135,16 @@ For structural variants, the CHM13 panel of normals is a merged panel combining For tumour-only small variants, ClairS-TO separates germline from somatic calls with a panel of normals and with its Verdict module, which tags each call as germline, somatic or subclonal somatic from tumour purity and allele-specific copy number. `--genome CHM13` supplies five CHM13 PON VCFs (gnomAD, dbSNP, 1000 Genomes, CoLoRSdb and ASAP), which **replace** the GRCh38 databases inside the container. Unless `--skip_ascat` is set, purity and copy number come from the pipeline's own ASCAT run (`CLAIRSTO_VERDICT_TAG`); only with `--skip_ascat` does ClairS-TO estimate them itself, from assembly-specific loci, allele and GC content files. A GRCh38 resource set on a CHM13 run leaves germline variants untagged. +When `--germline_var_keep` includes `deepvariant`, the tumour-only germline arm +runs DeepVariant on the **tumour** BAM, so on its own those calls mix germline and +clonal somatic variants. When `deepsomatic` is also in `--somatic_var_keep`, the +pipeline records DeepSomatic's verdict at each site in `INFO/DS_VERDICT` and keeps +only sites it calls `GERMLINE` or `PON`; `RefCall` and sites DeepSomatic never +evaluated are dropped. Without `deepsomatic`, the DeepVariant germline calls are +used without a verdict filter and may include somatic variants. Either way the +tumour-only germline arm is a tumour-derived proxy, not a call set from normal +tissue. + With `--genome CHM13 --skip_ascat` the pipeline builds a CHM13 resource set from the ASCAT files it already downloads, so no extra setup is needed. LogR correction is GC-only, as ClairS-TO recommends for CHM13: no replication timing file is published for the assembly. Without `--skip_ascat` nothing is built, because the tagging comes from ASCAT's own tables. To use a resource set of your own — another assembly, or a CHM13 set carrying an `RT_.txt` for replication timing correction — pass `--clairsto_cna_resources` together with `--skip_ascat`. Without `--skip_ascat` it is ignored, with a warning. Expected layout: @@ -382,14 +392,61 @@ Both tools run from `ghcr.io/ljwharbers/sigprofiler`, which adds CHM13 support n These options control how variants from multiple callers are filtered and merged. -| Parameter | Description | -| ------------------------------ | --------------------------------------------------------------------------------------------------- | -| `--germline_var_keep` | Expression or threshold for retaining germline variants after calling. Default = `null` | -| `--somatic_var_keep` | Expression or threshold for retaining somatic variants after calling. Default = `null` | -| `--germline_var_combine` | Strategy for combining germline variant caller outputs (e.g. union, intersection). Default = `null` | -| `--somatic_var_combine` | Strategy for combining somatic variant caller outputs (e.g. union, intersection). Default = `null` | -| `--prioritize_caller_germline` | Comma-separated caller priority order used when combining germline calls. Default = `null` | -| `--prioritize_caller_somatic` | Comma-separated caller priority order used when combining somatic calls. Default = `null` | +| Parameter | Description | +| ------------------------------ | ------------------------------------------------------------------------------------------------------------- | +| `--germline_var_keep` | Comma-separated germline callers to run: `deepvariant`, `clair`. Default = `clair` | +| `--somatic_var_keep` | Comma-separated somatic callers to run: `deepsomatic`, `clair`. Default = `clair` | +| `--germline_var_combine` | How to combine germline caller outputs: `consensus` (shared calls only) or `all` (union). Default = `all` | +| `--somatic_var_combine` | How to combine somatic caller outputs: `consensus` (shared calls only) or `all` (union). Default = `all` | +| `--prioritize_caller_germline` | Whose record to use for variants called by both germline callers: `deepvariant` or `clair`. Default = `clair` | +| `--prioritize_caller_somatic` | Whose record to use for variants called by both somatic callers: `deepsomatic` or `clair`. Default = `clair` | +| `--smallvar_filter_pass` | Keep only PASS records from each small variant caller downstream. Default = `true` | + +DeepVariant and DeepSomatic emit a record for every site they evaluate, not only +for the variants they call, so most of their records are `RefCall`, `GERMLINE` or +`PON`; Clair3 and ClairS also keep their `LowQual` and `NonSomatic` records. Without +a `PASS` filter, `*_var_combine = 'all'` would carry all of these into the phased +VCFs and the mutation burden. + +`--smallvar_filter_pass` (`true` by default) restricts the copy of each caller's +VCF that is handed to the caller consensus, phasing, VEP and the report. In +tumor-only mode ClairS-TO is unaffected by the setting: `VCFSPLIT` already +restricts its somatic split to `PASS`, and its germline split is `PASS`-rewritten +rather than `PASS`-filtered. The per-caller VCFs published under +`//variants/` are never filtered, so no calls are lost +from the results directory. + +Set it to `false` to restore the previous unfiltered behaviour: each caller's +records are passed on with their original `FILTER`. Only the ClairS-TO germline +split is normalised to `PASS`, with its original value kept in +`INFO/ORIG_FILTER`. + +`consensus` keeps only variants called by both callers; `all` keeps the union, i.e. +every variant called by either. In both modes `--prioritize_caller_*` chooses only +whose record represents a variant that both callers found -- it never decides which +variants are kept. Multi-allelic records are split so they can be matched across +callers, and rejoined before phasing. + +#### Germline and somatic provenance + +Germline and somatic small variants are merged into one VCF for somatic phasing, +because Longphase needs all variant sites in a single file to produce consistent +phase blocks. The somatic arm is then recovered from the phased result by an +`INFO/SOMATIC` flag stamped on each arm before the merge, not by position, since a +germline record at the same coordinate as a somatic call would otherwise be kept. +Tagging leaves `FILTER` unchanged. + +Three INFO fields carry this provenance: + +| Field | Meaning | +| ------------- | --------------------------------------------------------------------------------------------------------------------- | +| `SOMATIC` | Record came from the somatic call set | +| `GERMLINE` | Record came from the germline call set | +| `ORIG_FILTER` | Original `FILTER` of ClairS-TO germline records, before normalisation to `PASS`. Multiple filters are joined with `,` | + +Germline calls dropped from `variants/phased/somatic_smallvariants.vcf.gz` are not +lost: they remain in `variants/phased/germline_smallvariants.vcf.gz`, in +`vep/germline/`, and in the unfiltered per-caller VCFs under `variants//`. #### PON Options diff --git a/modules/local/bcftools/view/main.nf b/modules/local/bcftools/view/main.nf index 652da9ac..63905dd2 100644 --- a/modules/local/bcftools/view/main.nf +++ b/modules/local/bcftools/view/main.nf @@ -8,7 +8,7 @@ process BCFTOOLS_VIEW { : 'community.wave.seqera.io/library/bcftools_htslib:0a3fa2654b52006f'}" input: - tuple val(meta), path(vcf), path(tbi), path(targets), path(targets_tbi) + tuple val(meta), path(vcf), path(tbi) output: tuple val(meta), path("*.vcf.gz"), emit: vcf @@ -23,7 +23,6 @@ process BCFTOOLS_VIEW { def prefix = task.ext.prefix ?: "${meta.id}" """ bcftools view \\ - -T ${targets} \\ -Oz \\ -W=tbi \\ ${args} \\ diff --git a/modules/local/bcftools/view/meta.yml b/modules/local/bcftools/view/meta.yml index 28c0fc48..29e3646b 100644 --- a/modules/local/bcftools/view/meta.yml +++ b/modules/local/bcftools/view/meta.yml @@ -1,5 +1,5 @@ name: bcftools_view -description: Filter VCF to positions defined by a targets file using bcftools view -T +description: Filter a VCF with bcftools view, with the filter expression supplied via ext.args; outputs a bgzipped VCF and tbi index keywords: - filtering - VCF @@ -25,14 +25,6 @@ input: type: file description: Tabix index of the input VCF pattern: "*.tbi" - - targets: - type: file - description: VCF file used as position filter (-T) - pattern: "*.{vcf.gz,vcf,bcf}" - - targets_tbi: - type: file - description: Tabix index of the targets VCF - pattern: "*.tbi" output: vcf: - - meta: @@ -50,7 +42,28 @@ output: type: file description: Tabix index of filtered VCF pattern: "*.tbi" + versions_bcftools: + - - ${task.process}: + type: string + description: The process the versions were collected from + - bcftools: + type: string + description: The tool name + - "bcftools --version | sed '1!d; s/^.*bcftools //'": + type: string + description: The command used to generate the version of the tool +topics: + versions: + - - ${task.process}: + type: string + description: The process the versions were collected from + - bcftools: + type: string + description: The tool name + - "bcftools --version | sed '1!d; s/^.*bcftools //'": + type: string + description: The command used to generate the version of the tool authors: - - "@rforsyth" + - "@robert-a-forsyth" maintainers: - - "@rforsyth" + - "@robert-a-forsyth" diff --git a/modules/local/vcfsplit/main.nf b/modules/local/vcfsplit/main.nf index f6156d34..55cb6300 100644 --- a/modules/local/vcfsplit/main.nf +++ b/modules/local/vcfsplit/main.nf @@ -38,10 +38,21 @@ process VCFSPLIT { bcftools concat -a -Oz -o germline_tmp.vcf.gz indels_filtered.vcf.gz snv_filtered.vcf.gz tabix -p vcf germline_tmp.vcf.gz - bcftools view germline_tmp.vcf.gz | awk 'BEGIN{FS=OFS="\t"} /^#/ {print} !/^#/ { \$7="PASS"; print }' | \ - bgzip -c > germline.vcf.gz + # Normalise FILTER to PASS, keeping the original in INFO/ORIG_FILTER (";" stored as ","). + bcftools view germline_tmp.vcf.gz | awk -v q='"' 'BEGIN{FS=OFS="\t"} + /^##/ { print; next } + /^#CHROM/ { print "##INFO="; print; next } + { of = \$7; gsub(/;/, ",", of) + \$8 = (\$8 == "." || \$8 == "") ? "ORIG_FILTER=" of : \$8 ";ORIG_FILTER=" of + \$7 = "PASS" + print } + ' | bgzip -c > germline.vcf.gz tabix -p vcf germline.vcf.gz + # Fail here, not downstream, if either header does not parse. + bcftools view -h somatic.vcf.gz > /dev/null + bcftools view -h germline.vcf.gz > /dev/null + # Cleanup intermediate files rm indels_pass.vcf.gz snv_pass.vcf.gz rm indels_pass.vcf.gz.tbi snv_pass.vcf.gz.tbi diff --git a/modules/local/vcftag/environment.yml b/modules/local/vcftag/environment.yml new file mode 100644 index 00000000..b276efd9 --- /dev/null +++ b/modules/local/vcftag/environment.yml @@ -0,0 +1,7 @@ +--- +# yaml-language-server: $schema=https://raw.githubusercontent.com/nf-core/modules/master/modules/environment-schema.json +channels: + - conda-forge + - bioconda +dependencies: + - bioconda::bcftools=1.20 diff --git a/modules/local/vcftag/main.nf b/modules/local/vcftag/main.nf new file mode 100644 index 00000000..390122c6 --- /dev/null +++ b/modules/local/vcftag/main.nf @@ -0,0 +1,53 @@ +process VCFTAG { + tag "$meta.id" + label 'process_single' + + conda "${moduleDir}/environment.yml" + container "${ workflow.containerEngine == 'singularity' && !task.ext.singularity_pull_docker_container ? + 'https://depot.galaxyproject.org/singularity/bcftools:1.20--h8b25389_0': + 'biocontainers/bcftools:1.20--h8b25389_0' }" + + input: + tuple val(meta), path(vcf), path(tbi) + val flag + + output: + tuple val(meta), path("${prefix}.vcf.gz") , emit: vcf + tuple val(meta), path("${prefix}.vcf.gz.tbi") , emit: tbi + + tuple val("${task.process}"), val('bcftools'), eval("bcftools --version |& sed '1!d ; s/bcftools //'"), topic: versions, emit: versions_bcftools + + when: + task.ext.when == null || task.ext.when + + script: + prefix = task.ext.prefix ?: "${meta.id}_${flag.toLowerCase()}" + """ + # Stamp a constant INFO flag marking the call set; FILTER is left as the caller emitted it. + # bcftools annotate cannot set a constant INFO field without an annotation file, hence awk. + bcftools view ${vcf} | awk -v flag="${flag}" -v q='"' 'BEGIN{FS=OFS="\t"} + /^##/ { print; next } + /^#CHROM/ { + print "##INFO=" + print + next + } + { + \$8 = (\$8 == "." || \$8 == "") ? flag : \$8 ";" flag + print + } + ' | bgzip -c > ${prefix}.vcf.gz + + # Fail here, not downstream, if the header does not parse. + bcftools view -h ${prefix}.vcf.gz > /dev/null + + tabix -p vcf ${prefix}.vcf.gz + """ + + stub: + prefix = task.ext.prefix ?: "${meta.id}_${flag.toLowerCase()}" + """ + echo "" | gzip > ${prefix}.vcf.gz + touch ${prefix}.vcf.gz.tbi + """ +} diff --git a/modules/local/vcftag/meta.yml b/modules/local/vcftag/meta.yml new file mode 100644 index 00000000..648ca689 --- /dev/null +++ b/modules/local/vcftag/meta.yml @@ -0,0 +1,75 @@ +name: vcftag +description: Stamp a constant INFO flag, named by the flag input, on every record of a VCF; FILTER is left untouched +keywords: + - vcf + - annotation + - INFO + - variant calling +tools: + - bcftools: + description: Tools for variant calling and manipulating VCFs and BCFs + homepage: http://samtools.github.io/bcftools/bcftools.html + documentation: http://www.htslib.org/doc/bcftools.html + tool_dev_url: https://github.com/samtools/bcftools + doi: "10.1093/gigascience/giab008" + licence: ["MIT"] + identifier: biotools:bcftools +input: + - - meta: + type: map + description: Groovy Map containing sample information e.g. [ id:'test' ] + - vcf: + type: file + description: Input VCF file to tag + pattern: "*.vcf.gz" + - tbi: + type: file + description: Tabix index of the input VCF + pattern: "*.tbi" + - flag: + type: string + description: | + Name of the INFO flag added to every record (e.g. `GERMLINE`); + its lowercase form is also used in the default output prefix +output: + vcf: + - - meta: + type: map + description: Groovy Map containing sample information + - ${prefix}.vcf.gz: + type: file + description: VCF with the INFO flag added to every record + pattern: "*.vcf.gz" + tbi: + - - meta: + type: map + description: Groovy Map containing sample information + - ${prefix}.vcf.gz.tbi: + type: file + description: Tabix index of the tagged VCF + pattern: "*.vcf.gz.tbi" + versions_bcftools: + - - ${task.process}: + type: string + description: The process the versions were collected from + - bcftools: + type: string + description: The tool name + - "bcftools --version |& sed '1!d ; s/bcftools //'": + type: string + description: The command used to generate the version of the tool +topics: + versions: + - - ${task.process}: + type: string + description: The process the versions were collected from + - bcftools: + type: string + description: The tool name + - "bcftools --version |& sed '1!d ; s/bcftools //'": + type: string + description: The command used to generate the version of the tool +authors: + - "@robert-a-forsyth" +maintainers: + - "@robert-a-forsyth" diff --git a/modules/local/vcftag/tests/main.nf.test b/modules/local/vcftag/tests/main.nf.test new file mode 100644 index 00000000..c0ec94a6 --- /dev/null +++ b/modules/local/vcftag/tests/main.nf.test @@ -0,0 +1,76 @@ +nextflow_process { + + name "Test Process VCFTAG" + script "../main.nf" + process "VCFTAG" + + tag "modules" + tag "modules_local" + tag "vcftag" + + // Runs for real (no -stub): the tagging is an awk program embedded in the Nextflow script + // block, so the escaping only holds if it is actually executed. + test("stamps the flag and declares its header without touching FILTER") { + + when { + process { + """ + input[0] = [ + [ id:'test' ], + file("\${projectDir}/tests/fixtures/vcftag_input.vcf.gz", checkIfExists: true), + file("\${projectDir}/tests/fixtures/vcftag_input.vcf.gz.tbi", checkIfExists: true) + ] + input[1] = 'SOMATIC' + """ + } + } + + then { + assert process.success + + // linesGzip is built into nf-test; the .vcf accessor needs an nft-vcf plugin that + // nf-test.config does not load, so this reads the records directly. + def all = path(process.out.vcf[0][1]).linesGzip + def lines = all.findAll { line -> !line.startsWith('#') } + def header = all.findAll { line -> line.startsWith('##') }.join('\n') + + def filters = lines.collectEntries { l -> def f = l.split('\t'); [ (f[1]): f[6] ] } + def infos = lines.collectEntries { l -> def f = l.split('\t'); [ (f[1]): f[7] ] } + + assertAll( + // every input record survives -- tagging must not filter + { assert lines.size() == 5 }, + // the flag is declared, so bcftools can query it downstream + { assert header.contains('ID=SOMATIC') }, + // the flag is stamped exactly once per record, including the one whose INFO was '.' + { assert infos.values().every { it.split(';').count('SOMATIC') == 1 } }, + { assert infos['300'] == 'SOMATIC' }, + // FILTER is left exactly as the caller emitted it + { assert filters == [ '100':'PASS', '200':'NonSomatic', '300':'RefCall', '400':'.', '500':'LowQual;RefCall' ] }, + // pre-existing INFO is preserved, with the flag appended + { assert infos['200'] == 'CALLER=clairs-to;EXISTING;SOMATIC' }, + { assert infos['500'] == 'CALLER=clair3;EXISTING;SOMATIC' }, + // VCFTAG no longer adds ORIG_FILTER + { assert !all.any { it.contains('ORIG_FILTER') } } + ) + } + } + + test("stub") { + + options "-stub" + + when { + process { + """ + input[0] = [ [ id:'test' ], [], [] ] + input[1] = 'GERMLINE' + """ + } + } + + then { + assert process.success + } + } +} diff --git a/modules/local/wakhan/environment.yml b/modules/local/wakhan/environment.yml index 6b0cdb43..33c3c873 100644 --- a/modules/local/wakhan/environment.yml +++ b/modules/local/wakhan/environment.yml @@ -4,4 +4,4 @@ channels: - conda-forge - bioconda dependencies: - - "bioconda::wakhan=0.4.3" + - "bioconda::wakhan=0.4.4" diff --git a/modules/local/wakhan/main.nf b/modules/local/wakhan/main.nf index ad7aba5e..517ae995 100644 --- a/modules/local/wakhan/main.nf +++ b/modules/local/wakhan/main.nf @@ -4,8 +4,8 @@ process WAKHAN { conda "${moduleDir}/environment.yml" container "${ workflow.containerEngine == 'singularity' && !task.ext.singularity_pull_docker_container ? - 'https://depot.galaxyproject.org/singularity/wakhan:0.4.3--pyhdfd78af_0': - 'biocontainers/wakhan:0.4.3--pyhdfd78af_0' }" + 'https://depot.galaxyproject.org/singularity/wakhan:0.4.4--pyhdfd78af_0': + 'biocontainers/wakhan:0.4.4--pyhdfd78af_0' }" input: tuple val(meta), path(tumor_input), path(tumor_index), path(normal_input), path(normal_index), path(vcf), path(breakpoints) @@ -38,9 +38,9 @@ process WAKHAN { tuple val(meta), path("solutions_ranks.tsv") , emit: solutions_ranks // Whole directories, not the plots inside: every solution's plot has the same basename, // and LRSOMATICREPORT resolves them by solution_/ path - tuple val(meta), path("solution_*", type: 'dir') , emit: solution_dirs, optional: true + tuple val(meta), path("solution_*", type: 'dir') , emit: solution_dirs // WARN: Manually update version information as tool does not provide on CLI - tuple val("${task.process}"), val('wakhan'), val("0.4.3"), topic: versions, emit: versions_wakhan + tuple val("${task.process}"), val('wakhan'), val("0.4.4"), topic: versions, emit: versions_wakhan when: task.ext.when == null || task.ext.when diff --git a/nextflow.config b/nextflow.config index 696d5161..4230dd67 100644 --- a/nextflow.config +++ b/nextflow.config @@ -20,6 +20,7 @@ params { somatic_var_combine = 'all' prioritize_caller_germline = 'clair' prioritize_caller_somatic = 'clair' + smallvar_filter_pass = true generate_gvcf = false // Longphase options diff --git a/nextflow_schema.json b/nextflow_schema.json index 713e2baa..a4efeb21 100644 --- a/nextflow_schema.json +++ b/nextflow_schema.json @@ -101,6 +101,13 @@ "default": "clair", "enum": ["deepsomatic", "clair"] }, + "smallvar_filter_pass": { + "type": "boolean", + "default": true, + "description": "Keep only PASS records from each small variant caller for downstream use.", + "help_text": "DeepVariant and DeepSomatic emit a record for every site they evaluate, so most records are RefCall (or GERMLINE/PON) rather than calls, and Clair3/ClairS keep their LowQual and NonSomatic calls. Those records otherwise flow into the caller consensus, phasing, VEP and the report. Set to false to consume each caller's unfiltered output. The per-caller VCFs published under `//variants/` are unaffected either way.", + "fa_icon": "fas fa-filter" + }, "generate_gvcf": { "type": "boolean" } diff --git a/subworkflows/local/paired/paired_smallvar_germline.nf b/subworkflows/local/paired/paired_smallvar_germline.nf index 3b473006..3164e33b 100644 --- a/subworkflows/local/paired/paired_smallvar_germline.nf +++ b/subworkflows/local/paired/paired_smallvar_germline.nf @@ -1,4 +1,6 @@ // IMPORT MODULES +include { BCFTOOLS_VIEW as CLAIR3_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' +include { BCFTOOLS_VIEW as DEEPVARIANT_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' include { CLAIR3 } from '../../../modules/local/clair3/main.nf' // IMPORT SUBWORKFLOWS @@ -73,8 +75,15 @@ workflow PAIRED_SMALLVAR_GERMLINE { fai ) - CLAIR3.out.vcf - .join(CLAIR3.out.tbi) + // PASS-only copy for downstream steps; published VCFs are untouched. + def clair3_vcf = CLAIR3.out.vcf.join(CLAIR3.out.tbi) + if (params.smallvar_filter_pass) { + CLAIR3_PASS_FILTER ( clair3_vcf, [], [], [] ) + clair3_vcf = CLAIR3_PASS_FILTER.out.vcf + .join(CLAIR3_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + clair3_vcf .map { meta, vcf , tbi -> def new_meta = meta + [caller:'clair3'] return [new_meta, vcf, tbi] @@ -119,8 +128,15 @@ workflow PAIRED_SMALLVAR_GERMLINE { [[:],[]] // GFF annotation (not used) ) - DEEPVARIANT.out.vcf - .join(DEEPVARIANT.out.vcf_index) + // PASS-only copy for downstream steps; published VCFs are untouched. + def deepvariant_vcf = DEEPVARIANT.out.vcf.join(DEEPVARIANT.out.vcf_index) + if (params.smallvar_filter_pass) { + DEEPVARIANT_PASS_FILTER ( deepvariant_vcf, [], [], [] ) + deepvariant_vcf = DEEPVARIANT_PASS_FILTER.out.vcf + .join(DEEPVARIANT_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + deepvariant_vcf .map{ meta, vcf, tbi -> def new_meta = meta + [caller:'deepvariant'] return [new_meta, vcf, tbi] diff --git a/subworkflows/local/paired/paired_smallvar_somatic.nf b/subworkflows/local/paired/paired_smallvar_somatic.nf index cf5a749d..ab296b99 100644 --- a/subworkflows/local/paired/paired_smallvar_somatic.nf +++ b/subworkflows/local/paired/paired_smallvar_somatic.nf @@ -1,4 +1,6 @@ // IMPORT MODULES +include { BCFTOOLS_VIEW as CLAIRS_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' +include { BCFTOOLS_VIEW as DEEPSOMATIC_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' include { CLAIRS } from '../../../modules/local/clairs/main.nf' include { BCFTOOLS_CONCAT } from '../../../modules/nf-core/bcftools/concat' include { BCFTOOLS_SORT } from '../../../modules/nf-core/bcftools/sort' @@ -71,8 +73,15 @@ workflow PAIRED_SMALLVAR_SOMATIC { BCFTOOLS_CONCAT.out.vcf ) - BCFTOOLS_SORT.out.vcf - .join(BCFTOOLS_SORT.out.tbi) + // PASS-only copy for downstream steps; published VCFs are untouched. + def clairs_vcf = BCFTOOLS_SORT.out.vcf.join(BCFTOOLS_SORT.out.tbi) + if (params.smallvar_filter_pass) { + CLAIRS_PASS_FILTER ( clairs_vcf, [], [], [] ) + clairs_vcf = CLAIRS_PASS_FILTER.out.vcf + .join(CLAIRS_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + clairs_vcf .map { meta, vcf , tbi -> def new_meta = meta + [caller:'clairs'] return [new_meta, vcf, tbi] @@ -109,8 +118,15 @@ workflow PAIRED_SMALLVAR_SOMATIC { ds_pon_channel ) - DEEPSOMATIC.out.vcf - .join(DEEPSOMATIC.out.vcf_index) + // PASS-only copy for downstream steps; published VCFs are untouched. + def deepsomatic_vcf = DEEPSOMATIC.out.vcf.join(DEEPSOMATIC.out.vcf_index) + if (params.smallvar_filter_pass) { + DEEPSOMATIC_PASS_FILTER ( deepsomatic_vcf, [], [], [] ) + deepsomatic_vcf = DEEPSOMATIC_PASS_FILTER.out.vcf + .join(DEEPSOMATIC_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + deepsomatic_vcf .map{ meta, vcf, tbi -> def new_meta = meta + [caller:'deepsomatic'] return [new_meta, vcf, tbi] diff --git a/subworkflows/local/phasing_haplotyping.nf b/subworkflows/local/phasing_haplotyping.nf index e80561c8..0f101586 100644 --- a/subworkflows/local/phasing_haplotyping.nf +++ b/subworkflows/local/phasing_haplotyping.nf @@ -8,6 +8,8 @@ include { SAMTOOLS_INDEX } from '../../module include { BCFTOOLS_CONCAT } from '../../modules/nf-core/bcftools/concat/main' include { BCFTOOLS_SORT } from '../../modules/nf-core/bcftools/sort/main' include { BCFTOOLS_VIEW } from '../../modules/local/bcftools/view/main.nf' +include { VCFTAG as TAG_SOMATIC } from '../../modules/local/vcftag/main.nf' +include { VCFTAG as TAG_GERMLINE } from '../../modules/local/vcftag/main.nf' workflow PHASING_HAPLOTYPING { @@ -138,11 +140,27 @@ workflow PHASING_HAPLOTYPING { } + // + // MODULE: VCFTAG (label: process_single), aliased TAG_SOMATIC / TAG_GERMLINE + // Stamp each arm with an INFO provenance flag before the merge; LongPhase keeps it through phasing. + // + TAG_SOMATIC ( somatic_vcf, 'SOMATIC' ) + TAG_GERMLINE( germline_vcf, 'GERMLINE' ) + + TAG_SOMATIC.out.vcf + .join(TAG_SOMATIC.out.tbi, failOnMismatch: true, failOnDuplicate: true) + .set{ tagged_somatic_vcf } + TAG_GERMLINE.out.vcf + .join(TAG_GERMLINE.out.tbi, failOnMismatch: true, failOnDuplicate: true) + .set{ tagged_germline_vcf } + // tagged_*_vcf: [meta, vcf, tbi] + // Somatic phasing needs germline and somatic sites in one VCF for consistent phase blocks - germline_vcf - .join(somatic_vcf) + tagged_germline_vcf + .join(tagged_somatic_vcf) .map { meta, germ_vcf, germ_tbi, som_vcf, som_tbi -> - def vcfs = [som_vcf, germ_vcf] // somatic first (higher priority in phasing) + // Order is cosmetic: BCFTOOLS_CONCAT sorts its inputs by name. + def vcfs = [som_vcf, germ_vcf] def tbis = [som_tbi, germ_tbi] return [ meta, vcfs, tbis] } @@ -174,7 +192,7 @@ workflow PHASING_HAPLOTYPING { if (!params.skip_modcall) { // With modcall: include base-modification VCF as additional phasing evidence normal_bams_w_tumoronly_ch - .join(germline_vcf) + .join(tagged_germline_vcf) .join(LONGPHASE_MODCALL_GERMLINE.out.mod_vcf) .map { meta, bam, bai, vcf, _tbi, mods-> def svs = [] // SVs for phasing are not used here @@ -196,7 +214,7 @@ workflow PHASING_HAPLOTYPING { else { // Without modcall: empty lists for SVs and mods normal_bams_w_tumoronly_ch - .join(germline_vcf) + .join(tagged_germline_vcf) .map { meta, bam, bai, vcf, _tbi -> def svs = [] def mods = [] @@ -253,20 +271,13 @@ workflow PHASING_HAPLOTYPING { // phased_somatic_germline_vcf: [meta, vcf, tbi] -- Longphase-phased somatic+germline VCF (unfiltered) // - // MODULE: BCFTOOLS_VIEW (label: process_medium) -- back to somatic-only, with the original somatic VCF as -T targets; PS/HP tags survive - // Input: [meta, phased_combined_vcf, phased_combined_tbi, somatic_vcf, somatic_tbi] + // MODULE: BCFTOOLS_VIEW (label: process_medium) + // Keep the somatic arm by its INFO/SOMATIC flag, not by position; PS/HP tags survive. + // Input: [meta, phased_combined_vcf, phased_combined_tbi] // Output: .vcf -- [meta, vcf.gz] -- phased somatic-only VCF // .tbi -- [meta, tbi] // - phased_somatic_germline_vcf - .join(somatic_vcf) - .map { meta, phased_vcf, phased_tbi, som_vcf, som_tbi -> - return [ meta, phased_vcf, phased_tbi, som_vcf, som_tbi ] - } - .set { bcftools_view_input_ch } - // bcftools_view_input_ch: [meta, phased_combined_vcf, tbi, somatic_vcf, somatic_tbi] - - BCFTOOLS_VIEW ( bcftools_view_input_ch ) + BCFTOOLS_VIEW ( phased_somatic_germline_vcf ) BCFTOOLS_VIEW.out.vcf .join(BCFTOOLS_VIEW.out.tbi) diff --git a/subworkflows/local/small_variant_consensus.nf b/subworkflows/local/small_variant_consensus.nf index 7e541aaa..6444a492 100644 --- a/subworkflows/local/small_variant_consensus.nf +++ b/subworkflows/local/small_variant_consensus.nf @@ -1,8 +1,8 @@ include { BCFTOOLS_NORM } from '../../modules/nf-core/bcftools/norm/main' +include { BCFTOOLS_NORM as BCFTOOLS_NORM_REJOIN } from '../../modules/nf-core/bcftools/norm/main' include { BCFTOOLS_ISEC } from '../../modules/nf-core/bcftools/isec/main' include { BCFTOOLS_QUERY } from '../../modules/nf-core/bcftools/query/main' include { BCFTOOLS_ANNOTATE } from '../../modules/nf-core/bcftools/annotate/main' -include { BCFTOOLS_ANNOTATE as STANDARDIZE_AF } from '../../modules/nf-core/bcftools/annotate/main' include { BCFTOOLS_CONCAT } from '../../modules/nf-core/bcftools/concat/main' include { BCFTOOLS_SORT } from '../../modules/nf-core/bcftools/sort/main' include { BCFTOOLS_SORT as SORT_POST_NORM } from '../../modules/nf-core/bcftools/sort/main' @@ -17,12 +17,12 @@ workflow SMALL_VARIANT_CONSENSUS { fasta // [[:], fasta] _fai // [[:], fai] prioritize_caller // str: which caller's calls take priority ('deepvariant'/'deepsomatic' or 'clair') - combine_method // str: 'consensus' (intersection only) or 'all' (intersection + private calls from priority caller) + combine_method // str: 'consensus' (shared calls only) or 'all' (union of both callers' calls) main: // - // MODULE: BCFTOOLS_NORM (label: process_medium) -- left-align and normalise; sorted after, since left-alignment can reorder records + // MODULE: BCFTOOLS_NORM (label: process_medium) -- left-align and split multi-allelics for isec; rejoined before phasing // Input: [meta, vcf, tbi] -- per-caller VCF // Output: .vcf -- [meta, vcf] -- left-aligned, normalised VCF (unsorted) // @@ -42,31 +42,10 @@ workflow SMALL_VARIANT_CONSENSUS { // normalized_vcfs: [meta(+caller), vcf.gz, tbi] -- normalised, sorted per-caller VCF // - // MODULE: STANDARDIZE_AF (BCFTOOLS_ANNOTATE alias, label: process_low) -- rename the AF FORMAT field to the priority caller's: + // ALLELE FREQUENCY KEY -- BCFTOOLS_ANNOTATE below renames the AF FORMAT field to the priority caller's: // FORMAT/AF -> FORMAT/VAF when prioritize_caller is 'deepvariant'/'deepsomatic' // FORMAT/VAF -> FORMAT/AF when prioritize_caller is 'clair' - // - if (combine_method == 'all') { - normalized_vcfs - .map { meta, vcf, tbi -> - def rename_to = prioritize_caller in ['deepvariant', 'deepsomatic'] ? 'VAF' : 'AF' - def new_meta = meta + [rename_to: rename_to] - return [new_meta, vcf, tbi, [], [], [], [], []] - } - .set { standardize_input } - - STANDARDIZE_AF(standardize_input) - - STANDARDIZE_AF.out.vcf - .join(STANDARDIZE_AF.out.tbi) - .map { meta, vcf, tbi -> - def clean_meta = meta.findAll { k, _v -> k != 'rename_to' } - return [clean_meta, vcf, tbi] - } - .set { normalized_vcfs } - // normalized_vcfs: [meta(+caller), vcf, tbi] -- normalised, AF-standardized per-caller VCF - } - // In 'consensus' mode, normalized_vcfs comes from SORT_POST_NORM (post-BCFTOOLS_NORM re-sorting) + // Only 'all' mode renames: it merges both callers, so the merged VCF needs one AF key for WAKHAN. // // MODULE: BCFTOOLS_QUERY (label: process_single) @@ -79,13 +58,17 @@ workflow SMALL_VARIANT_CONSENSUS { // Prepare BCFTOOLS_ANNOTATE input: VCF + caller-name annotation file normalized_vcfs - .join(BCFTOOLS_QUERY.out.output) - .join(BCFTOOLS_QUERY.out.index) + .join(BCFTOOLS_QUERY.out.output, failOnMismatch: true, failOnDuplicate: true) + .join(BCFTOOLS_QUERY.out.index, failOnMismatch: true, failOnDuplicate: true) .map{ meta, vcf, tbi, annotations, annotations_index -> def columns = [] // no extra column specs def header_lines = [] // no extra header lines def rename_chrs = [] // no chromosome renaming - return [ meta, vcf, tbi, annotations, annotations_index, columns, header_lines, rename_chrs ] + // 'all' mode merges both callers, so unify the AF key; 'consensus' needs no rename. + def new_meta = combine_method == 'all' + ? meta + [rename_to: (prioritize_caller in ['deepvariant', 'deepsomatic'] ? 'VAF' : 'AF')] + : meta + return [ new_meta, vcf, tbi, annotations, annotations_index, columns, header_lines, rename_chrs ] } .set{annotate_input} // annotate_input: [meta, vcf, tbi, annotations_tsv, annotations_tbi, [], [], []] @@ -100,17 +83,28 @@ workflow SMALL_VARIANT_CONSENSUS { BCFTOOLS_ANNOTATE(annotate_input) BCFTOOLS_ANNOTATE.out.vcf - .join(BCFTOOLS_ANNOTATE.out.tbi) + .join(BCFTOOLS_ANNOTATE.out.tbi, failOnMismatch: true, failOnDuplicate: true) + .map { meta, vcf, tbi -> + def clean_meta = meta.findAll { k, _v -> k != 'rename_to' } + return [clean_meta, vcf, tbi] + } .set{annotated_vcfs} // annotated_vcfs: [meta(+caller), vcf, tbi] -- VCF with CALLER INFO tag // Branch annotated VCFs by caller family for the intersection step + // `other` errors on an unrecognised meta.caller instead of silently dropping the sample. annotated_vcfs .branch { meta, _vcfs, _tbi -> deepvariant: meta.caller in [ 'deepvariant', 'deepsomatic' ] clair: meta.caller in ['clair3','clairs-to','clairs'] + other: true } .set{annotated_vcfs_branched} + + annotated_vcfs_branched.other + .map { meta, _vcfs, _tbi -> + error("SMALL_VARIANT_CONSENSUS: unrecognised meta.caller '${meta.caller}' for sample '${meta.id}'; expected one of [deepvariant, deepsomatic, clair3, clairs-to, clairs]") + } // annotated_vcfs_branched.deepvariant: [meta(caller=deepvariant/deepsomatic), vcf, tbi] // annotated_vcfs_branched.clair: [meta(caller=clair3/clairs-to/clairs), vcf, tbi] @@ -153,8 +147,9 @@ workflow SMALL_VARIANT_CONSENSUS { // deepvariant_ch: [meta (no caller), vcf, tbi] // Join DeepVariant and Clair VCFs per sample into a single tuple for BCFTOOLS_ISEC + // failOnMismatch: a sample missing one caller would otherwise be dropped silently. deepvariant_ch - .join(clair_ch) + .join(clair_ch, failOnMismatch: true, failOnDuplicate: true) .map { meta, deepvar_vcf, deepvar_tbi, clair_vcf, clair_tbi -> def vcfs = [deepvar_vcf, clair_vcf] def tbis = [deepvar_tbi, clair_tbi] @@ -193,6 +188,9 @@ workflow SMALL_VARIANT_CONSENSUS { BCFTOOLS_ISEC.out.clair_consensus_vcf .set{isec_consensus_vcf} } + else { + error("prioritize_caller must be one of [deepvariant, deepsomatic, clair], got '${prioritize_caller}'") + } // ISEC always writes 0002.vcf.gz, so germline and somatic would collide by basename in BCFTOOLS_CONCAT; // BCFTOOLS_SORT_CONSENSUS renames it per sample (conf/modules.config) BCFTOOLS_SORT_CONSENSUS(isec_consensus_vcf) @@ -202,48 +200,75 @@ workflow SMALL_VARIANT_CONSENSUS { } else if (combine_method == 'all') { - // Intersection plus the priority caller's private calls + // Union: shared calls (prioritized caller's record) plus both callers' private calls. + // The three isec sets are disjoint, so BCFTOOLS_CONCAT needs no -d. if (prioritize_caller in ['deepvariant', 'deepsomatic']) { - // consensus (DeepVariant record) + DeepVariant-private variants + // shared (DeepVariant record) + DeepVariant-private + Clair-private BCFTOOLS_ISEC.out.deepvar_consensus_vcf .join(BCFTOOLS_ISEC.out.deepvar_consensus_tbi) + .join(BCFTOOLS_ISEC.out.deepvar_private_vcf) + .join(BCFTOOLS_ISEC.out.deepvar_private_tbi) .join(BCFTOOLS_ISEC.out.clair_private_vcf) .join(BCFTOOLS_ISEC.out.clair_private_tbi) - .map{ meta, deepvar_vcf, deepvar_tbi, clair_vcf, clair_tbi -> - return[meta, [deepvar_vcf, clair_vcf], [deepvar_tbi, clair_tbi]] + .map{ meta, shared_vcf, shared_tbi, deepvar_vcf, deepvar_tbi, clair_vcf, clair_tbi -> + return[meta, [shared_vcf, deepvar_vcf, clair_vcf], [shared_tbi, deepvar_tbi, clair_tbi]] } .set{concat_input} - // concat_input: [meta, [consensus_vcf, private_vcf], [consensus_tbi, private_tbi]] - BCFTOOLS_CONCAT(concat_input) - BCFTOOLS_CONCAT.out.vcf - .set{concat_out} } else if (prioritize_caller == 'clair') { - // consensus (Clair record) + Clair-private variants - BCFTOOLS_ISEC.out.deepvar_private_vcf - .join(BCFTOOLS_ISEC.out.deepvar_private_tbi) - .join(BCFTOOLS_ISEC.out.clair_consensus_vcf) + // shared (Clair record) + DeepVariant-private + Clair-private + BCFTOOLS_ISEC.out.clair_consensus_vcf .join(BCFTOOLS_ISEC.out.clair_consensus_tbi) - .map{ meta, deepvar_vcf, deepvar_tbi, clair_vcf, clair_tbi -> - return[meta, [deepvar_vcf, clair_vcf], [deepvar_tbi, clair_tbi]] + .join(BCFTOOLS_ISEC.out.deepvar_private_vcf) + .join(BCFTOOLS_ISEC.out.deepvar_private_tbi) + .join(BCFTOOLS_ISEC.out.clair_private_vcf) + .join(BCFTOOLS_ISEC.out.clair_private_tbi) + .map{ meta, shared_vcf, shared_tbi, deepvar_vcf, deepvar_tbi, clair_vcf, clair_tbi -> + return[meta, [shared_vcf, deepvar_vcf, clair_vcf], [shared_tbi, deepvar_tbi, clair_tbi]] } .set{concat_input} - // concat_input: [meta, [private_vcf, consensus_vcf], [private_tbi, consensus_tbi]] - BCFTOOLS_CONCAT(concat_input) - BCFTOOLS_CONCAT.out.vcf - .set{concat_out} } - // concat_out: [meta, vcf] -- unsorted concatenated VCF (consensus + priority-caller-private) + else { + error("prioritize_caller must be one of [deepvariant, deepsomatic, clair], got '${prioritize_caller}'") + } + // concat_input: [meta, [shared_vcf, deepvar_private_vcf, clair_private_vcf], [tbis...]] + BCFTOOLS_CONCAT(concat_input) + BCFTOOLS_CONCAT.out.vcf + .set{concat_out} + // concat_out: [meta, vcf] -- unsorted union of both callers' calls BCFTOOLS_SORT(concat_out) BCFTOOLS_SORT.out.vcf .set{vcf} BCFTOOLS_SORT.out.tbi .set{tbi} - // vcf/tbi: [meta, vcf/tbi] -- sorted combined VCF + // vcf/tbi: [meta, vcf/tbi] -- sorted union VCF + } + + else { + error("combine_method must be 'consensus' or 'all', got '${combine_method}'") } + // + // MODULE: BCFTOOLS_NORM_REJOIN (BCFTOOLS_NORM alias) -- rejoin split sites (-m +any) so LongPhase and Wakhan see one record per position + // Input: [meta, vcf, tbi] -- sorted consensus/union VCF + // Output: .vcf -- [meta, vcf.gz] + // .tbi -- [meta, tbi] + // + BCFTOOLS_NORM_REJOIN( + vcf.join(tbi, failOnMismatch: true, failOnDuplicate: true), + fasta + ) + + BCFTOOLS_NORM_REJOIN.out.vcf + .join(BCFTOOLS_NORM_REJOIN.out.tbi, failOnMismatch: true, failOnDuplicate: true) + .multiMap { meta, rejoined_vcf, rejoined_tbi -> + vcf: [meta, rejoined_vcf] + tbi: [meta, rejoined_tbi] + } + .set { rejoined } + emit: - vcf // [meta, vcf] -- final consensus/combined VCF - tbi // [meta, tbi] + vcf = rejoined.vcf // [meta, vcf] -- final consensus/combined VCF, multi-allelics rejoined + tbi = rejoined.tbi // [meta, tbi] } diff --git a/subworkflows/local/tumor_only/tumoronly_smallvar.nf b/subworkflows/local/tumor_only/tumoronly_smallvar.nf index b37beaf2..ab726ceb 100644 --- a/subworkflows/local/tumor_only/tumoronly_smallvar.nf +++ b/subworkflows/local/tumor_only/tumoronly_smallvar.nf @@ -1,4 +1,6 @@ // IMPORT MODULES +include { BCFTOOLS_VIEW as DEEPVARIANT_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' +include { BCFTOOLS_VIEW as DEEPSOMATIC_PASS_FILTER } from '../../../modules/nf-core/bcftools/view/main' include { CLAIRSTO } from '../../../modules/local/clairsto/main.nf' include { CLAIRSTO_VERDICT_TAG } from '../../../modules/local/clairsto/verdict_tag/main.nf' include { VCFSPLIT } from '../../../modules/local/vcfsplit/main.nf' @@ -9,6 +11,11 @@ include { DEEPSOMATIC } from '../../../subwork include { SMALL_VARIANT_CONSENSUS as GERMLINE_CONSENSUS } from '../../../subworkflows/local/small_variant_consensus.nf' include { SMALL_VARIANT_CONSENSUS as SOMATIC_CONSENSUS } from '../../../subworkflows/local/small_variant_consensus.nf' +// Germline verdict transfer: DeepSomatic adjudicates DeepVariant's tumor-derived germline calls. +include { BCFTOOLS_QUERY as DS_VERDICT_QUERY } from '../../../modules/nf-core/bcftools/query/main' +include { BCFTOOLS_ANNOTATE as DS_VERDICT_ANNOTATE } from '../../../modules/nf-core/bcftools/annotate/main' +include { BCFTOOLS_VIEW as DS_GERMLINE_SELECT } from '../../../modules/nf-core/bcftools/view/main' + workflow TUMORONLY_SMALLVAR { @@ -130,6 +137,50 @@ workflow TUMORONLY_SMALLVAR { // clairsto_somatic_ch: [meta(+caller:'clairs-to'), vcf, tbi] -- somatic variants } + // DEEPSOMATIC in tumor-only mode: normal BAM/BAI are empty lists + if(somatic_var_keep.contains('deepsomatic')) { + tumor_bams + .map { meta, tumor_bam, tumor_bai -> + def normal_bam = [] + def normal_bai = [] + return [meta,normal_bam,normal_bai,tumor_bam,tumor_bai] + } + .set{deepsomatic_input_ch} + // deepsomatic_input_ch: [meta, [], [], tumor_bam, tumor_bai] + // empty normal_bam/bai signals tumor-only mode to DEEPSOMATIC subworkflow + + // + // SUBWORKFLOW: DEEPSOMATIC (local) + // Input: [meta, [], [], tumor_bam, tumor_bai] -- tumor-only (no normal) + // [[:],[]] / fasta / fai / [[:],[]] + // Output: .vcf -- [meta, vcf] + // .vcf_index -- [meta, tbi] + // + DEEPSOMATIC ( + deepsomatic_input_ch, + [[:],[]], // intervals (empty = genome-wide) + fasta, + fai, + [[:],[]], // GZI (empty if FASTA is uncompressed) + ds_pon_channel + ) + // PASS-only copy for downstream steps; published VCFs are untouched. + def deepsomatic_vcf = DEEPSOMATIC.out.vcf.join(DEEPSOMATIC.out.vcf_index) + if (params.smallvar_filter_pass) { + DEEPSOMATIC_PASS_FILTER ( deepsomatic_vcf, [], [], [] ) + deepsomatic_vcf = DEEPSOMATIC_PASS_FILTER.out.vcf + .join(DEEPSOMATIC_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + deepsomatic_vcf + .map{ meta, vcf, tbi -> + def new_meta = meta + [caller:'deepsomatic'] + return [new_meta, vcf, tbi] + } + .set{deepsomatic_ch} + // deepsomatic_ch: [meta(+caller:'deepsomatic'), vcf, tbi] + } + // DEEPVARIANT: germline-only variant calling (no somatic mode for tumor-only) if(germline_var_keep.contains('deepvariant')) { @@ -156,14 +207,62 @@ workflow TUMORONLY_SMALLVAR { [[:],[]] // GFF annotation (not used) ) - DEEPVARIANT.out.vcf - .join(DEEPVARIANT.out.vcf_index) + // PASS-only copy for downstream steps; published VCFs are untouched. + def deepvariant_vcf = DEEPVARIANT.out.vcf.join(DEEPVARIANT.out.vcf_index) + if (params.smallvar_filter_pass) { + DEEPVARIANT_PASS_FILTER ( deepvariant_vcf, [], [], [] ) + deepvariant_vcf = DEEPVARIANT_PASS_FILTER.out.vcf + .join(DEEPVARIANT_PASS_FILTER.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + // Keep only DeepSomatic-adjudicated germline sites; skipped when deepsomatic isn't selected. + def deepvariant_germline = deepvariant_vcf + if (somatic_var_keep.contains('deepsomatic')) { + // GERMLINE VERDICT TRANSFER: DeepVariant on the tumor BAM cannot tell germline from somatic. + // + // MODULE: DS_VERDICT_QUERY (BCFTOOLS_QUERY alias, label: process_single) + // Input: [meta, deepsomatic_vcf, tbi] -- the RAW DeepSomatic VCF, before its PASS filter + // Output: .output/.index -- [meta, tsv.gz/tbi] -- CHROM POS REF ALT FILTER, non-PASS/RefCall rows only + // + DS_VERDICT_QUERY ( DEEPSOMATIC.out.vcf.join(DEEPSOMATIC.out.vcf_index), [], [], [] ) + + // + // MODULE: DS_VERDICT_ANNOTATE (BCFTOOLS_ANNOTATE alias, label: process_medium) + // Stamps INFO/DS_VERDICT on each DeepVariant record from the DeepSomatic verdict table. + // + deepvariant_vcf + .join(DS_VERDICT_QUERY.out.output, failOnMismatch: true, failOnDuplicate: true) + .join(DS_VERDICT_QUERY.out.index, failOnMismatch: true, failOnDuplicate: true) + .map { meta, vcf, tbi, annotations, annotations_index -> + def columns = [] // no extra column specs + def header_lines = [] // no extra header lines + def rename_chrs = [] // no chromosome renaming + return [ meta, vcf, tbi, annotations, annotations_index, columns, header_lines, rename_chrs ] + } + .set{ ds_verdict_annotate_input } + + DS_VERDICT_ANNOTATE ( ds_verdict_annotate_input ) + + // + // MODULE: DS_GERMLINE_SELECT (BCFTOOLS_VIEW alias, label: process_medium) + // Keeps only the positively-adjudicated germline records (see ext.args in conf/modules.config). + // + DS_GERMLINE_SELECT ( + DS_VERDICT_ANNOTATE.out.vcf.join(DS_VERDICT_ANNOTATE.out.tbi, failOnMismatch: true, failOnDuplicate: true), + [], [], [] + ) + + deepvariant_germline = DS_GERMLINE_SELECT.out.vcf + .join(DS_GERMLINE_SELECT.out.index, failOnMismatch: true, failOnDuplicate: true) + } + + deepvariant_germline .map{ meta, vcf, tbi -> def new_meta = meta + [caller:'deepvariant'] return [new_meta, vcf, tbi] } .set{deepvariant_ch} - // deepvariant_ch: [meta(+caller:'deepvariant'), vcf, tbi] + // deepvariant_ch: [meta(+caller:'deepvariant'), vcf, tbi] -- germline-adjudicated if deepsomatic ran } // COMBINE GERMLINE VARIANTS @@ -196,42 +295,6 @@ workflow TUMORONLY_SMALLVAR { .set{germline_vcf} } - // DEEPSOMATIC in tumor-only mode: normal BAM/BAI are empty lists - if(somatic_var_keep.contains('deepsomatic')) { - tumor_bams - .map { meta, tumor_bam, tumor_bai -> - def normal_bam = [] - def normal_bai = [] - return [meta,normal_bam,normal_bai,tumor_bam,tumor_bai] - } - .set{deepsomatic_input_ch} - // deepsomatic_input_ch: [meta, [], [], tumor_bam, tumor_bai] - // empty normal_bam/bai signals tumor-only mode to DEEPSOMATIC subworkflow - - // - // SUBWORKFLOW: DEEPSOMATIC (local) - // Input: [meta, [], [], tumor_bam, tumor_bai] -- tumor-only (no normal) - // [[:],[]] / fasta / fai / [[:],[]] - // Output: .vcf -- [meta, vcf] - // .vcf_index -- [meta, tbi] - // - DEEPSOMATIC ( - deepsomatic_input_ch, - [[:],[]], // intervals (empty = genome-wide) - fasta, - fai, - [[:],[]], // GZI (empty if FASTA is uncompressed) - ds_pon_channel - ) - DEEPSOMATIC.out.vcf - .join(DEEPSOMATIC.out.vcf_index) - .map{ meta, vcf, tbi -> - def new_meta = meta + [caller:'deepsomatic'] - return [new_meta, vcf, tbi] - } - .set{deepsomatic_ch} - // deepsomatic_ch: [meta(+caller:'deepsomatic'), vcf, tbi] - } // COMBINE SOMATIC VARIATION if (somatic_var_keep.size() > 1) { diff --git a/tests/chm13.nf.test.snap b/tests/chm13.nf.test.snap index a267d392..4e196867 100644 --- a/tests/chm13.nf.test.snap +++ b/tests/chm13.nf.test.snap @@ -14,6 +14,9 @@ "CLAIR3": { "clair3": "1.2.0" }, + "CLAIR3_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CLAIRS": { "clairs": "0.4.4" }, @@ -23,6 +26,9 @@ "CLAIRSTO_CNA_RESOURCES": { "coreutils": 9.5 }, + "CLAIRS_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CRAMINO_POST": { "cramino": "1.3.0" }, @@ -101,6 +107,12 @@ "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -136,9 +148,9 @@ } ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:19:37.348390574" + "timestamp": "2026-09-25T12:34:22.954806822" } } \ No newline at end of file diff --git a/tests/clair_only.nf.test.snap b/tests/clair_only.nf.test.snap index ec1ff735..557714f7 100644 --- a/tests/clair_only.nf.test.snap +++ b/tests/clair_only.nf.test.snap @@ -24,12 +24,18 @@ "CLAIR3": { "clair3": "1.2.0" }, + "CLAIR3_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CLAIRS": { "clairs": "0.4.4" }, "CLAIRSTO": { "clairsto": "0.5.1" }, + "CLAIRS_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CRAMINO_POST": { "cramino": "1.3.0" }, @@ -114,6 +120,12 @@ "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -749,38 +761,38 @@ "sample5/vep/somatic/sample5_SOMATIC_VEP.vcf.gz_summary.html" ], [ - "sample1_normal.bam:md5,772e41f7cd86c03a22afbe5ec0592a6b", - "sample1_normal.bam.bai:md5,1b501f6a11efe5d2e6f47b7f1523220b", - "sample1_tumor.bam:md5,c8315c80dc92dfb5d874aef3f5dd46fb", - "sample1_tumor.bam.bai:md5,bc35f807be4b93fc795a14d701469367", + "sample1_normal.bam:md5,9a73c3f90bc4d9a140bcf8f652b0f269", + "sample1_normal.bam.bai:md5,37a866c569f24ed2b38f093f9475b9d6", + "sample1_tumor.bam:md5,dc15bf0e9ff1d401491347c15c408b6d", + "sample1_tumor.bam.bai:md5,fa57db4206692079d9b2084a866bf411", "sample1_normal.flagstat:md5,1c41ea9923945501eb7e41f83a90502d", "sample1_normal.idxstats:md5,902e503387799123ea59255e3fca172c", "sample1_normal.stats:md5,a8b3fba9c54efbc0934d6eacc1807140", "sample1_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample1_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample1_tumor.stats:md5,1c60a1d249d2e503b0678c72e851ea93", - "sample1_whatshap_stats.gtf:md5,eff050a68e36e778b06e0ec19435c569", - "sample1_whatshap_stats.log:md5,76b73731f74fe32ef2d11f6bb0a0f71a", - "sample1_whatshap_stats.tsv:md5,f566ae25b3c5a8f7e94b3d6c1b0417f8", + "sample1_whatshap_stats.gtf:md5,36bda647c08358df0eac0be321c24b20", + "sample1_whatshap_stats.log:md5,938f792fd22bb658a2c7cb1b7653035f", + "sample1_whatshap_stats.tsv:md5,cf8917be389dfef1344eeb6b7e99c7c5", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,47cb0e0bbe71abdbf4f40217dfda43f9", "read_qual.txt:md5,78247dfa2ea336eac0e128eba5e9eef4", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample2_normal.bam:md5,3157bd11ba095a884c7951aafcfcfb1c", - "sample2_normal.bam.bai:md5,edebda44c4383173caea728acde4ac43", - "sample2_tumor.bam:md5,47b2c5f86e0493ba94ff72cea77eeae3", - "sample2_tumor.bam.bai:md5,abf2c290c815f54c2b3f8179f717d9bd", + "sample2_normal.bam:md5,c18bbb1bc05b0ec830fa1ac3c6ef542f", + "sample2_normal.bam.bai:md5,de77684ef264476b3e0531b61ca23740", + "sample2_tumor.bam:md5,8b8cb5ac7668b8ac2f4097f584481d35", + "sample2_tumor.bam.bai:md5,3b46456b0e7b00688519dfb25c8e3c88", "sample2_normal.flagstat:md5,714d0cc0c213e2640e54a16f3d0e6e7e", "sample2_normal.idxstats:md5,72eb83bb11748dc863fef1a0a5497e4b", "sample2_normal.stats:md5,20c47cb94f9ac739d69c57be6daf82c5", "sample2_tumor.flagstat:md5,4344a8745efef9cc2a017024218d61c6", "sample2_tumor.idxstats:md5,69467fc02c83a30084736aeea8b785fb", "sample2_tumor.stats:md5,8635df10132c85a13f2d9878b7cf90a2", - "sample2_whatshap_stats.gtf:md5,4d8f4393e3aebe4e945c0b8236cf3b3e", - "sample2_whatshap_stats.log:md5,10bba7bae6dd99b989ece5e5dac7a8f9", - "sample2_whatshap_stats.tsv:md5,bb46226e486af9026ab76e014624e903", + "sample2_whatshap_stats.gtf:md5,35cd28699c298d99d01cee1c24c6d61b", + "sample2_whatshap_stats.log:md5,75a69a8e651979e25467d6e3c84cdf90", + "sample2_whatshap_stats.tsv:md5,92d7234e355833b1ec6f54951c38c09d", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,48baac86492026a4a7947bc708c47e6e", @@ -830,9 +842,9 @@ ] ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:27:45.728407577" + "timestamp": "2026-09-25T12:28:35.531071958" } } \ No newline at end of file diff --git a/tests/consensus.nf.test.snap b/tests/consensus.nf.test.snap index 6bf3a01d..03293479 100644 --- a/tests/consensus.nf.test.snap +++ b/tests/consensus.nf.test.snap @@ -14,6 +14,9 @@ "BCFTOOLS_NORM": { "bcftools": 1.22 }, + "BCFTOOLS_NORM_REJOIN": { + "bcftools": 1.22 + }, "BCFTOOLS_QUERY": { "bcftools": 1.22 }, @@ -29,12 +32,18 @@ "CLAIR3": { "clair3": "1.2.0" }, + "CLAIR3_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CLAIRS": { "clairs": "0.4.4" }, "CLAIRSTO": { "clairsto": "0.5.1" }, + "CLAIRS_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CRAMINO_POST": { "cramino": "1.3.0" }, @@ -47,6 +56,9 @@ "DEEPSOMATIC_MAKEEXAMPLES": { "deepsomatic": "1.7.0" }, + "DEEPSOMATIC_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPSOMATIC_POSTPROCESSVARIANTS": { "deepsomatic": "1.7.0" }, @@ -56,9 +68,21 @@ "DEEPVARIANT_MAKEEXAMPLES": { "deepvariant": "1.9.0" }, + "DEEPVARIANT_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPVARIANT_POSTPROCESSVARIANTS": { "deepvariant": "1.9.0" }, + "DS_GERMLINE_SELECT": { + "bcftools": "1.23.1" + }, + "DS_VERDICT_ANNOTATE": { + "bcftools": 1.22 + }, + "DS_VERDICT_QUERY": { + "bcftools": 1.22 + }, "GERMLINE_VEP": { "ensemblvep": 115.2, "perl-math-cdf": 0.1, @@ -134,6 +158,12 @@ "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -607,38 +637,38 @@ "sample3/vep/somatic/sample3_SOMATIC_VEP.vcf.gz_summary.html" ], [ - "sample1_normal.bam:md5,93dd8ac8b67eb4eb4bf27e09c8f5f99b", - "sample1_normal.bam.bai:md5,75402ef1cc35229cc131155d9ec973e0", - "sample1_tumor.bam:md5,69cba03cad51bcc1d1ee8c48da042527", - "sample1_tumor.bam.bai:md5,5d633ed05021ad81ce24b1f18cbf38b4", + "sample1_normal.bam:md5,5d3f0615b0d8a9748e6b6bc7d76ab577", + "sample1_normal.bam.bai:md5,3ef784f2536c7506a50d0ed20c3d78fb", + "sample1_tumor.bam:md5,7317ffe61fe3c4f73961db60329adfa9", + "sample1_tumor.bam.bai:md5,6aaeaebe40c422358bb55296c508db2b", "sample1_normal.flagstat:md5,1c41ea9923945501eb7e41f83a90502d", "sample1_normal.idxstats:md5,902e503387799123ea59255e3fca172c", "sample1_normal.stats:md5,a8b3fba9c54efbc0934d6eacc1807140", "sample1_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample1_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample1_tumor.stats:md5,1c60a1d249d2e503b0678c72e851ea93", - "sample1_whatshap_stats.gtf:md5,9f09f9ad1a788384cb8e46a933f77b3b", - "sample1_whatshap_stats.log:md5,20135b4e9965a31d3f9bb0df7d2cec90", - "sample1_whatshap_stats.tsv:md5,264d2d76a9b8d34ea4933aee325ce36e", + "sample1_whatshap_stats.gtf:md5,30bde8f88b7d4e88b935e88e00997ce7", + "sample1_whatshap_stats.log:md5,a3cf683c728ce63a5c2be33edd20f9cc", + "sample1_whatshap_stats.tsv:md5,cca55ca6f99eb0857759a4007c10b346", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,47cb0e0bbe71abdbf4f40217dfda43f9", "read_qual.txt:md5,78247dfa2ea336eac0e128eba5e9eef4", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample2_normal.bam:md5,f2ba30d007c521d479c6158e1e22367a", - "sample2_normal.bam.bai:md5,c3096f52115ec1e24c46fedc41f1f3d3", - "sample2_tumor.bam:md5,1c0287d24fa5b25b86e48024f2f55031", - "sample2_tumor.bam.bai:md5,62849cea5a005e3d8dbe8f9edcefaf60", + "sample2_normal.bam:md5,ece242f245fc4e2210e13e6de93b8fdc", + "sample2_normal.bam.bai:md5,d96d0071ab25ca8dc2327acba4395515", + "sample2_tumor.bam:md5,0c538dc0e27566677313d1f9374fa2b1", + "sample2_tumor.bam.bai:md5,ab2fca59e6729e011c30310eaee2ef1c", "sample2_normal.flagstat:md5,714d0cc0c213e2640e54a16f3d0e6e7e", "sample2_normal.idxstats:md5,72eb83bb11748dc863fef1a0a5497e4b", "sample2_normal.stats:md5,20c47cb94f9ac739d69c57be6daf82c5", "sample2_tumor.flagstat:md5,4344a8745efef9cc2a017024218d61c6", "sample2_tumor.idxstats:md5,69467fc02c83a30084736aeea8b785fb", "sample2_tumor.stats:md5,8635df10132c85a13f2d9878b7cf90a2", - "sample2_whatshap_stats.gtf:md5,f15fb43f0af73d02fc73b66fdc12d5d8", - "sample2_whatshap_stats.log:md5,ca87088fc2f11665eca3fb9c80489085", - "sample2_whatshap_stats.tsv:md5,ca53f81e39bf5d46aa4f604216add1f6", + "sample2_whatshap_stats.gtf:md5,2e5ace4cac0b42bb6132513062781e47", + "sample2_whatshap_stats.log:md5,0979359e459a14037a704733f32d3b95", + "sample2_whatshap_stats.tsv:md5,8c974019af5c839c72604e7526ae8a3d", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,48baac86492026a4a7947bc708c47e6e", @@ -662,9 +692,9 @@ ] ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:40:36.642665738" + "timestamp": "2026-09-25T12:21:59.090749802" } } \ No newline at end of file diff --git a/tests/deep_only.nf.test.snap b/tests/deep_only.nf.test.snap index 7887aaca..9a25aad9 100644 --- a/tests/deep_only.nf.test.snap +++ b/tests/deep_only.nf.test.snap @@ -23,6 +23,9 @@ "DEEPSOMATIC_MAKEEXAMPLES": { "deepsomatic": "1.7.0" }, + "DEEPSOMATIC_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPSOMATIC_POSTPROCESSVARIANTS": { "deepsomatic": "1.7.0" }, @@ -32,9 +35,21 @@ "DEEPVARIANT_MAKEEXAMPLES": { "deepvariant": "1.9.0" }, + "DEEPVARIANT_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPVARIANT_POSTPROCESSVARIANTS": { "deepvariant": "1.9.0" }, + "DS_GERMLINE_SELECT": { + "bcftools": "1.23.1" + }, + "DS_VERDICT_ANNOTATE": { + "bcftools": 1.22 + }, + "DS_VERDICT_QUERY": { + "bcftools": 1.22 + }, "GERMLINE_VEP": { "ensemblvep": 115.2, "perl-math-cdf": 0.1, @@ -107,6 +122,12 @@ "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -563,8 +584,8 @@ "sample1_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample1_tumor.stats:md5,1c60a1d249d2e503b0678c72e851ea93", "sample1_whatshap_stats.gtf:md5,e1d0e87353a5f9aed8a9ac4bf7973427", - "sample1_whatshap_stats.log:md5,bd6b83a062e22cd3201523dc4c2c13e7", - "sample1_whatshap_stats.tsv:md5,7a1508751cb1daa841a577ae25f55586", + "sample1_whatshap_stats.log:md5,43811a1aa4726c2ef62646ff22d1fad0", + "sample1_whatshap_stats.tsv:md5,a821c5d0e3451f82327645fee89e68eb", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,47cb0e0bbe71abdbf4f40217dfda43f9", @@ -582,22 +603,22 @@ "sample2_tumor.idxstats:md5,69467fc02c83a30084736aeea8b785fb", "sample2_tumor.stats:md5,8635df10132c85a13f2d9878b7cf90a2", "sample2_whatshap_stats.gtf:md5,af33281699a1d0da83fbe7eaff198d03", - "sample2_whatshap_stats.log:md5,bbd9ab2ce07a009d9348a1d78bc6fc70", - "sample2_whatshap_stats.tsv:md5,c65436f930c23ddbfd568532d07dce70", + "sample2_whatshap_stats.log:md5,0cd536df69e244a5271e9c5440ea0f3f", + "sample2_whatshap_stats.tsv:md5,c7cc47024ef622a72e7f4f0d5fa119eb", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,48baac86492026a4a7947bc708c47e6e", "read_qual.txt:md5,8b92ff7dc4536188be159b95525511cd", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample3_tumor.bam:md5,965264ef8436cb887ad02c92d40fb50e", - "sample3_tumor.bam.bai:md5,cfc6329667a3c6c66e3c0ca0ace4c6e9", + "sample3_tumor.bam:md5,9449340acf09825767e1f5b4a3f0e5c2", + "sample3_tumor.bam.bai:md5,3be46bae7402c3865f27881457b7a466", "sample3_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample3_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample3_tumor.stats:md5,ecd5ea4fee37379dd5c5ae3e89dfddda", - "sample3_whatshap_stats.gtf:md5,f47156e18c490ff9a4e6efd04d43acc5", - "sample3_whatshap_stats.log:md5,4f7648e763004ab764143cb4f8b6499e", - "sample3_whatshap_stats.tsv:md5,4cb58bb3b663aaba23da004d69adab3e", + "sample3_whatshap_stats.gtf:md5,5a9b20b6ed25ef2ede4271156a83ce88", + "sample3_whatshap_stats.log:md5,a62c149286176f863d4293d0748a6dec", + "sample3_whatshap_stats.tsv:md5,90c20ce05eef472d9f5bb77754636d31", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,56e899f85876cee082788927d0f89c5f", @@ -607,9 +628,9 @@ ] ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:46:42.562085661" + "timestamp": "2026-09-25T12:56:04.936282901" } } \ No newline at end of file diff --git a/tests/default.nf.test.snap b/tests/default.nf.test.snap index 22e72e6f..62439fb0 100644 --- a/tests/default.nf.test.snap +++ b/tests/default.nf.test.snap @@ -14,12 +14,18 @@ "CLAIR3": { "clair3": "1.2.0" }, + "CLAIR3_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CLAIRS": { "clairs": "0.4.4" }, "CLAIRSTO": { "clairsto": "0.5.1" }, + "CLAIRS_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CRAMINO_POST": { "cramino": "1.3.0" }, @@ -98,6 +104,12 @@ "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -553,38 +565,38 @@ "sample3/vep/somatic/sample3_SOMATIC_VEP.vcf.gz_summary.html" ], [ - "sample1_normal.bam:md5,772e41f7cd86c03a22afbe5ec0592a6b", - "sample1_normal.bam.bai:md5,1b501f6a11efe5d2e6f47b7f1523220b", - "sample1_tumor.bam:md5,c8315c80dc92dfb5d874aef3f5dd46fb", - "sample1_tumor.bam.bai:md5,bc35f807be4b93fc795a14d701469367", + "sample1_normal.bam:md5,9a73c3f90bc4d9a140bcf8f652b0f269", + "sample1_normal.bam.bai:md5,37a866c569f24ed2b38f093f9475b9d6", + "sample1_tumor.bam:md5,dc15bf0e9ff1d401491347c15c408b6d", + "sample1_tumor.bam.bai:md5,fa57db4206692079d9b2084a866bf411", "sample1_normal.flagstat:md5,1c41ea9923945501eb7e41f83a90502d", "sample1_normal.idxstats:md5,902e503387799123ea59255e3fca172c", "sample1_normal.stats:md5,a8b3fba9c54efbc0934d6eacc1807140", "sample1_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample1_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample1_tumor.stats:md5,1c60a1d249d2e503b0678c72e851ea93", - "sample1_whatshap_stats.gtf:md5,eff050a68e36e778b06e0ec19435c569", - "sample1_whatshap_stats.log:md5,76b73731f74fe32ef2d11f6bb0a0f71a", - "sample1_whatshap_stats.tsv:md5,f566ae25b3c5a8f7e94b3d6c1b0417f8", + "sample1_whatshap_stats.gtf:md5,36bda647c08358df0eac0be321c24b20", + "sample1_whatshap_stats.log:md5,938f792fd22bb658a2c7cb1b7653035f", + "sample1_whatshap_stats.tsv:md5,cf8917be389dfef1344eeb6b7e99c7c5", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,47cb0e0bbe71abdbf4f40217dfda43f9", "read_qual.txt:md5,78247dfa2ea336eac0e128eba5e9eef4", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample2_normal.bam:md5,3157bd11ba095a884c7951aafcfcfb1c", - "sample2_normal.bam.bai:md5,edebda44c4383173caea728acde4ac43", - "sample2_tumor.bam:md5,47b2c5f86e0493ba94ff72cea77eeae3", - "sample2_tumor.bam.bai:md5,abf2c290c815f54c2b3f8179f717d9bd", + "sample2_normal.bam:md5,c18bbb1bc05b0ec830fa1ac3c6ef542f", + "sample2_normal.bam.bai:md5,de77684ef264476b3e0531b61ca23740", + "sample2_tumor.bam:md5,8b8cb5ac7668b8ac2f4097f584481d35", + "sample2_tumor.bam.bai:md5,3b46456b0e7b00688519dfb25c8e3c88", "sample2_normal.flagstat:md5,714d0cc0c213e2640e54a16f3d0e6e7e", "sample2_normal.idxstats:md5,72eb83bb11748dc863fef1a0a5497e4b", "sample2_normal.stats:md5,20c47cb94f9ac739d69c57be6daf82c5", "sample2_tumor.flagstat:md5,4344a8745efef9cc2a017024218d61c6", "sample2_tumor.idxstats:md5,69467fc02c83a30084736aeea8b785fb", "sample2_tumor.stats:md5,8635df10132c85a13f2d9878b7cf90a2", - "sample2_whatshap_stats.gtf:md5,4d8f4393e3aebe4e945c0b8236cf3b3e", - "sample2_whatshap_stats.log:md5,10bba7bae6dd99b989ece5e5dac7a8f9", - "sample2_whatshap_stats.tsv:md5,bb46226e486af9026ab76e014624e903", + "sample2_whatshap_stats.gtf:md5,35cd28699c298d99d01cee1c24c6d61b", + "sample2_whatshap_stats.log:md5,75a69a8e651979e25467d6e3c84cdf90", + "sample2_whatshap_stats.tsv:md5,92d7234e355833b1ec6f54951c38c09d", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,48baac86492026a4a7947bc708c47e6e", @@ -608,9 +620,9 @@ ] ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:10:56.589241561" + "timestamp": "2026-09-25T11:57:04.096725193" } } \ No newline at end of file diff --git a/tests/fixtures/vcftag_input.vcf b/tests/fixtures/vcftag_input.vcf new file mode 100644 index 00000000..a1945635 --- /dev/null +++ b/tests/fixtures/vcftag_input.vcf @@ -0,0 +1,13 @@ +##fileformat=VCFv4.2 +##INFO= +##INFO= +##FILTER= +##FILTER= +##FILTER= +##contig= +#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT testsample +chr1 100 . A C 30 PASS CALLER=clairs-to GT:DP 0/1:20 +chr1 200 . G T 20 NonSomatic CALLER=clairs-to;EXISTING GT:DP 0/1:18 +chr1 300 . T A 10 RefCall . GT:DP 0/0:12 +chr1 400 . C G 40 . CALLER=deepsomatic GT:DP 1/1:25 +chr1 500 . G A 5 LowQual;RefCall CALLER=clair3;EXISTING GT:DP 0/1:9 diff --git a/tests/fixtures/vcftag_input.vcf.gz b/tests/fixtures/vcftag_input.vcf.gz new file mode 100644 index 00000000..c956c3c5 Binary files /dev/null and b/tests/fixtures/vcftag_input.vcf.gz differ diff --git a/tests/fixtures/vcftag_input.vcf.gz.tbi b/tests/fixtures/vcftag_input.vcf.gz.tbi new file mode 100644 index 00000000..75130768 Binary files /dev/null and b/tests/fixtures/vcftag_input.vcf.gz.tbi differ diff --git a/tests/union.nf.test b/tests/union.nf.test index 2fe8925b..8acb4956 100644 --- a/tests/union.nf.test +++ b/tests/union.nf.test @@ -115,4 +115,45 @@ nextflow_pipeline { ) } } + + test("-profile test, union combine mode, smallvar_filter_pass=false") { + + when { + params { + outdir = "$outputDir" + germline_var_combine = 'all' + somatic_var_combine = 'all' + germline_var_keep = 'clair, deepvariant' + somatic_var_keep = 'clair, deepsomatic' + smallvar_filter_pass = false + } + } + + then { + assertAll( + { assert workflow.success }, + + // ── No PASS filter process may run ─────────────────────────── + // With the param false no _PASS_FILTER process runs or reports a version. + { + def versions = file("$outputDir/pipeline_info/lrsomatic_software_mqc_versions.yml") + assert versions.exists() + assert !versions.text.contains('_PASS_FILTER') + }, + + // ── Phased VCFs still exist and have data ──────────────────── + { + ['sample1', 'sample2', 'sample3'].each { s -> + def germline = file("$launchDir/output/${s}/variants/phased/germline_smallvariants.vcf.gz") + def somatic = file("$launchDir/output/${s}/variants/phased/somatic_smallvariants.vcf.gz") + assert germline.exists() + assert somatic.exists() + assert germline.size() > 0 + assert somatic.size() > 0 + } + } + ) + } + } + } diff --git a/tests/union.nf.test.snap b/tests/union.nf.test.snap index 654a82d7..42bea83c 100644 --- a/tests/union.nf.test.snap +++ b/tests/union.nf.test.snap @@ -14,6 +14,9 @@ "BCFTOOLS_NORM": { "bcftools": 1.22 }, + "BCFTOOLS_NORM_REJOIN": { + "bcftools": 1.22 + }, "BCFTOOLS_QUERY": { "bcftools": 1.22 }, @@ -26,12 +29,18 @@ "CLAIR3": { "clair3": "1.2.0" }, + "CLAIR3_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CLAIRS": { "clairs": "0.4.4" }, "CLAIRSTO": { "clairsto": "0.5.1" }, + "CLAIRS_PASS_FILTER": { + "bcftools": "1.23.1" + }, "CRAMINO_POST": { "cramino": "1.3.0" }, @@ -44,6 +53,9 @@ "DEEPSOMATIC_MAKEEXAMPLES": { "deepsomatic": "1.7.0" }, + "DEEPSOMATIC_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPSOMATIC_POSTPROCESSVARIANTS": { "deepsomatic": "1.7.0" }, @@ -53,9 +65,21 @@ "DEEPVARIANT_MAKEEXAMPLES": { "deepvariant": "1.9.0" }, + "DEEPVARIANT_PASS_FILTER": { + "bcftools": "1.23.1" + }, "DEEPVARIANT_POSTPROCESSVARIANTS": { "deepvariant": "1.9.0" }, + "DS_GERMLINE_SELECT": { + "bcftools": "1.23.1" + }, + "DS_VERDICT_ANNOTATE": { + "bcftools": 1.22 + }, + "DS_VERDICT_QUERY": { + "bcftools": 1.22 + }, "GERMLINE_VEP": { "ensemblvep": 115.2, "perl-math-cdf": 0.1, @@ -126,14 +150,17 @@ "SORT_POST_NORM": { "bcftools": 1.22 }, - "STANDARDIZE_AF": { - "bcftools": 1.22 - }, "SV_VEP": { "ensemblvep": 115.2, "perl-math-cdf": 0.1, "tabix": 1.21 }, + "TAG_GERMLINE": { + "bcftools": 1.2 + }, + "TAG_SOMATIC": { + "bcftools": 1.2 + }, "UNTAR": { "untar": 1.34 }, @@ -607,52 +634,52 @@ "sample3/vep/somatic/sample3_SOMATIC_VEP.vcf.gz_summary.html" ], [ - "sample1_normal.bam:md5,cbfc940a38c74cbe8435c18b9da5dd32", - "sample1_normal.bam.bai:md5,c1498328929d45b2898fa2265b0d617c", - "sample1_tumor.bam:md5,b78866edf991393806d37505d16f7e3d", - "sample1_tumor.bam.bai:md5,f613de14ab19fc3a85403661a4f6188c", + "sample1_normal.bam:md5,f0809c8be190289f21339696883dea6f", + "sample1_normal.bam.bai:md5,cbff669c086761e893a84ba3ec032b5f", + "sample1_tumor.bam:md5,6267bf3e86a69536e45968b908d1cc6f", + "sample1_tumor.bam.bai:md5,1ec6d68c92182ff9051b8707c4975445", "sample1_normal.flagstat:md5,1c41ea9923945501eb7e41f83a90502d", "sample1_normal.idxstats:md5,902e503387799123ea59255e3fca172c", "sample1_normal.stats:md5,a8b3fba9c54efbc0934d6eacc1807140", "sample1_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample1_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample1_tumor.stats:md5,1c60a1d249d2e503b0678c72e851ea93", - "sample1_whatshap_stats.gtf:md5,9ae556e13516dd47d4108acf2104bddb", - "sample1_whatshap_stats.log:md5,eaddcf6a1666d4a3c1ad3316dac24139", - "sample1_whatshap_stats.tsv:md5,c2773e011c2781160fd9a7741b10546b", + "sample1_whatshap_stats.gtf:md5,f5d331899db63b2ae21e51436bad2ddd", + "sample1_whatshap_stats.log:md5,111538b7e305189fb6de95ea978d06c8", + "sample1_whatshap_stats.tsv:md5,fa0fd5ce2b0919098ffa7d960cf177de", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,47cb0e0bbe71abdbf4f40217dfda43f9", "read_qual.txt:md5,78247dfa2ea336eac0e128eba5e9eef4", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample2_normal.bam:md5,f2ba30d007c521d479c6158e1e22367a", - "sample2_normal.bam.bai:md5,c3096f52115ec1e24c46fedc41f1f3d3", - "sample2_tumor.bam:md5,1c0287d24fa5b25b86e48024f2f55031", - "sample2_tumor.bam.bai:md5,62849cea5a005e3d8dbe8f9edcefaf60", + "sample2_normal.bam:md5,c18bbb1bc05b0ec830fa1ac3c6ef542f", + "sample2_normal.bam.bai:md5,de77684ef264476b3e0531b61ca23740", + "sample2_tumor.bam:md5,8b8cb5ac7668b8ac2f4097f584481d35", + "sample2_tumor.bam.bai:md5,3b46456b0e7b00688519dfb25c8e3c88", "sample2_normal.flagstat:md5,714d0cc0c213e2640e54a16f3d0e6e7e", "sample2_normal.idxstats:md5,72eb83bb11748dc863fef1a0a5497e4b", "sample2_normal.stats:md5,20c47cb94f9ac739d69c57be6daf82c5", "sample2_tumor.flagstat:md5,4344a8745efef9cc2a017024218d61c6", "sample2_tumor.idxstats:md5,69467fc02c83a30084736aeea8b785fb", "sample2_tumor.stats:md5,8635df10132c85a13f2d9878b7cf90a2", - "sample2_whatshap_stats.gtf:md5,f15fb43f0af73d02fc73b66fdc12d5d8", - "sample2_whatshap_stats.log:md5,a6767b3490cafdcbaf3b7114644028de", - "sample2_whatshap_stats.tsv:md5,570796e5e291229e8872733425e0b133", + "sample2_whatshap_stats.gtf:md5,35cd28699c298d99d01cee1c24c6d61b", + "sample2_whatshap_stats.log:md5,9308d0bd8dc9a4a86359f1ec926e6cb1", + "sample2_whatshap_stats.tsv:md5,b81282a609eca306912d4f7b7fdfeebb", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,48baac86492026a4a7947bc708c47e6e", "read_qual.txt:md5,8b92ff7dc4536188be159b95525511cd", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", - "sample3_tumor.bam:md5,116b6944da4aa833a8d21c46b5f5ecfe", - "sample3_tumor.bam.bai:md5,ea9eca53bbaba26d40b791a2ea1aadf6", + "sample3_tumor.bam:md5,94e7f18f0a22d16f0868780df7975447", + "sample3_tumor.bam.bai:md5,a2cc1013d6fde4fec50608fa9681a7a7", "sample3_tumor.flagstat:md5,8ff32d733c62c4910bf185ef24bf27cf", "sample3_tumor.idxstats:md5,2de140e61f9e86c9c10af20dd565cc93", "sample3_tumor.stats:md5,ecd5ea4fee37379dd5c5ae3e89dfddda", - "sample3_whatshap_stats.gtf:md5,f47156e18c490ff9a4e6efd04d43acc5", - "sample3_whatshap_stats.log:md5,679dcfa209888a9e69a07e4c4e4b049e", - "sample3_whatshap_stats.tsv:md5,035d5aa0425ba3fc32d65268b793b424", + "sample3_whatshap_stats.gtf:md5,5a9b20b6ed25ef2ede4271156a83ce88", + "sample3_whatshap_stats.log:md5,dc9a34518331417b7b5811462b722791", + "sample3_whatshap_stats.tsv:md5,26a07a6db5f1e5bf7844aa0a8c190cb7", "breakpoint_clusters.tsv:md5,d36a70de292ee130ef30da4a58bced18", "breakpoint_clusters_list.tsv:md5,0c0ce62e329f8de492487e8414c30a50", "breakpoints_double.csv:md5,56e899f85876cee082788927d0f89c5f", @@ -662,9 +689,9 @@ ] ], "meta": { - "nf-test": "0.9.0", - "nextflow": "26.04.6" + "nf-test": "0.9.3", + "nextflow": "25.10.4" }, - "timestamp": "2026-09-21T11:54:50.88115842" + "timestamp": "2026-09-25T12:04:10.544092583" } } \ No newline at end of file diff --git a/workflows/lrsomatic.nf b/workflows/lrsomatic.nf index 37bd1c4c..03517069 100644 --- a/workflows/lrsomatic.nf +++ b/workflows/lrsomatic.nf @@ -696,10 +696,11 @@ workflow LRSOMATIC { // ascat_tumoronly_ch: [meta, purityploidy, segments] // All ASCAT files per sample for the report module, which globs by suffix - ch_ascat_files = ASCAT.out.segments_raw - .mix(ASCAT.out.purityploidy, ASCAT.out.png) - .groupTuple() - .map { meta, files -> [meta, files.flatten()] } + // Joined per sample; segments_raw is optional, so it arrives as null when absent + ch_ascat_files = ASCAT.out.purityploidy + .join(ASCAT.out.png) + .join(ASCAT.out.segments_raw, remainder: true) + .map { meta, purityploidy, png, segments_raw -> [meta, [purityploidy, png, segments_raw ?: []].flatten()] } // ch_ascat_files: [meta, [file, file, ...]] } @@ -1255,10 +1256,11 @@ workflow LRSOMATIC { ) // The WAKHAN outputs the report renders: ranked solutions, heatmap, per-solution plots + // Joined per sample; all three outputs are required ch_wakhan_files = WAKHAN.out.solutions_ranks - .mix(WAKHAN.out.heatmap_html, WAKHAN.out.solution_dirs) - .groupTuple() - .map { meta, files -> [meta, files.flatten()] } // solution_dirs contributes a list + .join(WAKHAN.out.heatmap_html) + .join(WAKHAN.out.solution_dirs) + .map { meta, ranks, heatmap, dirs -> [meta, [ranks, heatmap, dirs].flatten()] } // dirs may be a list // ch_wakhan_files: [meta, [file_or_dir, ...]] } @@ -1276,13 +1278,19 @@ workflow LRSOMATIC { .set { report_id_meta } // report_id_meta: [id, meta] - ch_somatic_vep_vcf - .map { meta, vcf -> [meta.id, vcf] } - .set { report_vep_ch } + // A skipped module leaves an empty leg, and remainder: true then defers every sample to + // channel close. One [] per sample keeps each leg matched so samples report independently. + def report_empty_slot = { -> report_id_meta.map { id, _meta -> [id, []] } } - ch_sv_vep_vcf - .map { meta, vcf -> [meta.id, vcf] } - .set { report_sv_vep_ch } + def report_vep_ch = params.skip_vep + ? report_empty_slot.call() + : ch_somatic_vep_vcf.map { meta, vcf -> [meta.id, vcf] } + // report_vep_ch: [id, vcf] + + def report_sv_vep_ch = params.skip_vep + ? report_empty_slot.call() + : ch_sv_vep_vcf.map { meta, vcf -> [meta.id, vcf] } + // report_sv_vep_ch: [id, vcf] SEVERUS.out.somatic_vcf .map { meta, vcf -> [meta.id, vcf] } @@ -1292,29 +1300,63 @@ workflow LRSOMATIC { .map { meta, vcf, _tbi -> [meta.id, vcf] } .set { report_somatic_ch } - ch_ascat_files - .map { meta, files -> [meta.id, files] } - .set { report_ascat_ch } + def report_ascat_ch = params.skip_ascat + ? report_empty_slot.call() + : ch_ascat_files.map { meta, files -> [meta.id, files] } + // report_ascat_ch: [id, [files]] - ch_wakhan_files - .map { meta, files -> [meta.id, files] } - .set { report_wakhan_ch } + def report_wakhan_ch = params.skip_wakhan + ? report_empty_slot.call() + : ch_wakhan_files.map { meta, files -> [meta.id, files] } + // report_wakhan_ch: [id, [files]] + + // One emission per sample per tool, none optional: mosdepth 2, cramino 1, samtools 2. + // Adding another per-sample emission to either mix below must bump this count. + def qc_files_per_sample = params.skip_qc + ? 0 + : (params.skip_mosdepth ? 0 : 2) + (params.skip_cramino ? 0 : 1) + (params.skip_bamstats ? 0 : 2) // Tumor-side QC, keyed by the sample id (= report id) + // groupKey: emit a sample's bundle on its own files; toString() restores a plain String key ch_mosdepth_summary .mix(ch_mosdepth_global, ch_cramino_post_txt, ch_bam_stats, ch_bam_flagstat) .filter { meta, _f -> meta.type == 'tumor' } - .map { meta, f -> [meta.id, f] } + .map { meta, f -> [groupKey(meta.id, qc_files_per_sample), f] } .groupTuple() - .set { report_qc_tumor_ch } + .map { key, files -> [key.toString(), files] } + .set { report_qc_tumor_grouped } + // report_qc_tumor_grouped: [id, [qc_file, ...]] // Normal-side QC (matched mode): a pair shares meta.id, so already keyed by the report id ch_mosdepth_summary .mix(ch_mosdepth_global, ch_cramino_post_txt, ch_bam_stats, ch_bam_flagstat) .filter { meta, _f -> meta.type == 'normal' } - .map { meta, f -> [meta.id, f] } + .map { meta, f -> [groupKey(meta.id, qc_files_per_sample), f] } .groupTuple() - .set { report_qc_normal_ch } + .map { key, files -> [key.toString(), files] } + .set { report_qc_normal_grouped } + // report_qc_normal_grouped: [id, [qc_file, ...]] -- paired samples only + + def report_qc_tumor_ch = qc_files_per_sample == 0 + ? report_empty_slot.call() + : report_qc_tumor_grouped + + // Normal-side QC covers paired samples only; meta.paired_data gives the tumor-only arm + // its [] up front instead of waiting out channel close for a match that never arrives. + report_id_meta + .branch { _id, meta -> + paired: meta.paired_data + tumor_only: true + } + .set { report_roster } + + def report_qc_normal_ch = qc_files_per_sample == 0 + ? report_empty_slot.call() + : report_roster.paired + .join(report_qc_normal_grouped) + .map { id, _meta, files -> [id, files] } + .mix(report_roster.tumor_only.map { id, _meta -> [id, []] }) + // report_qc_normal_ch: [id, [files] | []] -- full roster report_id_meta .join(report_vep_ch, remainder: true)