Skip to content

New module: xengsort/classify - #12864

Merged
SPPearce merged 14 commits into
nf-core:masterfrom
LeonHornich:xengsort_classify
Sep 25, 2026
Merged

SPPearce merged 14 commits into
nf-core:masterfrom
LeonHornich:xengsort_classify

Conversation

@LeonHornich

@LeonHornich LeonHornich commented Sep 1, 2026 •

Copy link
Copy Markdown
Contributor

Added new module xengsort/classify for classifying reads using a existing index file.

Notes:

  • apart from versions_xengsort, all output files are declared optional. This is due to the optional parameters --filter and --count. When specifying those a different set of files is produced.
  • in the nf-tests i only check for file creation since i got md5sum for empty file found error otherwise. I am assuming this is due to the minimum and small nature of the testing data

PR checklist

  • This comment contains a description of changes (with reason).
  • If you've fixed a bug or added code that should be tested, add tests!
  • If you've added a new tool - have you followed the module conventions in the contribution docs
  • If necessary, include test data in your PR.
  • Remove all TODO statements.
  • Broadcast software version numbers to topic: versions - See version_topics
  • Follow the naming conventions.
  • Follow the parameters requirements.
  • Follow the input/output options guidelines.
  • Add a resource label
  • Use BioConda and BioContainers if possible to fulfil software requirements.
  • Ensure that the test works with either Docker / Singularity. Conda CI tests can be quite flaky:
    • For modules:
      • nf-core modules test <MODULE> --profile docker
      • nf-core modules test <MODULE> --profile singularity
      • nf-core modules test <MODULE> --profile conda

@SPPearce SPPearce left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are all your output files unstable? That doesn't seem great for the tool...
Can you use sanitizeOutput from nft-utils please.

@SPPearce SPPearce left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are all the output files unstable?

@LeonHornich
LeonHornich force-pushed the xengsort_classify branch 2 times, most recently from e47802d to f3682fd Compare September 13, 2026 18:37
@LeonHornich

LeonHornich commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor Author

Sorry for the auto ping on all the people, not intended. In the process of pushing some changes something went wrong and i had to reset. The current commit should be fine with what i want it to be.

*have been removed from the review list

@LeonHornich

Copy link
Copy Markdown
Contributor Author

@SPPearce i looked into it a bit more. When i just run:


... i get empty files.

My understanding is that due to the small nature of the test data, some of the output files host, graft, both, neither, ambiguous stay empty and this causes the error.

@SPPearce

Copy link
Copy Markdown
Contributor

If some of the files are empty, then yes that makes sense as to why you need to mark them as unstable. But are ANY of the files non-empty?

@LeonHornich

Copy link
Copy Markdown
Contributor Author

Alright, finally got to look into this again. I checked where the test reads are being written to:

0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-ambiguous.1.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-ambiguous.2.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-both.1.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-both.2.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-graft.1.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-graft.2.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-host.1.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-host.2.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-neither.1.fq.gz
     0 reads  .nf-test/tests/48f34874e238eb6d9d9dfbaf43e01ff9/work/cd/9563c52ff3c8f9b0bf6dbefbcf85c1/test-neither.2.fq.gz
     0 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-ambiguous.fq.gz
     0 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-both.fq.gz
   100 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-graft.fq.gz
     0 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-host.fq.gz
     0 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-neither.fq.gz
     0 reads  .nf-test/tests/63ca06643019f3f6d133a0b3c84102b3/work/dd/b0adac5c921857896747e68b2616fb/test-sites.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-ambiguous.1.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-ambiguous.2.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-both.1.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-both.2.fq.gz
   100 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-graft.1.fq.gz
   100 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-graft.2.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-host.1.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-host.2.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-neither.1.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-neither.2.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-sites.1.fq.gz
     0 reads  .nf-test/tests/c4fc34643a2860da654ec49f35bec2d9/work/1f/f805be6a762bd19fa3b0187db5a186/test-sites.2.fq.gz

Looks like they are all going in the graft output. I have therefore removed 'graft' from unstableKeys and now the tests and linting finally passes. Tests and linting is passing now

@SPPearce
SPPearce added this pull request to the merge queue Sep 25, 2026
Merged via the queue into nf-core:master with commit 7fb83dc Sep 25, 2026
26 checks passed
vagkaratzas pushed a commit that referenced this pull request Sep 29, 2026
* initiated module template

* added xengsort classify module

* added xengsort classify tests

* updated testing to only check for file existance

* trimmed some whitespaces

* minor formatting correction

* updated testing

* classify tests are now run on single thread

* actual single thread adjustments

* updated snapshot

* defining unstable outputs

* nf-test single-end graft output is empty and therefore unstable

* removed graft from unstable keys

* removed params.classify_args from nextflow.config
vagkaratzas pushed a commit that referenced this pull request Sep 29, 2026
* initiated module template

* added xengsort classify module

* added xengsort classify tests

* updated testing to only check for file existance

* trimmed some whitespaces

* minor formatting correction

* updated testing

* classify tests are now run on single thread

* actual single thread adjustments

* updated snapshot

* defining unstable outputs

* nf-test single-end graft output is empty and therefore unstable

* removed graft from unstable keys

* removed params.classify_args from nextflow.config
famosab pushed a commit to Lukecele/modules that referenced this pull request Sep 29, 2026
* Bump mgnifam/generatefamilies to 3.1.0 and add stats output

mgnifam 3.1.0 writes <chunk>_stats.json, a MultiQC-ready run summary.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* pigz/compress stub fix (nf-core#13015)

.gz files should be properly created

* Read eggnogmapper's version from package metadata (nf-core#13016)

eggnog-mapper's get_version() runs `git describe --tags` with cwd set to its
own installed package directory, and only falls back to __VERSION__ when that
fails. Under conda the package is installed inside the pipeline's own checkout,
so `emapper.py --version` prints that repository's tag rather than the tool's
version:

    emapper-1.4.1-184-g1b3550e / Expected eggNOG DB version: 5.0.2 / ...

The module's grep then reported 1.4.1 (a metatdenovo release tag) instead of
2.1.13. Container installs sit outside any repository, and the biocontainer has
no git, so they were unaffected and the drift only showed up in a pipeline's
conda CI.

importlib.metadata reads the installed distribution's own metadata, which no
surrounding repository can influence. The package name is passed through
sys.argv to keep the expression free of nested quotes, which nf-core modules
lint's main.nf parser does not handle.


Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* checkm2/predict: pass database via CHECKM2DB and use small test DB (nf-core#13019)

CheckM2 enforces a hardcoded checksum on databases passed with
--database_path, but not on the one set through the CHECKM2DB
environment variable. Using the env var allows reduced databases,
so the test now uses the small DB from test-datasets instead of
downloading the full CheckM2 database.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* gffread: accept gzipped gff and fasta input (nf-core#13024)

* gffread: accept gzipped gff and fasta input

gffread cannot read gzip: a gzipped gff silently yields no features and a
gzipped fasta fails with "sequence lines in a FASTA record must have the
same length". Stream a gzipped gff through stdin, and decompress a gzipped
fasta to a temporary file, removed afterwards so it does not match the
*.fasta output glob.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvGYU6kuXJeVv9ffdSzdzH

* Update modules/nf-core/gffread/main.nf

Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>

* Bump foldseek/createdb and foldseek/easysearch to 10.941cd33 (nf-core#13027)

* Bump foldseek/createdb and foldseek/easysearch to 10.941cd33

- Update conda env and biocontainers to foldseek 10.941cd33
- easysearch: discover the main DB basename via its .lookup file instead
  of misusing ext.prefix2/meta2.id
- Add maintainers and .m8 EDAM ontology to meta.yml
- Fix stub test name and regenerate snapshots

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* foldseek/easysearch: select DB via meta2.id, fail fast on lookup fallback

Use ${db}/${meta2.id} when it exists so pipelines can still pick a
database in a multi-DB directory; otherwise fall back to the single
.lookup file and exit with an error if zero or several are found.
Add a test covering the fallback path.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat: update trna to add all optional outputs (nf-core#13029)

* feat: update trna to add all optional outputs

* fix: snapshots

* Add taxonkit/lca module (nf-core#13023)

* Add taxonkit/lca module

* Restructure taxonkit/lca input error checking*

*Add tests to confirm the module appropriately fails if neither taxids, nor a file containing taxids are passed and if both are passed

* Fix precommit format issue in taxonkit/lca

* Fix trailing whitespace issue for taxonkit/lca

* Remove `MINIMAC4` sites input (nf-core#13022)

* Fix miniconda version

* Remove sites options from minimac4

* Update test

* Add --sites argument as output

* Set to optional

* Add ontology

---------

Co-authored-by: Kevin-Brockers <57921086+Kevin-Brockers@users.noreply.github.com>

* New module: xengsort/classify (nf-core#12864)

* initiated module template

* added xengsort classify module

* added xengsort classify tests

* updated testing to only check for file existance

* trimmed some whitespaces

* minor formatting correction

* updated testing

* classify tests are now run on single thread

* actual single thread adjustments

* updated snapshot

* defining unstable outputs

* nf-test single-end graft output is empty and therefore unstable

* removed graft from unstable keys

* removed params.classify_args from nextflow.config

* New module: samtools/trimheader (nf-core#13035)

* New module: samtools/trimheader

Removes @sq lines for references without reads, re-encoding every record
against the smaller header. featureCounts (Subread 2.1.1) cannot read a BAM
header over 2 GiB, which a fragmented assembly of about 75M contigs produces.
samtools reheader cannot do this, since it leaves record reference ids as
they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

* Update modules/nf-core/samtools/trimheader/tests/main.nf.test

Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>

* Key readsMD5 on the bam channel and regenerate the snapshot

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>

* Fix interproscan ignoring staged database (nf-core#13032)

interproscan.sh cd's into its install dir before reading
INTERPROSCAN_CONF, so the relative path resolved to the container's
own properties file and the bundled sample data was always used.
Use an absolute INTERPROSCAN_CONF, point data.directory at the staged
data/ folder and bin.directory at the InterProScan bin folder.

Drop Hamap from the database test: the test data ships hamap/2023_05,
whereas InterProScan 5.59-91 expects hamap/2021_04. Document the
database requirements in meta.yml.

Closes nf-core#13009

Generated by Claude Opus 5.5

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Update conda-incubator/setup-miniconda digest to be893c9

* fix(hmmer/eslsfetchindex): use GNU sort for large indexes (nf-core#13041)

* fix(hmmer/eslsfetchindex): use GNU sort for large indexes

Add Coreutils to the module environment and Wave containers so sort can spill to TMPDIR. Report the Coreutils version and document temporary disk use.

Generated by Codex

* remove empty config file

* feat(rsem,dupradar): make temp-file scratch location overridable via ext.args (nf-core#13043)

* feat(rsem,dupradar): make temp-file scratch location overridable via ext.args

Both RSEM_CALCULATEEXPRESSION/SENTIEON_RSEMCALCULATEEXPRESSION's
--temporary-folder and DUPRADAR's featureCounts tmpDir are currently
hardcoded to a path relative to the task's working directory, with no
way to point them elsewhere. On some network/object-storage-backed work
directories (Fusion, NFS, Lustre), the tools' own scratch I/O can be a
meaningful cost, and pipelines that have a genuine local scratch mount
available have no way to use it without forking the module.

This only adds an opt-in override, defaulting exactly as before:
- RSEM: ext.args can now include --temporary-folder, guarded the same
  way many other nf-core modules guard a default CLI flag.
- dupradar: a new tmp_dir option, parsed by the template's existing
  parse_args() mechanism (same pattern already used for feature_type),
  forwarded to analyzeDuprates()'s tmpDir, which dupRadar itself passes
  through to Rsubread::featureCounts.

Context: nf-core/rnaseq#1957, nf-core/rnaseq#1962

* fix(dupradar): create tmp_dir if it doesn't already exist

featureCounts doesn't create its own tmpDir - the default '.' always
already exists so this was never visible before, but a custom tmp_dir
override fails outright ("temporary directory is not writable") unless
something creates it first. Verified locally: without this, overriding
tmp_dir fails; with it, featureCounts picks it up and uses it correctly.

* fix(dupradar): don't suppress dir.create warnings for tmp_dir

showWarnings=FALSE hid the one diagnostic (dir.create's own warning
naming the actual OS-level reason) that would explain why a custom
tmp_dir couldn't be created, leaving only featureCounts' own less
specific "not writable" error. Only call dir.create when the path
doesn't already exist (so the default "." case never touches this
at all), then explicitly verify existence and writability with a
clear error naming the path.

* fix(rsem/calculateexpression): remove unused publishDir causing setup failure

The "homo_sapiens - bam" test's alignment.config set a publishDir
referencing params.outdir, but the test's `then` block only ever
asserts on process.out directly via snapshot() - nothing reads from
a published path. That publishDir also applied to RSEM_PREPAREREFERENCE,
run in this test's `setup` block before `when.params.outdir` is set,
so every run failed with "Access to undefined parameter `outdir`"
regardless of anything this PR changes.

Checked the other modules sharing this same publishDir/outdir config
pattern (bcl2fastq, lima, pycoqc, rseqc/bamstat, rseqc/inferexperiment,
spring/decompress, subread/featurecounts) - none combine it with a
setup dependency, so this exact conflict hasn't come up before.

* fix(rsem/calculateexpression): remove second unused publishDir/outdir instance

Same issue as the previous commit, in nextflow.config (used by the
"fastq" and "stub" tests) rather than alignment.config. The "stub"
test's own when block never set params.outdir at all - it only
happened to pass when run together with the other tests in this file,
because nf-test carries params state across sequential tests in the
same run. CI shards tests individually, which exposed it. Removed the
now-pointless params.outdir assignments from "bam" and "fastq" too,
since nothing reads a published path in any of these tests' assertions.

* Add octopusv submodules - filter + subset (nf-core#13040)

* Add octopusv submodules - filter + subset

* update tests

* address review feedback

* Update `bcftools_mpileup` (nf-core#13047)

* Update bcftools mpileup

* Update meta

* Fix test

* Fix test

* Add parabricks/deepsomatic (nf-core#12797)

* Add parabricks/deepsomatic

* fix tests, use data with accepted readgroup values

* remove hallucinated field

---------

Co-authored-by: Friederike Hanssen <Friederike.hanssen@qbic.uni-tuebingen.de>
Co-authored-by: Friederike Hanssen <friederike.hanssen@seqera.io>

* Bump mgnifam/generatefamilies to 4.0.0

mgnifam 4.0.0 renames the MultiQC summary to <chunk>_mgnifam_stats.json
and reorders/renames the metadata CSV columns. Point the documentation
link to the new docs site.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Matthieu Muffato <mm49@sanger.ac.uk>
Co-authored-by: Daniel Lundin <erik.rikard.daniel@gmail.com>
Co-authored-by: Diego Alvarez S. <dialvarezs@gmail.com>
Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>
Co-authored-by: Jim Downie <19718667+prototaxites@users.noreply.github.com>
Co-authored-by: jcbioinformatics <60235622+jcbioinformatics@users.noreply.github.com>
Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>
Co-authored-by: Kevin-Brockers <57921086+Kevin-Brockers@users.noreply.github.com>
Co-authored-by: Leon Hornich <82643611+LeonHornich@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Jonathan Manning <jonathan.manning@seqera.io>
Co-authored-by: Manas <manas.sehgal@dkfz.de>
Co-authored-by: Simon Pearce <24893913+SPPearce@users.noreply.github.com>
Co-authored-by: Friederike Hanssen <Friederike.hanssen@qbic.uni-tuebingen.de>
Co-authored-by: Friederike Hanssen <friederike.hanssen@seqera.io>
famosab pushed a commit to Lukecele/modules that referenced this pull request Sep 29, 2026
* Bump mgnifam/updatefamilies to 3.1.0 and add stats output

mgnifam 3.1.0 writes <chunk>_updated_stats.json, a MultiQC-ready run summary.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* pigz/compress stub fix (nf-core#13015)

.gz files should be properly created

* Read eggnogmapper's version from package metadata (nf-core#13016)

eggnog-mapper's get_version() runs `git describe --tags` with cwd set to its
own installed package directory, and only falls back to __VERSION__ when that
fails. Under conda the package is installed inside the pipeline's own checkout,
so `emapper.py --version` prints that repository's tag rather than the tool's
version:

    emapper-1.4.1-184-g1b3550e / Expected eggNOG DB version: 5.0.2 / ...

The module's grep then reported 1.4.1 (a metatdenovo release tag) instead of
2.1.13. Container installs sit outside any repository, and the biocontainer has
no git, so they were unaffected and the drift only showed up in a pipeline's
conda CI.

importlib.metadata reads the installed distribution's own metadata, which no
surrounding repository can influence. The package name is passed through
sys.argv to keep the expression free of nested quotes, which nf-core modules
lint's main.nf parser does not handle.


Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* checkm2/predict: pass database via CHECKM2DB and use small test DB (nf-core#13019)

CheckM2 enforces a hardcoded checksum on databases passed with
--database_path, but not on the one set through the CHECKM2DB
environment variable. Using the env var allows reduced databases,
so the test now uses the small DB from test-datasets instead of
downloading the full CheckM2 database.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* gffread: accept gzipped gff and fasta input (nf-core#13024)

* gffread: accept gzipped gff and fasta input

gffread cannot read gzip: a gzipped gff silently yields no features and a
gzipped fasta fails with "sequence lines in a FASTA record must have the
same length". Stream a gzipped gff through stdin, and decompress a gzipped
fasta to a temporary file, removed afterwards so it does not match the
*.fasta output glob.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvGYU6kuXJeVv9ffdSzdzH

* Update modules/nf-core/gffread/main.nf

Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>

* Bump foldseek/createdb and foldseek/easysearch to 10.941cd33 (nf-core#13027)

* Bump foldseek/createdb and foldseek/easysearch to 10.941cd33

- Update conda env and biocontainers to foldseek 10.941cd33
- easysearch: discover the main DB basename via its .lookup file instead
  of misusing ext.prefix2/meta2.id
- Add maintainers and .m8 EDAM ontology to meta.yml
- Fix stub test name and regenerate snapshots

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* foldseek/easysearch: select DB via meta2.id, fail fast on lookup fallback

Use ${db}/${meta2.id} when it exists so pipelines can still pick a
database in a multi-DB directory; otherwise fall back to the single
.lookup file and exit with an error if zero or several are found.
Add a test covering the fallback path.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat: update trna to add all optional outputs (nf-core#13029)

* feat: update trna to add all optional outputs

* fix: snapshots

* Add taxonkit/lca module (nf-core#13023)

* Add taxonkit/lca module

* Restructure taxonkit/lca input error checking*

*Add tests to confirm the module appropriately fails if neither taxids, nor a file containing taxids are passed and if both are passed

* Fix precommit format issue in taxonkit/lca

* Fix trailing whitespace issue for taxonkit/lca

* Remove `MINIMAC4` sites input (nf-core#13022)

* Fix miniconda version

* Remove sites options from minimac4

* Update test

* Add --sites argument as output

* Set to optional

* Add ontology

---------

Co-authored-by: Kevin-Brockers <57921086+Kevin-Brockers@users.noreply.github.com>

* New module: xengsort/classify (nf-core#12864)

* initiated module template

* added xengsort classify module

* added xengsort classify tests

* updated testing to only check for file existance

* trimmed some whitespaces

* minor formatting correction

* updated testing

* classify tests are now run on single thread

* actual single thread adjustments

* updated snapshot

* defining unstable outputs

* nf-test single-end graft output is empty and therefore unstable

* removed graft from unstable keys

* removed params.classify_args from nextflow.config

* New module: samtools/trimheader (nf-core#13035)

* New module: samtools/trimheader

Removes @sq lines for references without reads, re-encoding every record
against the smaller header. featureCounts (Subread 2.1.1) cannot read a BAM
header over 2 GiB, which a fragmented assembly of about 75M contigs produces.
samtools reheader cannot do this, since it leaves record reference ids as
they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

* Update modules/nf-core/samtools/trimheader/tests/main.nf.test

Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>

* Key readsMD5 on the bam channel and regenerate the snapshot

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LydWahEnWZexTULEub5bfX

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>

* Fix interproscan ignoring staged database (nf-core#13032)

interproscan.sh cd's into its install dir before reading
INTERPROSCAN_CONF, so the relative path resolved to the container's
own properties file and the bundled sample data was always used.
Use an absolute INTERPROSCAN_CONF, point data.directory at the staged
data/ folder and bin.directory at the InterProScan bin folder.

Drop Hamap from the database test: the test data ships hamap/2023_05,
whereas InterProScan 5.59-91 expects hamap/2021_04. Document the
database requirements in meta.yml.

Closes nf-core#13009

Generated by Claude Opus 5.5

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Update conda-incubator/setup-miniconda digest to be893c9

* fix(hmmer/eslsfetchindex): use GNU sort for large indexes (nf-core#13041)

* fix(hmmer/eslsfetchindex): use GNU sort for large indexes

Add Coreutils to the module environment and Wave containers so sort can spill to TMPDIR. Report the Coreutils version and document temporary disk use.

Generated by Codex

* remove empty config file

* feat(rsem,dupradar): make temp-file scratch location overridable via ext.args (nf-core#13043)

* feat(rsem,dupradar): make temp-file scratch location overridable via ext.args

Both RSEM_CALCULATEEXPRESSION/SENTIEON_RSEMCALCULATEEXPRESSION's
--temporary-folder and DUPRADAR's featureCounts tmpDir are currently
hardcoded to a path relative to the task's working directory, with no
way to point them elsewhere. On some network/object-storage-backed work
directories (Fusion, NFS, Lustre), the tools' own scratch I/O can be a
meaningful cost, and pipelines that have a genuine local scratch mount
available have no way to use it without forking the module.

This only adds an opt-in override, defaulting exactly as before:
- RSEM: ext.args can now include --temporary-folder, guarded the same
  way many other nf-core modules guard a default CLI flag.
- dupradar: a new tmp_dir option, parsed by the template's existing
  parse_args() mechanism (same pattern already used for feature_type),
  forwarded to analyzeDuprates()'s tmpDir, which dupRadar itself passes
  through to Rsubread::featureCounts.

Context: nf-core/rnaseq#1957, nf-core/rnaseq#1962

* fix(dupradar): create tmp_dir if it doesn't already exist

featureCounts doesn't create its own tmpDir - the default '.' always
already exists so this was never visible before, but a custom tmp_dir
override fails outright ("temporary directory is not writable") unless
something creates it first. Verified locally: without this, overriding
tmp_dir fails; with it, featureCounts picks it up and uses it correctly.

* fix(dupradar): don't suppress dir.create warnings for tmp_dir

showWarnings=FALSE hid the one diagnostic (dir.create's own warning
naming the actual OS-level reason) that would explain why a custom
tmp_dir couldn't be created, leaving only featureCounts' own less
specific "not writable" error. Only call dir.create when the path
doesn't already exist (so the default "." case never touches this
at all), then explicitly verify existence and writability with a
clear error naming the path.

* fix(rsem/calculateexpression): remove unused publishDir causing setup failure

The "homo_sapiens - bam" test's alignment.config set a publishDir
referencing params.outdir, but the test's `then` block only ever
asserts on process.out directly via snapshot() - nothing reads from
a published path. That publishDir also applied to RSEM_PREPAREREFERENCE,
run in this test's `setup` block before `when.params.outdir` is set,
so every run failed with "Access to undefined parameter `outdir`"
regardless of anything this PR changes.

Checked the other modules sharing this same publishDir/outdir config
pattern (bcl2fastq, lima, pycoqc, rseqc/bamstat, rseqc/inferexperiment,
spring/decompress, subread/featurecounts) - none combine it with a
setup dependency, so this exact conflict hasn't come up before.

* fix(rsem/calculateexpression): remove second unused publishDir/outdir instance

Same issue as the previous commit, in nextflow.config (used by the
"fastq" and "stub" tests) rather than alignment.config. The "stub"
test's own when block never set params.outdir at all - it only
happened to pass when run together with the other tests in this file,
because nf-test carries params state across sequential tests in the
same run. CI shards tests individually, which exposed it. Removed the
now-pointless params.outdir assignments from "bam" and "fastq" too,
since nothing reads a published path in any of these tests' assertions.

* Add octopusv submodules - filter + subset (nf-core#13040)

* Add octopusv submodules - filter + subset

* update tests

* address review feedback

* Update `bcftools_mpileup` (nf-core#13047)

* Update bcftools mpileup

* Update meta

* Fix test

* Fix test

* Add parabricks/deepsomatic (nf-core#12797)

* Add parabricks/deepsomatic

* fix tests, use data with accepted readgroup values

* remove hallucinated field

---------

Co-authored-by: Friederike Hanssen <Friederike.hanssen@qbic.uni-tuebingen.de>
Co-authored-by: Friederike Hanssen <friederike.hanssen@seqera.io>

* Bump mgnifam/updatefamilies to 4.0.0

mgnifam 4.0.0 renames the MultiQC summary to
<chunk>_updated_mgnifam_stats.json, reorders/renames the metadata CSV
columns and leaves `converged` empty under --skip_refine. Point the
documentation link to the new docs site.

Generated by Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Matthieu Muffato <mm49@sanger.ac.uk>
Co-authored-by: Daniel Lundin <erik.rikard.daniel@gmail.com>
Co-authored-by: Diego Alvarez S. <dialvarezs@gmail.com>
Co-authored-by: Paolo Inglese <26252284+piplus2@users.noreply.github.com>
Co-authored-by: Jim Downie <19718667+prototaxites@users.noreply.github.com>
Co-authored-by: jcbioinformatics <60235622+jcbioinformatics@users.noreply.github.com>
Co-authored-by: Louis Le Nézet <58640615+LouisLeNezet@users.noreply.github.com>
Co-authored-by: Kevin-Brockers <57921086+Kevin-Brockers@users.noreply.github.com>
Co-authored-by: Leon Hornich <82643611+LeonHornich@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Jonathan Manning <jonathan.manning@seqera.io>
Co-authored-by: Manas <manas.sehgal@dkfz.de>
Co-authored-by: Simon Pearce <24893913+SPPearce@users.noreply.github.com>
Co-authored-by: Friederike Hanssen <Friederike.hanssen@qbic.uni-tuebingen.de>
Co-authored-by: Friederike Hanssen <friederike.hanssen@seqera.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants