Skip to content

Repository files navigation

Mutational Signature Attribution pipeline

logo

Mutational signature attribution analysis, featuring automated optimisation study with simulated data.

Introduction

The purpose of this README is to provide a guide for the quick start of using MSA. An extensive Wiki page detailing the usage of this tool can be found here.

Running with Nextflow

The best way to run the code is by using Nextflow. Once you have installed Nextflow, run the test job locally or on your favourite cluster:

nextflow run https://gitlab.com/s.senkin/MSA -profile conda,test

Important: The pipeline has been updated to Nextflow DSL2 and requires Nextflow version 21.04.0 or later. The version can be changed using the following command:

export NXF_VER=25.04.8

Nextflow 25+/26+: newer Nextflow ships a stricter config/DSL parser that rejects some constructs this pipeline still uses (e.g. for loops and the mix(*list) spread). Until the pipeline is migrated to the strict syntax, select the legacy parser via an environment variable before running (e.g. in your job script or shell profile):

export NXF_SYNTAX_PARSER=v1

It is recommended to use docker or singularity profiles as these normally provide greater stability and reproducibility than conda.

The pipeline should run and produce all the results automatically. You can also retrieve the code (see below) in order to adjust all the inputs and parameters. In the nextflow.config file various parameters can be specified.

Running on SigProfiler output

MSA natively supports SigProfilerExtractor and SigProfilerMatrixGenerator outputs.

The simplest way to run is as follows:

nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
    --dataset SP_test \
    --SP_extractor_output_path /full/path/to/SP_extractor_output/

If SigProfilerMatrixGenerator output is provided, it will take priority over the SigProfilerExtractor one for input mutation matrices:

nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
    --dataset SP_test \
    --SP_matrix_generator_output_path /full/path/to/SP_matrix_generator_output/ \
    --SP_extractor_output_path /full/path/to/SP_extractor_output/

Running with specific input files

MSA now supports direct specification of individual mutation tables and signature files:

nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
    --dataset my_data \
    --input_mutation_table /path/to/mutations.txt \
    --signatures_file /path/to/signatures.txt \
    --mutation_types SBS \
    --SBS_context 96

Note: When using specific files, you must specify exactly ONE mutation type and (for SBS) ONE context.

Options

All parameters are described in the dedicated wiki page. Most general parameters are listed below.

Input Priority

The pipeline processes inputs in the following priority order:

  1. Specific files (highest): --input_mutation_table and --signatures_file
  2. SigProfiler outputs: --SP_extractor_output_path and --SP_matrix_generator_output_path
  3. Default directories (lowest): --input_tables and --signature_tables

General parameters

Parameter Default value Description
--help null Print usage and optional parameters
--input_mutation_table null Specific mutation table file to convert and use (overrides all other inputs)
--signatures_file null Specific signature file to convert and use (overrides all other inputs)
--SP_matrix_generator_output_path null Use SigProfilerMatrixGenerator output from specified full path
--SP_extractor_output_path null Use SigProfilerExtractor output from specified full path to attribute signatures
--dataset SIM_test Set the name of the dataset
--input_tables $baseDir/input_mutation_tables Full path to input mutation tables directory
--signature_tables $baseDir/signature_tables Full path to input signature tables directory
--signature_prefix sigProfiler Prefix of signature files (e.g. sigProfiler, sigRandom)
--output_path . Output path for plots and tables
--temp_path $output_path/temp Temporary path for converted inputs
--mutation_types ['SBS'] Mutation types to analyse. Only ONE can be specified from command line; use list in script for multiple
--number_of_samples -1 Number of samples to analyse (-1 means all available)
--SBS_context 96 SBS context to use (96, 192, 288, 1536, or 4608)
--COSMIC_signatures false If true, use COSMIC signatures from SigProfiler output; otherwise use de-novo

Output structure

The pipeline organizes outputs as follows:

output_path/
├── output_tables/              # Final NNLS results
│   └── {dataset}/
│       ├── output_*_mutations_table.csv
│       ├── output_*_weights_table.csv
│       ├── signatures_prevalences_*.csv
│       └── bootstrap_output/
├── output_tables_unoptimised/  # Unoptimized results
├── plots/                      # All generated plots
└── temp/                       # Temporary conversions
    ├── input_tables/
    └── signature_tables/

GPU acceleration (experimental)

The greedy NNLS signature-optimisation hot path can be offloaded to the GPU via cuML's batched NNLS solver. This is opt-in and does not affect the default CPU behaviour.

  1. Build the GPU environment (once per cluster + GPU architecture, as a build job on a node with an NVIDIA driver — not the login node). This builds the source-only cuML into a conda env and tops it up with MSA's dependencies:

    ./setup_gpu_env.sh                              # CUDA 13.3 (default)
    ./setup_gpu_env.sh -c all_cuda-129_arch-x86_64 -e all_cuda-129_arch-x86_64   # if the GPU driver only supports CUDA 12.x

    Match the CUDA version to the driver on your GPU nodes (nvidia-smi shows the maximum supported CUDA). The script prints the full env path to use in the next step.

  2. Enable it in nextflow.config (or on the command line). gpu_conda_env must be the full path to the built env — a bare name is treated by Nextflow's conda directive as a package to install:

    use_GPU       = true
    gpu_conda_env = '/home/you/miniforge3/envs/all_cuda-133_arch-x86_64'
  3. Route the NNLS jobs to a GPU queue. Cluster-specific queue/GPU-request settings are deliberately not committed; put them in your own config (e.g. gpu.config) and pass it with -c gpu.config only for GPU runs. See gpu.config.example for a template scoped to the gpu_nnls-labelled processes.

Then run with a conda-enabled profile:

nextflow run run_auto_optimised_analysis.nf -profile conda,test_GPU -c gpu.config

Running manually

Getting started

Retrieve the code:

git clone https://gitlab.com/s.senkin/MSA.git
cd MSA

Run the fully automated pipeline with optimisation on the test dataset:

nextflow run run_auto_optimised_analysis.nf -profile conda,test \
    --output_path test

Alternatively, to run the pipeline without optimisation, using fixed penalties (DSL1 only, please downgrade Nextflow with export NXF_VER=22.10.4):

nextflow run run_analysis.nf -profile conda \
    --dataset SIM_test \
    --weak_threshold 0.02 \
    --output_path test

Using resume

The pipeline partially supports Nextflow's -resume functionality to restart from cached processes:

nextflow run run_auto_optimised_analysis.nf -profile conda -resume \
    --dataset my_data \
    --output_path results

Setting up dependencies

If you cannot run nextflow, you can still run some basic analysis manually (scripts in the ./bin folder). Dependencies: pandas, numpy, scipy, matplotlib and seaborn.

Set up the virtual environment using conda:

conda env create -f environment.yml

This only needs to be done once. Afterwards, activate the environment:

conda activate msa

Alternatively, you can use docker with the provided Dockerfile, or use ready-made images (docker or singularity).

Simulating data

  • input_mutation_tables/SIM folder contains a set produced with existing (PCAWG or COSMIC) signatures using normal distributions of mutational burdens.

  • To reproduce the simulated set of samples with reshuffled SBS1/5/22/40 PCAWG signatures:

python bin/simulate_data.py -t SBS -c 96 -n 100 -s signature_tables
  • input_mutation_tables/SIMrand folder contains simulated samples where each contains contributions from 5 randomly selected signatures out of 100 Poisson-generated signatures. To reproduce:
python bin/generate_random_signatures.py -t SBS -c 96 -n 100
python bin/simulate_data.py -r -t SBS -c 96 -n 100 -d SIMrand

Both scripts support normal distribution noise using -z option with standard deviation set by -Z option (2 by default). Additional flags: use -h option.

Example workflows

Basic analysis with repository test data

nextflow run run_auto_optimised_analysis.nf -profile conda

Analysis with SigProfiler extractor output

nextflow run run_auto_optimised_analysis.nf -profile conda \
    --dataset my_cohort \
    --SP_extractor_output_path /data/sigprofiler_output \
    --output_path results

Analysis with specific mutation table

nextflow run run_auto_optimised_analysis.nf -profile conda \
    --dataset my_sample \
    --input_mutation_table data/my_mutations_SBS96.txt \
    --mutation_types SBS \
    --SBS_context 96 \
    --output_path results

Multiple mutation types (must be set in script)

Edit run_auto_optimised_analysis.nf:

params.mutation_types = ['SBS', 'DBS', 'ID']

Then run:

nextflow run run_auto_optimised_analysis.nf -profile conda \
    --dataset my_data \
    --output_path results

Changes from DSL1

Key improvements in the DSL2 version:

  • Modern Nextflow syntax (DSL2)
  • Improved caching and resume functionality
  • Support for specific file inputs
  • Cleaner module organization
  • All temporary files in configurable temp directory
  • Better error handling and validation
  • Unified output directory structure

Citation

If you use this pipeline, please cite the original publication and the MSA repository.

About

A mirror of the original gitlab.com/s.senkin/MSA repository

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages