Mutational signature attribution analysis, featuring automated optimisation study with simulated data.
The purpose of this README is to provide a guide for the quick start of using MSA. An extensive Wiki page detailing the usage of this tool can be found here.
The best way to run the code is by using Nextflow. Once you have installed Nextflow, run the test job locally or on your favourite cluster:
nextflow run https://gitlab.com/s.senkin/MSA -profile conda,testImportant: The pipeline has been updated to Nextflow DSL2 and requires Nextflow version 21.04.0 or later. The version can be changed using the following command:
export NXF_VER=25.04.8Nextflow 25+/26+: newer Nextflow ships a stricter config/DSL parser that rejects some constructs this pipeline still uses (e.g. for loops and the mix(*list) spread). Until the pipeline is migrated to the strict syntax, select the legacy parser via an environment variable before running (e.g. in your job script or shell profile):
export NXF_SYNTAX_PARSER=v1It is recommended to use docker or singularity profiles as these normally provide greater stability and reproducibility than conda.
The pipeline should run and produce all the results automatically. You can also retrieve the code (see below) in order to adjust all the inputs and parameters. In the nextflow.config file various parameters can be specified.
MSA natively supports SigProfilerExtractor and SigProfilerMatrixGenerator outputs.
The simplest way to run is as follows:
nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
--dataset SP_test \
--SP_extractor_output_path /full/path/to/SP_extractor_output/If SigProfilerMatrixGenerator output is provided, it will take priority over the SigProfilerExtractor one for input mutation matrices:
nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
--dataset SP_test \
--SP_matrix_generator_output_path /full/path/to/SP_matrix_generator_output/ \
--SP_extractor_output_path /full/path/to/SP_extractor_output/MSA now supports direct specification of individual mutation tables and signature files:
nextflow run https://gitlab.com/s.senkin/MSA -profile conda \
--dataset my_data \
--input_mutation_table /path/to/mutations.txt \
--signatures_file /path/to/signatures.txt \
--mutation_types SBS \
--SBS_context 96Note: When using specific files, you must specify exactly ONE mutation type and (for SBS) ONE context.
All parameters are described in the dedicated wiki page. Most general parameters are listed below.
The pipeline processes inputs in the following priority order:
- Specific files (highest):
--input_mutation_tableand--signatures_file - SigProfiler outputs:
--SP_extractor_output_pathand--SP_matrix_generator_output_path - Default directories (lowest):
--input_tablesand--signature_tables
| Parameter | Default value | Description |
|---|---|---|
| --help | null | Print usage and optional parameters |
| --input_mutation_table | null | Specific mutation table file to convert and use (overrides all other inputs) |
| --signatures_file | null | Specific signature file to convert and use (overrides all other inputs) |
| --SP_matrix_generator_output_path | null | Use SigProfilerMatrixGenerator output from specified full path |
| --SP_extractor_output_path | null | Use SigProfilerExtractor output from specified full path to attribute signatures |
| --dataset | SIM_test | Set the name of the dataset |
| --input_tables | $baseDir/input_mutation_tables | Full path to input mutation tables directory |
| --signature_tables | $baseDir/signature_tables | Full path to input signature tables directory |
| --signature_prefix | sigProfiler | Prefix of signature files (e.g. sigProfiler, sigRandom) |
| --output_path | . | Output path for plots and tables |
| --temp_path | $output_path/temp | Temporary path for converted inputs |
| --mutation_types | ['SBS'] | Mutation types to analyse. Only ONE can be specified from command line; use list in script for multiple |
| --number_of_samples | -1 | Number of samples to analyse (-1 means all available) |
| --SBS_context | 96 | SBS context to use (96, 192, 288, 1536, or 4608) |
| --COSMIC_signatures | false | If true, use COSMIC signatures from SigProfiler output; otherwise use de-novo |
The pipeline organizes outputs as follows:
output_path/
├── output_tables/ # Final NNLS results
│ └── {dataset}/
│ ├── output_*_mutations_table.csv
│ ├── output_*_weights_table.csv
│ ├── signatures_prevalences_*.csv
│ └── bootstrap_output/
├── output_tables_unoptimised/ # Unoptimized results
├── plots/ # All generated plots
└── temp/ # Temporary conversions
├── input_tables/
└── signature_tables/
The greedy NNLS signature-optimisation hot path can be offloaded to the GPU via cuML's batched NNLS solver. This is opt-in and does not affect the default CPU behaviour.
-
Build the GPU environment (once per cluster + GPU architecture, as a build job on a node with an NVIDIA driver — not the login node). This builds the source-only cuML into a conda env and tops it up with MSA's dependencies:
./setup_gpu_env.sh # CUDA 13.3 (default) ./setup_gpu_env.sh -c all_cuda-129_arch-x86_64 -e all_cuda-129_arch-x86_64 # if the GPU driver only supports CUDA 12.x
Match the CUDA version to the driver on your GPU nodes (
nvidia-smishows the maximum supported CUDA). The script prints the full env path to use in the next step. -
Enable it in nextflow.config (or on the command line).
gpu_conda_envmust be the full path to the built env — a bare name is treated by Nextflow'scondadirective as a package to install:use_GPU = true gpu_conda_env = '/home/you/miniforge3/envs/all_cuda-133_arch-x86_64'
-
Route the NNLS jobs to a GPU queue. Cluster-specific queue/GPU-request settings are deliberately not committed; put them in your own config (e.g.
gpu.config) and pass it with-c gpu.configonly for GPU runs. See gpu.config.example for a template scoped to thegpu_nnls-labelled processes.
Then run with a conda-enabled profile:
nextflow run run_auto_optimised_analysis.nf -profile conda,test_GPU -c gpu.configRetrieve the code:
git clone https://gitlab.com/s.senkin/MSA.git
cd MSARun the fully automated pipeline with optimisation on the test dataset:
nextflow run run_auto_optimised_analysis.nf -profile conda,test \
--output_path testAlternatively, to run the pipeline without optimisation, using fixed penalties (DSL1 only, please downgrade Nextflow with export NXF_VER=22.10.4):
nextflow run run_analysis.nf -profile conda \
--dataset SIM_test \
--weak_threshold 0.02 \
--output_path testThe pipeline partially supports Nextflow's -resume functionality to restart from cached processes:
nextflow run run_auto_optimised_analysis.nf -profile conda -resume \
--dataset my_data \
--output_path resultsIf you cannot run nextflow, you can still run some basic analysis manually (scripts in the ./bin folder). Dependencies: pandas, numpy, scipy, matplotlib and seaborn.
Set up the virtual environment using conda:
conda env create -f environment.ymlThis only needs to be done once. Afterwards, activate the environment:
conda activate msaAlternatively, you can use docker with the provided Dockerfile, or use ready-made images (docker or singularity).
-
input_mutation_tables/SIM folder contains a set produced with existing (PCAWG or COSMIC) signatures using normal distributions of mutational burdens.
-
To reproduce the simulated set of samples with reshuffled SBS1/5/22/40 PCAWG signatures:
python bin/simulate_data.py -t SBS -c 96 -n 100 -s signature_tables- input_mutation_tables/SIMrand folder contains simulated samples where each contains contributions from 5 randomly selected signatures out of 100 Poisson-generated signatures. To reproduce:
python bin/generate_random_signatures.py -t SBS -c 96 -n 100
python bin/simulate_data.py -r -t SBS -c 96 -n 100 -d SIMrandBoth scripts support normal distribution noise using -z option with standard deviation set by -Z option (2 by default). Additional flags: use -h option.
nextflow run run_auto_optimised_analysis.nf -profile condanextflow run run_auto_optimised_analysis.nf -profile conda \
--dataset my_cohort \
--SP_extractor_output_path /data/sigprofiler_output \
--output_path resultsnextflow run run_auto_optimised_analysis.nf -profile conda \
--dataset my_sample \
--input_mutation_table data/my_mutations_SBS96.txt \
--mutation_types SBS \
--SBS_context 96 \
--output_path resultsEdit run_auto_optimised_analysis.nf:
params.mutation_types = ['SBS', 'DBS', 'ID']Then run:
nextflow run run_auto_optimised_analysis.nf -profile conda \
--dataset my_data \
--output_path resultsKey improvements in the DSL2 version:
- Modern Nextflow syntax (DSL2)
- Improved caching and resume functionality
- Support for specific file inputs
- Cleaner module organization
- All temporary files in configurable temp directory
- Better error handling and validation
- Unified output directory structure
If you use this pipeline, please cite the original publication and the MSA repository.
