This repository contains the complete learning material for the Python for Bioinformatics workshop. It is organised as a sequence of hands-on Jupyter notebooks that take you from querying public biological databases, through downloading and parsing sequencing data, to running a full RNA-seq analysis pipeline and building an interactive data app.
The notebooks are sequential — each one builds on the concepts and techniques introduced in the previous ones.
Click on the following link to create your GitHub Codespace:
Once the Codespace is created, run the following in the terminal to install some missing packages:
pip install biopython
pip install dash
pip install dash-bio
pip install dash-bootstrap-componentsYou are ready to go 🚀!
The material is organised into five main components:
| Directory | Description |
|---|---|
notebooks/ |
The 5 core workshop notebooks (covered step by step below) |
automatize/ |
Reusable Python tools for downloading data from UniProt and plotting sequence distributions |
rnaseq/ |
A Nextflow RNA-seq pipeline (FASTQC → Trim Galore → HISAT2 → MultiQC) |
scripts/ |
Shell helpers, e.g. for downloading raw reads from SRA |
data/ |
Sample datasets used throughout the workshop (reads, CSV tables, references) |
solutions/ |
Reference implementations for the workshop exercises |
notebook1_uniprot.ipynb— Query and bulk-download protein sequences from the UniProt REST APInotebook2_sra.ipynb— Fetch raw sequencing data from the NCBI Sequence Read Archive (SRA)notebook3_rnaseq.ipynb— Explore an RNA-seq differential expression analysisnotebook4_dash.ipynb— Build interactive visualisation apps with Plotly Dashnotebook5_project.ipynb— Consolidate everything into a mini research project
By the end of this workshop you should be able to:
- Query biological databases through REST APIs (with the UniProt API as the worked example)
- Perform bulk downloads of sequences and parse paginated API responses
- Understand the difference between common formats such as FASTA, FASTQ and TSV
- Fetch raw sequencing reads from NCBI SRA using the SRA Toolkit
- Parse and manipulate FASTA/FASTQ sequence files with Biopython
- Load, filter and explore tabular datasets with pandas
- Extract meaningful structure from real datasets (e.g. a 19th-century census table)
- Produce publication-style static charts with matplotlib (histograms, distributions)
- Build rich interactive figures with plotly
- Turn a figure into a shareable static HTML file
- Design multi-component apps with Plotly Dash
- Connect charts to user interaction through Dash callbacks
- Critically evaluate when a full server-based app is worth the added complexity (vs. a static plotly HTML export)
- Be aware of the deployment and hosting options for Dash apps
- Understand what an RNA-seq analysis pipeline does at each stage: quality control (FASTQC), adapter trimming (Trim Galore), alignment (HISAT2) and reporting (MultiQC)
- Run a workflow with Nextflow and its different profiles (local, Docker, HPC cluster)
- Interpret the meaning of a volcano plot in a differential expression context
- Turn ad-hoc scripts into reusable command-line tools (
argparse) - Structure analysis code so it can be re-run on new data
- Reference the reference solutions in
solutions/when stuck on exercises
Reference implementations for the exercises are in solutions/:
download_tsv_uniprot.py— download records from UniProt and save as TSVgc_content.py— compute GC content from FASTAplotly_layout.py,plotly_radioitems.py— plotly chart/layout tricksvolcano_dash.py— a complete interactive volcano-plot Dash app (and seeimages/volcano_dash.giffor what it looks like)
Try the exercises yourself before consulting these.
- Python plotting scripts use
matplotlib.use("Agg")for headless rendering, so they work without a display server. - UniProt input files use the alternating name/URL plain-text format (not JSON).
- Nextflow pipeline behaviour is edited in
rnaseq/modules/, not inrnaseq.nf. - The notebooks are numbered and build on each other; complete them in order.
