Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Python for Bioinformatics — Workshop Materials

BIgMAG workflow overview

This repository contains the complete learning material for the Python for Bioinformatics workshop. It is organised as a sequence of hands-on Jupyter notebooks that take you from querying public biological databases, through downloading and parsing sequencing data, to running a full RNA-seq analysis pipeline and building an interactive data app.

The notebooks are sequential — each one builds on the concepts and techniques introduced in the previous ones.


Getting started

Click on the following link to create your GitHub Codespace:

Open in GitHub Codespaces

Once the Codespace is created, run the following in the terminal to install some missing packages:

 pip install biopython
 pip install dash
 pip install dash-bio
 pip install dash-bootstrap-components

You are ready to go 🚀!


Workshop structure

The material is organised into five main components:

Directory Description
notebooks/ The 5 core workshop notebooks (covered step by step below)
automatize/ Reusable Python tools for downloading data from UniProt and plotting sequence distributions
rnaseq/ A Nextflow RNA-seq pipeline (FASTQC → Trim Galore → HISAT2 → MultiQC)
scripts/ Shell helpers, e.g. for downloading raw reads from SRA
data/ Sample datasets used throughout the workshop (reads, CSV tables, references)
solutions/ Reference implementations for the workshop exercises

The notebooks

  1. notebook1_uniprot.ipynb — Query and bulk-download protein sequences from the UniProt REST API
  2. notebook2_sra.ipynb — Fetch raw sequencing data from the NCBI Sequence Read Archive (SRA)
  3. notebook3_rnaseq.ipynb — Explore an RNA-seq differential expression analysis
  4. notebook4_dash.ipynb — Build interactive visualisation apps with Plotly Dash
  5. notebook5_project.ipynb — Consolidate everything into a mini research project

Learning outcomes

By the end of this workshop you should be able to:

Working with public bioinformatics databases

  • Query biological databases through REST APIs (with the UniProt API as the worked example)
  • Perform bulk downloads of sequences and parse paginated API responses
  • Understand the difference between common formats such as FASTA, FASTQ and TSV
  • Fetch raw sequencing reads from NCBI SRA using the SRA Toolkit

Handling biological data with Python

  • Parse and manipulate FASTA/FASTQ sequence files with Biopython
  • Load, filter and explore tabular datasets with pandas
  • Extract meaningful structure from real datasets (e.g. a 19th-century census table)

Visualising results

  • Produce publication-style static charts with matplotlib (histograms, distributions)
  • Build rich interactive figures with plotly
  • Turn a figure into a shareable static HTML file

Building interactive data apps

  • Design multi-component apps with Plotly Dash
  • Connect charts to user interaction through Dash callbacks
  • Critically evaluate when a full server-based app is worth the added complexity (vs. a static plotly HTML export)
  • Be aware of the deployment and hosting options for Dash apps

Running bioinformatics pipelines

  • Understand what an RNA-seq analysis pipeline does at each stage: quality control (FASTQC), adapter trimming (Trim Galore), alignment (HISAT2) and reporting (MultiQC)
  • Run a workflow with Nextflow and its different profiles (local, Docker, HPC cluster)
  • Interpret the meaning of a volcano plot in a differential expression context

Reproducibility and best practice

  • Turn ad-hoc scripts into reusable command-line tools (argparse)
  • Structure analysis code so it can be re-run on new data
  • Reference the reference solutions in solutions/ when stuck on exercises

Solutions

Reference implementations for the exercises are in solutions/:

  • download_tsv_uniprot.py — download records from UniProt and save as TSV
  • gc_content.py — compute GC content from FASTA
  • plotly_layout.py, plotly_radioitems.py — plotly chart/layout tricks
  • volcano_dash.py — a complete interactive volcano-plot Dash app (and see images/volcano_dash.gif for what it looks like)

Try the exercises yourself before consulting these.


Repository conventions

  • Python plotting scripts use matplotlib.use("Agg") for headless rendering, so they work without a display server.
  • UniProt input files use the alternating name/URL plain-text format (not JSON).
  • Nextflow pipeline behaviour is edited in rnaseq/modules/, not in rnaseq.nf.
  • The notebooks are numbered and build on each other; complete them in order.

About

Learning material for the workshop Python for Bioinformatics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages