A Python implementation of the DynDom algorithm for analyzing protein domain movements between conformational states.
DynDom1D_Python is a Python implementation of the DynDom algorithm for analyzing protein conformational changes in terms of rigid-body domain movements. The software determines protein domains, hinge axes, and amino acid residues involved in hinge bending through fully automated analysis.
Given two conformational states of the same protein (e.g., from X-ray crystallography, NMR, or molecular dynamics simulations), DynDom1D_Python identifies how the protein moves by treating different regions as quasi-rigid bodies. This approach transforms complex conformational changes into an easily understood view of domain movements connected by flexible hinges.
The analysis reveals:
- Dynamic domains: Regions that move as quasi-rigid bodies
- Hinge regions: Flexible areas that allow domain movement
- Screw axes: Mathematical description of how domains rotate and translate relative to each other
- Automated Domain Detection: Identifies quasi-rigid domains using rotation vector clustering
- Hierarchical Analysis: Uses advanced hierarchical domain reference system for complex multi-domain proteins
- Hinge Analysis: Locates flexible regions and mechanical hinges between domains
- Screw Axis Calculation: Determines interdomain rotation axes, angles, and translations
- Motion Validation: Applies geometric constraints and interdomain/intradomain motion ratios
- Comprehensive Visualization: Generates PyMOL scripts with domain coloring and 3D motion arrows
- Flexible Input: Supports any two conformational states in PDB format
- Detailed Output: Provides motion statistics, hinge residue identification, and structural files
- Quality Control: Validates domain connectivity and motion significance
- Clustering Logs: Detailed logging of the clustering process for debugging and analysis
The article associated with this tool has been published in JOSS:
- Python 3.7 or higher
- Required packages listed in
requirements.txt - (Optional) PyMol (paid or open source). Some installation options are here.
python -m pip install pymol-open-sourcesometimes also works, or for conda/mamba usersmamba install -c conda-forge open-source-pymolworks more reliably.
-
Clone the repository:
git clone https://github.com/steven-hayward/DynDom1D_Python.git cd DynDom1D_Python/src/DynDom1D_Python -
Create a virtual environment (recommended):
python -m venv dyndom-env source dyndom-env/bin/activate # On Windows: dyndom-env\Scripts\activate
-
Install dependencies:
pip install -r requirements.txt
Or install manually:
pip install numpy scipy scikit-learn gemmi
-
Verify installation:
python -c "import numpy, scipy, sklearn, gemmi; print('All dependencies installed successfully')" -
Set up directories:
mkdir -p data output
-
Prepare input files edit the
command.txtfile in thedata/directory:data/command.txt:
input_path=data output_path=output filename1=4ake chain1id=A filename2=2eck chain2id=Bdata/param.txt:
window=5 domain=20 ratio=1.0 k_means_n_init=1 k_means_max_iter=500 atoms=backbone -
Prepare PDB files:
- Place your PDB files in the
data/directory, or - Use PDB codes (e.g.,
4ake,2eck) and the software will automatically download them from RCSB
- Place your PDB files in the
-
Run the analysis:
- Run main.py from the
DynDom1D_Python/src/directory
python main.py
- Run main.py from the
input_path: Directory containing PDB filesoutput_path: Directory for output filesfilename1,filename2: PDB IDs or filenames (without .pdb extension)chain1id,chain2id: Chain identifiers
window: Sliding window size for local motion analysis (default: 5)domain: Minimum domain size in residues (default: 20)ratio: Minimum ratio of interdomain to intradomain motion (default: 1.0)atoms: Atoms to use for analysis (backbonefor N,CA,C orcafor CA only)k_means_n_init: Number of K-means initializations (default: 1)k_means_max_iter: Maximum K-means iterations (default: 500)
The software generates several output files in the specified output directory:
{protein1}_{chain1}_{protein2}_{chain2}.pdb: Superimposed structures (two models){protein1}_{chain1}_{protein2}_{chain2}_arrows.pdb: 3D motion arrows for PyMOL visualization{protein1}_rot_vecs.pdb: Rotation vectors for each residue
{protein1}_{chain1}_{protein2}_{chain2}.pml: PyMOL visualization script with:- Domain coloring (different colors for each domain)
- Hinge regions highlighted in green
- 3D motion arrows showing screw axes
- Automatic camera positioning and lighting
{protein1}_{chain1}_{protein2}_{chain2}.w5_info: Detailed analysis results including:- Domain definitions and residue ranges
- Rotation angles and screw axis parameters
- Motion statistics and RMSD values
- Hierarchical domain relationships
clustering_logs/: Directory containing detailed clustering logsdomain_*_comparison.pdb: Individual domain motion comparison files (debug mode)
-
Load in PyMOL:
pymol output/{protein1}_{chain1}_{protein2}_{chain2}.pml -
Visualization features:
- Domain coloring: Different domains shown in different colors
- Global reference domain: Shown in blue
- Moving domains: Shown in red, yellow, pink, etc.
- Hinge regions: Flexible regions highlighted in green
- Motion arrows: 3D arrows showing screw axes and rotation directions
- Arrow coloring: Shaft shows reference domain color, tip shows moving domain color
DynDom1D_Python implements a 7-step automated workflow with hierarchical domain analysis:
- Global Structure Alignment: Performs whole-protein best-fit superposition
- Local Motion Detection: Analyzes rotation vectors using sliding windows
- Rotation Vector Clustering: Uses adaptive K-means clustering with validation
- Hierarchical Domain Construction: Builds domains with connectivity analysis
- Domain Validation: Applies size and motion ratio criteria
- Hierarchical Screw Axis Calculation: Determines motion parameters for each domain pair
- Hinge Residue Identification: Locates flexible regions using statistical analysis
- Hierarchical Analysis: Domains are analyzed relative to their most appropriate reference domain
- Global Reference Domain: The most connected domain serves as the global reference
- Analysis Pairs: Each domain is analyzed relative to its optimal reference domain
- Screw Axis: Mathematical description of rigid body motion (rotation + translation)
- Bending Residues: Residues with rotations outside the main domain distribution
-
"Sequence Identity less than 40%":
- Check that both PDB files contain the same protein
- Verify chain IDs are correct
-
"Too many fails. Increasing window size":
- Normal behavior - algorithm automatically adjusts parameters
-
No domains found:
- Reduce
domainparameter (try 15 or 10) - Reduce
ratioparameter (try 0.8 or 0.6) - but be aware the results may not be meaningful for smaller values of the ratio (see paper by Hayward and Berendsen below). - Increase
windowparameter
- Reduce
Original DynDom algorithm:
Hayward, S. and Berendsen, H.J.C. (1998) Systematic analysis of domain motions in proteins from conformational change: new results on citrate synthase and T4 lysozyme. Proteins 30, 144-154.
S. Hayward, R. A. Lee
"Improvements in the analysis of domain motions in proteins from conformational change: DynDom version 1.50" J Mol Graph Model, Dec, 21(3), 181-3, 2002.
Contributions are welcome! See CONTRIBUTING.md for how to report issues, suggest features, or submit a pull request.
BSD 3-Clause License Refer to LICENSE file