Evidence-guided paper · 2021 · outside-scope · full-text
CirPred, the first structure modeling and linker design system for circularly permuted proteins
Teng-Ruei Chen; Yen-Cheng Lin; Yu-Wei Huang; Chih-Chieh Chen; Wei-Cheng Lo. CirPred, the first structure modeling and linker design system for circularly permuted proteins. BMC Bioinformatics 22:494 (2021).
30-second read
CirPred handles CP structure modeling, ordinary collinear modeling, new-termini linker design, and CP-induced 3D domain swapping.
Central question
How can one build complete models of circularly permuted proteins, design linkers between old termini, and handle CP-induced domain reorientation?
Intuition
First rearrange the native template at both sequence and PDB levels into a pseudo-CP template, then perform comparative modeling; if domain orientation is wrong, use the CP site as a hinge, and if old termini are too distant, design a linker with a predict-and-refine strategy.
Why it matters
Conventional modeling tools often omit a CP-delimited region, stretch it into a long coil, or misorient domains. CirPred incorporates CP-specific topology, linker design, and domain swapping to reduce protein-engineering trial and error.
Prerequisites
- Understand circular permutation as sequence rearrangement
- Understand template-based comparative modeling and RMSD/coverage
- Know peptide linkers, DOPE scores, and 3D domain swapping
paper-specific guide · plain → technical → input → output → source
Method walkthrough
-
01 · Create a pseudo-CP template
At the CP site, move the native N-terminal fragment after the C-terminal fragment, rearrange sequence and PDB residue/atom numbering, and restore missing atoms.
Technical reading: reduce/teLeap restore atoms; target-to-pseudo-template alignment is selected from (PS)2, Smith-Waterman, and Stretcher by the largest aligned-residue count.
Input: A native structure, CP site, and optional target-CPM sequence
Output: A template/alignment whose topology order matches the target
Boundary: Template quality and alignment error still constrain low-identity cases.
PDF pp. 3 and 15–16, Fig. 1, 'Generating the pseudo...' and 'Comparative structure modeling'
-
02 · Design the old-termini linker
Estimate linker length and positions, predict residue candidates with machine learning, repeatedly model random candidates, and screen by DOPE/energy.
Technical reading: Training uses CPDB linker data; web mode 3 returns 30 candidate linkers/models, leaving final selection to protein expertise. Initial per-residue prediction accuracy is about 67.5%, with refinement improving utility.
Input: Native-termini geometry, linker length, and CPDB-derived propensities
Output: Energy-ranked linker sequences and CPM models
Boundary: The sub-angstrom result comes from particular redesign cases and is not a guarantee for every new linker.
PDF pp. 5–6, 11–16, Fig. 3 and Linker design Methods
-
03 · Refine domains around the CP-site hinge
If a coarse model has incorrect domain orientation, the algorithm searches orientations around the CP-site hinge, followed by energy minimization and optional MD.
Technical reading: This refinement correctly models the domain-swapped open form of betaB2-crystallin CPM87; a web query without MD typically takes under three minutes.
Input: A comparative coarse model and CP-delimited domains
Output: An orientation-refined, energy-minimized CPM model
Boundary: Capturing known cases does not mean every CP-induced conformational change can be represented as a hinge rotation.
PDF pp. 3, 9–13, Figs. 1 and 4, Discussion
Key result
The paper reports sub-ångström linker-design accuracy and support for low-identity and domain-swapped conformations.
Evidence-guided deep reading
Paper facts, project readings, and teaching models are labelled separately.
paper-fact
Why conventional modeling fails on CP
In the viable-DHFR-CPM scan, SWISS-MODEL omits a region, RaptorX produces a long coil, and Robetta builds two regions with incorrect orientation; CP changes the sequence start point and contiguous order.
CirPred is the only compared method to model all viable DHFR CPMs completely and correctly; setting the CP site to residue 1 recovers conventional collinear modeling.
Source locator: PDF pp. 2–5, Figs. 1–2 and Results 'Comparison...'
paper-fact
Large-scale performance and identity boundary
Engineered CPMs show a mean model-to-actual alignment ratio of 99.2% and RMSD 1.59 A; among 1,568 CPDB pairs, identity-at-least-20% groups average over 90% coverage and generally under 2.5 A RMSD.
The below-20%-identity group drops to average 74.0% coverage and 3.90 A RMSD, clearly marking template alignment as a failure boundary.
Source locator: PDF pp. 6–7, Tables 1–2
paper-fact
Linker accuracy and domain swapping are separate tasks
In beta-glucanase redesign, CirPred builds a structurally and sequentially similar candidate for a deleted 17-residue linker; for cases with termini under 10 A apart, the paper reports 0.26 A linker accuracy.
BetaB2-crystallin CPM87 tests domain orientation: SWISS-MODEL resembles the native closed form, while CirPred aligns to the actual domain-swapped CPM at 99.4% and 3.61 A RMSD. These metrics should not be collapsed into one 'accuracy.'
Source locator: PDF pp. 5–6 and 9–12, Figs. 3–4, linker/DS Discussions
Study design and evaluation
Data and samples
Viable DHFR CPMs, engineered CPM literature cases, 1,568 CPDB pairs at at least 10% identity, a CPDB linker set, and synthetic Dataset S with 1,802 proteins.
Baselines
- SWISS-MODEL, RaptorX, and Robetta; actual CPM structures; native or reconstructed linkers
Metrics
- alignment ratio
- The fraction of residues well matched between model and actual structure.
Boundary: High coverage does not guarantee low coordinate error. - RMSD
- Average spatial deviation of aligned atoms after superposition.
Boundary: Depends on selected/aligned residues and must be read with coverage. - linker sequence/structure accuracy
- Sequence similarity and geometric deviation of a designed linker against a known linker.
Boundary: A redesign benchmark is not expression/function success in a novel protein.
Reported result
CirPred resolves CP-topology failures of conventional tools, maintains high coverage/low RMSD for large-scale cases at at least 20% identity, and can design linkers and capture known CP-induced domain swapping.
PDF pp. 2–16, Figs. 1–5, Tables 1–2 and Methods
teaching-model · not a reported experiment
Teaching example (not a reported experiment)
Two-metric acceptance for a CPM model
Teaching model: model A has 98% coverage and RMSD 4.0 A, while model B has 75% coverage and RMSD 1.2 A.
- Ask the use first: whole-scaffold engineering may prioritize coverage, while an active-site study may prioritize local RMSD.
- Check whether the unaligned 25% contains the linker, CP site, or domain interface.
- Do not select a model from one number; add energy, stereochemistry, and wet-lab viability.
Takeaway: Coverage and RMSD are complementary coordinates; engineering utility also requires local-risk analysis and experiments.
outside-scope
Evidence boundary versus FAST
The primary outputs are CP models and linkers, not a same-benchmark comparison against FAST.
Lawful source and access
23 pages · SHA-256 deddb8974a8aabd60fd24d74cff7fabc940ca15b21fc2c629079036c18680c3a
Lawful open full text.
Limits and misreadings
- Model and linker accuracy depend on templates, target conformations, and experimental case distribution.
Source locator map
- PDF pp. 2–7, Figs. 1–3 and Tables 1–2
- PDF pp. 9–13, Fig. 4 and Discussion
- PDF pp. 13–16, Methods and Fig. 5
Check understanding
What does the pseudo-CP template solve?
Answer: It makes template residue order topologically compatible with the CPM target.
Ordinary collinear alignment otherwise often breaks at the CP boundary.
Is sub-angstrom linker accuracy universally guaranteed?
Answer: No.
It comes from particular redesign benchmarks and termini geometry.
What happens below 20% identity?
Answer: Mean coverage drops to 74% and RMSD rises to 3.90 A.
Target-template alignment becomes a major limit.
Completion task: Build a model-acceptance sheet for one CP-engineering target with columns for coverage, global/local RMSD, linker geometry, domain orientation, energy, and validation experiments.
Paper-specific glossary
- pseudo-CP template
- A template representation reordered at a CP site while retaining the native 3D scaffold.
- DOPE score
- A statistical potential used to rank structural plausibility in comparative modeling.
- domain swapping
- A conformational phenomenon in which protein subunits exchange equivalent structural parts to form oligomers.
- alignment ratio
- The fraction of residues included in a good structural correspondence.