Evidence-guided paper · 2021 · outside-scope · full-text

CirPred, the first structure modeling and linker design system for circularly permuted proteins

Teng-Ruei Chen; Yen-Cheng Lin; Yu-Wei Huang; Chih-Chieh Chen; Wei-Cheng Lo. CirPred, the first structure modeling and linker design system for circularly permuted proteins. BMC Bioinformatics 22:494 (2021).

30-second read

CirPred handles CP structure modeling, ordinary collinear modeling, new-termini linker design, and CP-induced 3D domain swapping.

Central question

How can one build complete models of circularly permuted proteins, design linkers between old termini, and handle CP-induced domain reorientation?

Intuition

First rearrange the native template at both sequence and PDB levels into a pseudo-CP template, then perform comparative modeling; if domain orientation is wrong, use the CP site as a hinge, and if old termini are too distant, design a linker with a predict-and-refine strategy.

Why it matters

Conventional modeling tools often omit a CP-delimited region, stretch it into a long coil, or misorient domains. CirPred incorporates CP-specific topology, linker design, and domain swapping to reduce protein-engineering trial and error.

Prerequisites

  • Understand circular permutation as sequence rearrangement
  • Understand template-based comparative modeling and RMSD/coverage
  • Know peptide linkers, DOPE scores, and 3D domain swapping

paper-specific guide · plain → technical → input → output → source

Method walkthrough

  1. 01 · Create a pseudo-CP template

    At the CP site, move the native N-terminal fragment after the C-terminal fragment, rearrange sequence and PDB residue/atom numbering, and restore missing atoms.

    Technical reading: reduce/teLeap restore atoms; target-to-pseudo-template alignment is selected from (PS)2, Smith-Waterman, and Stretcher by the largest aligned-residue count.

    Input: A native structure, CP site, and optional target-CPM sequence

    Output: A template/alignment whose topology order matches the target

    Boundary: Template quality and alignment error still constrain low-identity cases.

    PDF pp. 3 and 15–16, Fig. 1, 'Generating the pseudo...' and 'Comparative structure modeling'

  2. 02 · Design the old-termini linker

    Estimate linker length and positions, predict residue candidates with machine learning, repeatedly model random candidates, and screen by DOPE/energy.

    Technical reading: Training uses CPDB linker data; web mode 3 returns 30 candidate linkers/models, leaving final selection to protein expertise. Initial per-residue prediction accuracy is about 67.5%, with refinement improving utility.

    Input: Native-termini geometry, linker length, and CPDB-derived propensities

    Output: Energy-ranked linker sequences and CPM models

    Boundary: The sub-angstrom result comes from particular redesign cases and is not a guarantee for every new linker.

    PDF pp. 5–6, 11–16, Fig. 3 and Linker design Methods

  3. 03 · Refine domains around the CP-site hinge

    If a coarse model has incorrect domain orientation, the algorithm searches orientations around the CP-site hinge, followed by energy minimization and optional MD.

    Technical reading: This refinement correctly models the domain-swapped open form of betaB2-crystallin CPM87; a web query without MD typically takes under three minutes.

    Input: A comparative coarse model and CP-delimited domains

    Output: An orientation-refined, energy-minimized CPM model

    Boundary: Capturing known cases does not mean every CP-induced conformational change can be represented as a hinge rotation.

    PDF pp. 3, 9–13, Figs. 1 and 4, Discussion

Key result

The paper reports sub-ångström linker-design accuracy and support for low-identity and domain-swapped conformations.

Evidence-guided deep reading

Paper facts, project readings, and teaching models are labelled separately.

paper-fact

Why conventional modeling fails on CP

In the viable-DHFR-CPM scan, SWISS-MODEL omits a region, RaptorX produces a long coil, and Robetta builds two regions with incorrect orientation; CP changes the sequence start point and contiguous order.

CirPred is the only compared method to model all viable DHFR CPMs completely and correctly; setting the CP site to residue 1 recovers conventional collinear modeling.

Source locator: PDF pp. 2–5, Figs. 1–2 and Results 'Comparison...'

paper-fact

Large-scale performance and identity boundary

Engineered CPMs show a mean model-to-actual alignment ratio of 99.2% and RMSD 1.59 A; among 1,568 CPDB pairs, identity-at-least-20% groups average over 90% coverage and generally under 2.5 A RMSD.

The below-20%-identity group drops to average 74.0% coverage and 3.90 A RMSD, clearly marking template alignment as a failure boundary.

Source locator: PDF pp. 6–7, Tables 1–2

paper-fact

Linker accuracy and domain swapping are separate tasks

In beta-glucanase redesign, CirPred builds a structurally and sequentially similar candidate for a deleted 17-residue linker; for cases with termini under 10 A apart, the paper reports 0.26 A linker accuracy.

BetaB2-crystallin CPM87 tests domain orientation: SWISS-MODEL resembles the native closed form, while CirPred aligns to the actual domain-swapped CPM at 99.4% and 3.61 A RMSD. These metrics should not be collapsed into one 'accuracy.'

Source locator: PDF pp. 5–6 and 9–12, Figs. 3–4, linker/DS Discussions

Study design and evaluation

Data and samples

Viable DHFR CPMs, engineered CPM literature cases, 1,568 CPDB pairs at at least 10% identity, a CPDB linker set, and synthetic Dataset S with 1,802 proteins.

Baselines

  • SWISS-MODEL, RaptorX, and Robetta; actual CPM structures; native or reconstructed linkers

Metrics

alignment ratio
The fraction of residues well matched between model and actual structure.
Boundary: High coverage does not guarantee low coordinate error.
RMSD
Average spatial deviation of aligned atoms after superposition.
Boundary: Depends on selected/aligned residues and must be read with coverage.
linker sequence/structure accuracy
Sequence similarity and geometric deviation of a designed linker against a known linker.
Boundary: A redesign benchmark is not expression/function success in a novel protein.

Reported result

CirPred resolves CP-topology failures of conventional tools, maintains high coverage/low RMSD for large-scale cases at at least 20% identity, and can design linkers and capture known CP-induced domain swapping.

PDF pp. 2–16, Figs. 1–5, Tables 1–2 and Methods

teaching-model · not a reported experiment

Teaching example (not a reported experiment)

Two-metric acceptance for a CPM model

Teaching model: model A has 98% coverage and RMSD 4.0 A, while model B has 75% coverage and RMSD 1.2 A.

  1. Ask the use first: whole-scaffold engineering may prioritize coverage, while an active-site study may prioritize local RMSD.
  2. Check whether the unaligned 25% contains the linker, CP site, or domain interface.
  3. Do not select a model from one number; add energy, stereochemistry, and wet-lab viability.

Takeaway: Coverage and RMSD are complementary coordinates; engineering utility also requires local-risk analysis and experiments.

outside-scope

Evidence boundary versus FAST

The primary outputs are CP models and linkers, not a same-benchmark comparison against FAST.

Lawful source and access

23 pages · SHA-256 deddb8974a8aabd60fd24d74cff7fabc940ca15b21fc2c629079036c18680c3a

Lawful open full text.

Europe PMC open-access PDF

Limits and misreadings

  • Model and linker accuracy depend on templates, target conformations, and experimental case distribution.

Source locator map

  1. PDF pp. 2–7, Figs. 1–3 and Tables 1–2
  2. PDF pp. 9–13, Fig. 4 and Discussion
  3. PDF pp. 13–16, Methods and Fig. 5

Check understanding

  1. What does the pseudo-CP template solve?

    Answer: It makes template residue order topologically compatible with the CPM target.

    Ordinary collinear alignment otherwise often breaks at the CP boundary.

  2. Is sub-angstrom linker accuracy universally guaranteed?

    Answer: No.

    It comes from particular redesign benchmarks and termini geometry.

  3. What happens below 20% identity?

    Answer: Mean coverage drops to 74% and RMSD rises to 3.90 A.

    Target-template alignment becomes a major limit.

Completion task: Build a model-acceptance sheet for one CP-engineering target with columns for coverage, global/local RMSD, linker geometry, domain orientation, energy, and validation experiments.

Paper-specific glossary

pseudo-CP template
A template representation reordered at a CP site while retaining the native 3D scaffold.
DOPE score
A statistical potential used to rank structural plausibility in comparative modeling.
domain swapping
A conformational phenomenon in which protein subunits exchange equivalent structural parts to form oligomers.
alignment ratio
The fraction of residues included in a good structural correspondence.