Evidence-guided paper · 2009 · historical-ecosystem · full-text
CPDB: a database of circular permutation in proteins
Wei-Cheng Lo; Chi-Ching Lee; Che-Yu Lee; Ping-Chiang Lyu. CPDB: a database of circular permutation in proteins. Nucleic Acids Research 37:D328–D332 (2009).
30-second read
CPDB turns CPSARST screening, quality tiers, and manual verification into a searchable circular-permutation database.
Central question
How can scattered protein circular permutations—often missed by ordinary sequential alignment—be organized into a searchable, evidence-traceable database that supports new hypotheses?
Intuition
CPDB does not simply dump search hits into a table. CPSARST first screens candidates at high throughput, manual review removes false cases, exhaustive breakpoint shifts with FAST refine them, and accepted pairs are organized into clusters, folds, and classes with network and circular-alignment views.
Why it matters
A reliable database is simultaneously a query tool, benchmark, and hypothesis generator. CPDB gives CP prevalence, evolution, viable-breakpoint, and protein-engineering studies a shared large-scale data foundation.
Prerequisites
- Understand circular permutation and a CP pair.
- Know that automated candidate search and manual curation play different roles.
- Understand direct/indirect relationships and connected components in a graph.
- Distinguish database coverage, prediction hit rate, and alignment quality.
paper-specific guide · plain → technical → input → output → source
Method walkthrough
-
01 · Database-wide screening and manual curation
Use CPSARST to find possible CPs in the PDB, then inspect them manually and remove false cases.
Technical reading: All-against-all CPSARST search is run over 26,349 nonredundant PDB polypeptides. Candidates undergo visual inspection, then permutation-site shifting and possible-CP alignment enumeration with FAST as the structural-alignment engine.
Input: A nonredundant PDB polypeptide set.
Output: 4,169 curated CP pairs, 2,238 proteins, and refined CP sites.
Boundary: This is a collection under the contemporary PDB snapshot and decision rules, not the complete population of CPs in nature.
PDF p. 2, Contents and Methods — Identification of CP
-
02 · Organize pairs into clusters, folds, and classes
Group proteins linked directly or indirectly, then elevate structurally similar groups into folds and broad classes.
Technical reading: Connected proteins in the CP graph form a cluster. FAST similarities among representatives feed nearest-neighbor clustering plus manual adjustment to form folds, which are classified by secondary-structure content into mainly-α, mainly-β, and mixed α–β classes.
Input: The curated CP-pair graph and representative structures.
Output: A browsable cluster/fold/class hierarchy.
Boundary: An indirect relationship within a cluster does not mean every pair of nodes has equally strong direct CP evidence.
PDF p. 2, Categorization of circular permutants
-
03 · Expose evidence through exploratory views
Let users inspect termini, breakpoints, alignment text, structures, and relationship networks instead of a single score.
Technical reading: The interface provides circularized sequence/structure alignments, CP networks, structural-diversity-scaled star maps, protein pages, CPSARST/SARST search, and a literature list.
Input: Curated pairs, metadata, alignments, and structural coordinates.
Output: A web resource whose evidence can be inspected and traced.
Boundary: Visualization aids interpretation but is not an independent validation dataset.
PDF pp. 2–4, Figure 1, Web Interface and Figure 2
Key result
Its main contribution is a reusable benchmark and knowledge base, not a new pairwise superposition optimizer.
Evidence-guided deep reading
Paper facts, project readings, and teaching models are labelled separately.
paper-fact
Database quality comes from layered filtering
CPDB starts from 26,349 nonredundant PDB polypeptides. CPSARST provides scalable screening, manual visual inspection removes false cases, and a theoretically more accurate breakpoint-shifting procedure with FAST performs final refinement.
The result is 4,169 pairs among 2,238 proteins. This count is a product of the curation pipeline and cannot directly be treated as natural CP prevalence because sampling, detectability, and definitions all shape the denominator.
Source locator: PDF p. 2, Identification of CP
paper-fact
A relationship graph answers more than a flat list
A pair graph places A↔B and B↔C in one cluster even when A and C are related only indirectly; fold and class levels situate the local network in broader structural context.
Figure 1's circular alignment uses terminal spheres and color boundaries to reveal CP sites, the network view shows cluster connectivity, and the star map uses structural diversity to present a query's CP and linear homologs.
Source locator: PDF pp. 2–3, Figure 1
project-reading
Treat the database as an evidence layer, not universal truth
CPDB is useful for generating candidates, tracing cases, and building benchmarks; each relationship should still be read with its source structures, alignment, breakpoint, and curation method.
The paper states that the current release favors global CP and undercovers partial CP. When evaluating a new method, unrecorded database cases must not automatically be treated as true negatives.
Source locator: PDF p. 4, Future Works
Study design and evaluation
Data and samples
A set of 26,349 nonredundant PDB polypeptides undergoes CPSARST all-against-all search, manual inspection, and FAST breakpoint refinement; nonredundant CP sites in the database also evaluate closeness- and RSA-based breakpoint predictors.
Baselines
- Prior sequence- and structure-based CP detection (Uliel, SHEBA, SAMO, FASE) frames the resource gap; viable-breakpoint analysis compares closeness and relative side-chain area (RSA).
Metrics
- Curated pair/protein count
- The number of CP pairs and proteins accepted by the full pipeline.
Boundary: This is database yield, not alignment accuracy or natural prevalence. - CP-site hit rate
- Whether a residue measure identifies known nonredundant CP sites as viable positions.
Boundary: This is positive-case coverage; it lacks equally complete specificity over inviable sites. - Structural diversity
- A composite structural difference used to scale query–homolog relationships in star maps.
Boundary: It is a display/ranking quantity and cannot alone establish an evolutionary mechanism.
Reported result
CPDB records 4,169 nonredundant CP pairs among 2,238 proteins and organizes them into a browsable hierarchy. After adding hydrogens, closeness hits 66.5% of nonredundant CP sites and RSA 60.9%.
PDF pp. 2–3, Identification of CP and Prediction of viable circular permutants
teaching-model · not a reported experiment
Teaching example (not a reported experiment)
Build a cluster from three CP pairs
Project teaching model: curation accepts A↔B, B↔C, and D↔E, while A↔C lacks a direct accepted edge.
- Treat proteins as nodes and accepted CP pairs as edges.
- Connected components yield cluster {A,B,C} and cluster {D,E}.
- Compare representative structures across clusters; if structurally similar, place them in one fold while preserving the original direct edges.
Takeaway: A hierarchy compresses complex relationships, but direct-pair evidence and indirect membership must remain distinct.
historical-ecosystem
Evidence boundary versus FAST
The database inherits the CPSARST/FAST ecosystem, but database scale is not evidence of better alignment quality than FAST.
Lawful source and access
5 pages · SHA-256 3259418dbdb0a4e0fad07ec84696f14de602a63d0c07dd1c6442e5f79b9b2ec7
The formal issue is dated 2009; online publication occurred in 2008.
Limits and misreadings
- Coverage is bounded by the contemporary PDB release and decision thresholds.
Source locator map
- PDF p. 1, Abstract and Introduction
- PDF p. 2, Identification of CP
- PDF p. 2, Categorization of circular permutants
- PDF pp. 2–3, Figure 1 and Prediction of viable circular permutants
- PDF pp. 3–4, Web Interface, Figure 2 and Future Works
Check understanding
What does CPDB's number 4,169 count?
Answer: Nonredundant CP pairs accepted after CPSARST, manual review, and FAST refinement.
It is not an unbiased prevalence estimate for CP in the PDB.
Must A and C in one cluster share a direct CP edge?
Answer: No. They may be connected indirectly through B.
Cluster membership and pair-level evidence are different levels.
Can the 66.5% closeness result be read as 66.5% overall accuracy?
Answer: No. It is a hit rate over known nonredundant CP sites.
Without a complete negative set, specificity and overall accuracy do not follow.
Completion task: Design a modern CPDB evidence card containing at least source PDBs, direct/indirect relation, alignment, CP site, curation status, and database version, then explain which field prevents treating a missing record as a negative.
Paper-specific glossary
- Curation
- Human or semi-automated review, correction, and acceptance under explicit rules.
- CP cluster
- A connected component of proteins linked by direct or indirect CP edges.
- Fold
- A higher level grouping structurally similar CP clusters.
- Closeness
- A residue-network proximity measure used to estimate breakpoint viability.