Evidence-guided paper · 2009 · historical-ecosystem · full-text

CPDB:蛋白質環狀置換資料庫

Wei-Cheng Lo; Chi-Ching Lee; Che-Yu Lee; Ping-Chiang Lyu. CPDB: a database of circular permutation in proteins. Nucleic Acids Research 37:D328–D332 (2009).

30 秒理解

CPDB 把 CPSARST 大規模搜尋、品質分級與人工核對整理成可查詢的環狀置換關係資料庫。

核心問題

如何把分散、難以用一般順序比對找到的 protein circular permutations,整理成可查詢、可追蹤證據且可支援新假說的資料庫?

直覺

CPDB 不是把搜尋結果直接倒進資料表;它先由 CPSARST 高吞吐找候選,再人工排除假案例、用逐切點 FAST 精修,最後把 pair 組成 cluster、fold 與 class,並用網路圖與環狀 alignment 呈現關係。

為什麼重要

可靠資料庫同時是查詢工具、基準集與假說生成器。CPDB 讓 CP prevalence、演化機制、可行切點與蛋白工程研究第一次有大型共同資料底座。

閱讀前置

  • 理解 circular permutation 與 CP pair。
  • 知道自動搜尋候選與人工 curation 的角色不同。
  • 理解 graph 中 direct/indirect relationship 與 connected component。
  • 能區分資料庫 coverage、prediction hit rate 與 alignment quality。

paper-specific guide · plain → technical → input → output → source

逐步方法導讀

  1. 01 · 全庫搜尋與人工 curation

    先讓 CPSARST 從 PDB 找出可能 CP,再由人檢視並剔除假案例。

    論文語言: 對 26,349 個 nonredundant PDB polypeptides 做 all-against-all CPSARST search;候選經 visual inspection,接著移動 permutation site、枚舉可能 CP alignments,並以 FAST 作結構 alignment engine 精修。

    輸入: 非冗餘 PDB polypeptide set。

    輸出: 4,169 個 curated CP pairs、2,238 個 proteins 與 refined CP sites。

    邊界: 這是當時 PDB snapshot 與判定規則下的集合,不是自然界 CP 的完整母體。

    PDF p. 2, Contents and Methods — Identification of CP

  2. 02 · 由 pair 組成 cluster、fold 與 class

    先把直接或間接相連的蛋白放進同一群,再依結構相似性把群提升成 fold 與大類。

    論文語言: CP graph 的 connected proteins 構成 cluster;representative proteins 以 FAST 算相似性,再用 nearest-neighbor clustering 與人工調整形成 folds;folds 依 secondary-structure content 分為 mainly-α、mainly-β、α–β mixed。

    輸入: curated CP pair graph 與 representative structures。

    輸出: 可逐層瀏覽的 cluster/fold/class hierarchy。

    邊界: cluster 的間接關係不表示任兩節點都具有同樣強的直接 CP 證據。

    PDF p. 2, Categorization of circular permutants

  3. 03 · 把證據做成可探索介面

    讓使用者同時看到端點、切點、對齊文字、結構與關係網路,而不只看一個分數。

    論文語言: 介面提供 circularized sequence/structure alignment、CP network、以 structural diversity 為邊長的 star map、protein pages、CPSARST/SARST search 與文獻清單。

    輸入: curated pairs、metadata、alignment 與結構座標。

    輸出: 可驗證與追索的 Web 資源。

    邊界: 視覺化協助判讀,不等同新的獨立驗證資料。

    PDF pp. 2–4, Figure 1, Web Interface and Figure 2

關鍵結果

重點是建立可重用的基準與知識庫,而非提出新的 pairwise 疊合最佳化。

逐節證據導讀

論文事實、本站判讀與教學模型分開標示。

paper-fact

資料庫品質來自多層篩選

CPDB 的 source set 為 26,349 個非冗餘 PDB polypeptides。CPSARST 負責 scalable screening,人工 visual inspection 排除 false cases,最後以理論上較精確的逐切點方法與 FAST 精修。

最後得到 4,169 pairs/2,238 proteins。這個數字是 curation pipeline 的產物;不能直接當成 CP 的自然發生率,因為抽樣、檢出能力與定義都參與了分母。

原文定位: PDF p. 2, Identification of CP

paper-fact

關係網比平面清單多回答一層問題

pair graph 讓 A↔B、B↔C 形成同 cluster,即使 A 與 C 只有間接關係;fold 與 class 再把局部網路放進更大的結構分類脈絡。

Figure 1 的環狀 alignment 用端點球與顏色邊界顯示 CP site,network view 顯示 cluster connectivity,star map 則以 structural diversity 呈現 query 的 CP 與 linear homologs。

原文定位: PDF pp. 2–3, Figure 1

project-reading

把資料庫當成證據層,不是宇宙真值

CPDB 適合產生候選、回溯案例與建立 benchmark;每一筆關係仍應連同來源結構、對齊、切點與 curation 方法閱讀。

論文也明列 current release 偏向 global CP,partial CP 的涵蓋不足。使用者若拿 CPDB 評估新方法,必須避免把資料庫未收錄案例直接當成真 negative。

原文定位: PDF p. 4, Future Works

研究設計與評估

資料與樣本

26,349 個 nonredundant PDB polypeptides 經 CPSARST all-against-all、人工檢視與 FAST 切點精修;另以資料庫中的 nonredundant CP sites 評估 closeness 與 RSA 切點 predictor。

比較基準

  • 前期 sequence-based 與 structure-based CP detection(Uliel、SHEBA、SAMO、FASE)構成資源缺口背景;可行切點部分比較 closeness 與 relative side-chain area(RSA)。

指標

Curated pair/protein count
經完整 pipeline 接受的 CP pairs 與涉及 proteins 數。
邊界: 是 database yield,不是 alignment accuracy,也不是自然 prevalence。
CP-site hit rate
某 residue measure 是否把已知 nonredundant CP site 列為可行位置。
邊界: 只有正例覆蓋率;未提供同等完整的 inviable-site specificity。
Structural diversity
在 star map 中表達 query 與 homolog 間的複合結構差異。
邊界: 是呈現與排序量,不能單獨證成 evolutionary mechanism。

論文報告的結果

CPDB 收錄 4,169 個 nonredundant CP pairs、2,238 個 proteins,並把 pair 組織成可瀏覽 hierarchy;加入氫原子後,closeness 對 nonredundant CP sites 的 hit rate 為 66.5%,RSA 為 60.9%。

PDF pp. 2–3, Identification of CP and Prediction of viable circular permutants

teaching-model · not a reported experiment

教學例(不是論文實驗)

從三個 CP pairs 建一個 cluster

本站教學模型:curation 接受 A↔B、B↔C、D↔E;A↔C 沒有直接達標。

  1. 把 proteins 當節點、accepted CP pairs 當邊。
  2. connected-component 分群得到 cluster {A,B,C} 與 cluster {D,E}。
  3. 選代表結構比較兩 clusters;若結構相似,再把它們放進同一 fold,但仍保留原始 direct edges。

帶走什麼: 階層可壓縮複雜關係,但 direct pair evidence 與 indirect membership 必須分開呈現。

historical-ecosystem

相對 FAST 的證據邊界

資料源承接 CPSARST/FAST 生態,但 CPDB 本身是資料庫,不能拿資料庫規模宣稱 alignment quality 勝 FAST。

合法來源與取用

5 pages · SHA-256 3259418dbdb0a4e0fad07ec84696f14de602a63d0c07dd1c6442e5f79b9b2ec7

正式卷期為 2009,線上先於 2008 發表。

Europe PMC open-access PDF

限制與防誤讀

  • 資料內容受當時 PDB 版本與判定閾值限制。

證據定位清單

  1. PDF p. 1, Abstract and Introduction
  2. PDF p. 2, Identification of CP
  3. PDF p. 2, Categorization of circular permutants
  4. PDF pp. 2–3, Figure 1 and Prediction of viable circular permutants
  5. PDF pp. 3–4, Web Interface, Figure 2 and Future Works

理解檢查

  1. CPDB 的 4,169 是什麼數字?

    答案: 經 CPSARST、人工檢視與 FAST 精修後接受的 nonredundant CP pairs。

    不是 PDB 中 CP 的無偏 prevalence estimate。

  2. 同一 cluster 的 A 與 C 一定有 direct CP edge 嗎?

    答案: 不一定;它們可以透過 B 間接相連。

    cluster membership 與 pair-level evidence 是不同層級。

  3. closeness 66.5% 能否解讀為 66.5% overall accuracy?

    答案: 不能;它是已知 nonredundant CP sites 的 hit rate。

    沒有完整 negative set 時,無法由此得到 specificity 或 overall accuracy。

讀完標準: 替 CPDB 設計一張現代 evidence card:至少包含來源 PDB、direct/indirect 關係、alignment、CP site、curation 狀態與資料庫版本,並說明哪一欄可防止把 missing record 當 negative。

本篇詞彙表

Curation
依明確規則人工或半人工檢查、修正與接受資料。
CP cluster
由 direct 或 indirect CP edges 連成的 protein connected component。
Fold
把結構相似的 CP clusters 再向上聚合的層級。
Closeness
由 residue 在結構互動網路中的鄰近程度推估切點可行性的量。