Evidence-guided paper · 2023 · outside-scope · full-text
SeqCP:只用序列搜尋環狀置換蛋白
Chi-Chun Chen; Yu-Wei Huang; Hsuan-Cheng Huang; Wei-Cheng Lo; Ping-Chiang Lyu. SeqCP: A sequence-based algorithm for searching circularly permuted proteins. Computational and Structural Biotechnology Journal 21:185–201 (2023).
30 秒理解
SeqCP 把查詢序列複製後與正常查詢並行比對,從邊界位移、覆蓋、gap 與改善率辨認 CP 關係。
核心問題
沒有已知 3D structure 時,如何只用 normal/duplicated sequence alignments 找 circularly permuted proteins 並定位 CP site?
直覺
把 query 複製成 query+query:若 target 是 circular permutant,它的兩段可跨 duplicated-query 接縫連續對上;再比較 normal 與 duplicated alignment 的邊界位移、coverage、gaps 與改善幅度,就能把 CP pattern 與普通 similarity 分開。
為什麼重要
structure-based CP search 無法處理大量只有 sequence 的 proteins。SeqCP 以 sequence database search 擴大 discovery space,約比 CPSARST 快33倍;但 sequence classifier 的 AUC/threshold 不等於 structure alignment quality。
閱讀前置
- 理解 circular permutation 與 duplicated-query trick
- 理解 BLAST local alignment、identity、similarity 與 gaps
- 知道 sensitivity、specificity、MCC、ROC AUC
paper-specific guide · plain → technical → input → output → source
逐步方法導讀
-
01 · normal 與 duplicated query 雙搜尋
同一 target database 分別以 q 與 qq′ 搜尋;只有兩種 query 都能 align 的 targets 進候選池。
論文語言: blastp 採 word size 2、BLOSUM45、gap open 10/extend 3 等低 identity 設定;CP sites 由 qq′ alignment 的 first/second-copy boundary positions 計算。
輸入: query sequence 與 FASTA target database
輸出: 有 normal/duplicated alignments 與 candidate CP sites 的 hits
邊界: 極低 identity 或短 fragments 可能無法產生可辨認的兩段 alignment。
PDF pp. 2–5, Methods §§2.1–2.4, Figs. 1–2
-
02 · 九條規則做初篩
用 protein length ratio、可定位 CP site、normal/duplicate boundary shifts、identity consistency 與 duplicate alignment rate 移除不像 CP 的 pairs。
論文語言: 例如 average length≥100、short/long ratio≥0.8、兩種 boundary shifts≥20%、duplicate alignment rate≥0.8;九項 criteria 在 CPDB 中由2555 marginal 降到22,保留472/480 high-quality。
輸入: 雙 alignment 的 geometry 與 quality features
輸出: 小而高 recall 的 CP candidate pool
邊界: thresholds 由 CPDB-derived distributions 設定,可能承襲其 domain/structure bias。
PDF pp. 5–10, Methods §§2.6–2.7, Table 2 and Fig. 5
-
03 · 七項特徵評分與可調 threshold
把 duplicate coverage、gap density、similarity、相對 normal alignment 的 identity/similarity/coverage 改善與 stability factor 合成 SeqCP score。
論文語言: CPDB training 估 weights;2007 test AUC 0.89,2019 test AUC 0.88,independent homology-reduced test AUC 0.81。推薦 threshold 0.60;0.75 可犧牲 recall 換更高 precision。
輸入: 初篩後 candidates 的七項 features
輸出: ranked CPMs、candidate CP sites、CP identities 與 confidence
邊界: score 是 training-distribution 下的 discriminator;dataset shift 需重新 calibration。
PDF pp. 6–11 and 15–16, Eq. 15, Fig. 6, Table 3, Discussion §4.4
關鍵結果
不需要已知三維結構即可在大型序列庫搜尋 CP 候選,把 CPSARST/CPDB 的結構知識外推到序列尺度。
逐節證據導讀
論文事實、本站判讀與教學模型分開標示。
paper-fact
high-quality label 從哪裡來
CPDB 有4169 nonredundant CP pairs;paper 以 CE-CP structure alignment、length/block/site/gap/coverage/RMSD criteria 與 manual inspection 定義480 high-quality pairs,其餘為 marginal。
因此 SeqCP 學的是接近這個 structural-curation definition 的 sequence signature;AUC 並非『自然界所有 CP』的無偏 prevalence or sensitivity。
原文定位: PDF pp. 3–10, Methods §§2.2 and 2.5, Table 1
paper-fact
0.60 與 0.75 回答不同需求
0.60 在 test data 得最高 MCC;paper 概括 sensitivity/specificity 約86%。這適合 broad discovery。
若 positive follow-up 昂貴,0.75 可減少21% candidates 並移除96% unlikely cases。這不是演算法突然更準,而是 operating point 改變。
原文定位: PDF pp. 10–11 and 16, Fig. 6, Table 3, Discussion §4.4
paper-fact
CP search 如何改善 modeling
UPI00057E200D 的 normal sequence templates coverage 不完整;用 SeqCP 建議 CPM71 後,SWISS-MODEL 可建 full model,與 AlphaFold2 model superposition RMSD 1.20 Å/154 aligned Cα。
A0A388KU10 CPM475 使 template coverage 49→90%、global identity 約22→41%;YbeA CPM74 case 又顯示 sequence detection 可在 CP 改變 domain-swapped conformation 時補 structure-only methods。
原文定位: PDF pp. 11–15, Figs. 7–11, Results §3.6
研究設計與評估
資料與樣本
CPDB 2238 proteins/4169 pairs 作 training;NrPDB100-2007 26,349 proteins 作 threshold test;NrPDB100-2019 85,725 作 large-scale test;20,000 simulated CPMs 作 site test。
比較基準
- CE-CP structural standard、CPSARST site localization、existing motif/sequence methods、normal versus CPM-assisted modeling
指標
- ROC AUC
- SeqCP score 對 high-quality/marginal pairs 的 threshold-free discrimination。
邊界: 依 label definition、class balance 與 dataset relatedness。 - MCC
- 同時考慮 TP/TN/FP/FN 的 threshold quality,用來選0.60。
邊界: deployment prevalence 改變時 optimum 可能移動。 - CP-site accuracy
- simulated descendants 中 exact CP site 被找對的比例。
邊界: in silico mutations 不涵蓋所有 natural evolutionary processes。
論文報告的結果
SeqCP 在多個 tests 達約0.9 AUC、推薦 threshold 0.60,identity>20% 時 exact site accuracy 至少80%,掃描速度約1.75M pairs/min、為 CPSARST 約33倍。
PDF pp. 3–16, Figs. 1–11 and Tables 1–3
teaching-model · not a reported experiment
教學例(不是論文實驗)
threshold 不是信仰:按 follow-up cost 選
教學模型:你能實驗驗證100個 candidates,SeqCP 產生500個≥0.60 hits。
- 若目標是建立 discovery catalog,保留0.60 並分層抽樣驗證,估 calibration。
- 若每個 experiment 昂貴,先用0.75 high-confidence pool,再以 diversity criteria 避免只選同 family。
- 回報每個 threshold 的 candidate count、precision target 與遺漏風險,而非宣稱一個『正確 threshold』。
帶走什麼: score 排序、threshold 決策、實驗 budget 是同一 pipeline 的三個層次。
outside-scope
相對 FAST 的證據邊界
SeqCP 是序列式 CP 搜尋,不能用其分類表現推論 FAST 的 RMSD 或 coverage。
合法來源與取用
17 pages · SHA-256 7723ea237f46a1db5071d7d00e60364b751bc44e0638a0a1e1211ac704c6a76d
2022 年線上發表、正式卷期為 2023;合法開放全文。
限制與防誤讀
- 分數與閾值由 CPDB 案例訓練,資料偏差會影響泛化。
證據定位清單
- PDF pp. 2–6, Methods §§2.1–2.7 and Figs. 1–2
- PDF pp. 7–11, Results §§3.1–3.5, Fig. 6 and Tables 1–3
- PDF pp. 11–16, Results §3.6 and Discussion §§4.1–4.5
理解檢查
為何 duplicate query 能看 CP?
答案: target 的兩個重排片段可跨兩份 query 的接縫連續對齊。
normal alignment 往往只能抓到其中一段或有 boundary shift。
AUC 0.89 是 alignment RMSD 嗎?
答案: 不是,是 CP classification discrimination。
structure geometry 需另以 alignment coverage/RMSD 評估。
0.75 相對0.60 的 trade-off?
答案: precision/specifity 提高但 sensitivity 降低。
適合 follow-up cost 高的情境。
讀完標準: 用一個假想 CP search project,先寫 false positive 與 false negative 的成本,再選 threshold 並說明驗證 sampling plan。
本篇詞彙表
- duplicated query
- 將 query sequence 串接兩次,讓 circularly shifted target 可跨接縫對齊。
- boundary shift
- normal 與 duplicate alignments 的起訖位置差異,作 CP signature。
- MCC
- 在 class imbalance 下仍綜合四格 confusion matrix 的 binary metric。
- operating point
- 選定 threshold 後 sensitivity/specificity/precision 的特定組合。