Evidence-guided paper · 2021 · benchmark-methodology · full-text
資料集同源性如何扭曲二級結構預測評估
Teng-Ruei Chen; Chia-Hua Lo; Sheng-Hung Juan; Wei-Cheng Lo. The influence of dataset homology and a rigorous evaluation strategy on protein secondary structure prediction. PLOS ONE 16:e0254555 (2021).
30 秒理解
系統拆解 training、test、independent test 與 reference database 內外的同源冗餘,辨識哪些會造成成績高估或低估。
核心問題
training、testing、independent-test 與 PSSM reference datasets 之內/之間的 homology,哪些會讓 SSP accuracy 被高估或低估?
直覺
『train 與 test 低同源』不是充分條件;要把每個 dataset 內的 redundancy、query-reference leakage、independent set 與所有 development data 的關係分開控制。paper 的核心是把一團 homology 拆成可操縱的邊。
為什麼重要
模型差異可能小於資料切分偏差。這篇提供一個更嚴格的 evaluation blueprint:所有 sets 內外約 30% cutoff、真正時間外 independent tests、多次 random sampling,以及 length-weighted/micro averages。
閱讀前置
- 理解 training/test/independent test 的角色
- 理解 PSSM reference set 與 query set 不同
- 知道 overfitting、data leakage 與 sequence identity cutoff
paper-specific guide · plain → technical → input → output → source
逐步方法導讀
-
01 · 逐邊操縱 homology
固定其他關係,一次只改 train–test inter-set、各 query set inner-set、reference isolation 或 query–reference homology。
論文語言: 小型 experiments 用 NrPdbx-2015、250 train/250 test、10,000 reference,通常 random sampling 重複 20 次;TS115/CASP12 作 independent tests。
輸入: 具有明確 inner/inter identity cutoffs 的 dataset layouts
輸出: 每條 homology edge 對 apparent/practical accuracy 的影響
邊界: identity cutoff 是 heuristic proxy,不等於完整 evolutionary independence。
PDF pp. 3–8 and 17–19, Results, Figs. 1–5, Materials and methods
-
02 · 用真正新蛋白挑戰模型
把 training/testing 的 accuracy 視為 apparent,只有時間外、低同源 independent tests 的 accuracy 接近 practical performance。
論文語言: TS115/CASP12/13 為 2016 後 structures;TS416 則用 2019–2020 PDB 且對 2018 前資料 <25% identity。
輸入: 未被 model development 看過的 time-split query sets
輸出: 較可信的 practical accuracy 與 overfitting gap
邊界: 資料庫隨時間新增 homologs,曾經 independent 的 set 會逐漸失去隔離性。
PDF pp. 3–6 and 18–19, Results Figs. 1–3, Independent test datasets
-
03 · 建立完整低冗餘 layout
讓所有 query/reference sets 內外都高度 nonredundant,尤其 independent set 對所有其他資料要低同源。
論文語言: 作者建議約 30% identity cutoff、reference 約 5M 足以飽和 accuracy、multiple random repeats,並用 length-weighted SOV 與 residue micro-Q。
輸入: 原始 PDB/UniRef 與 homology-reduction tools
輸出: 可稽核且較少 leakage 的 development/evaluation protocol
邊界: paper 針對 PSSM-based SSP;轉移到 structure alignment 或 foundation models 需重新定義 leakage。
PDF pp. 15–20, Fig. 9, 'Details of the proposed strategy' and accuracy methods
關鍵結果
多數資料集內外冗餘會高估成績;參考資料庫內冗餘反而可能低估,整體高估效應較強。
逐節證據導讀
論文事實、本站判讀與教學模型分開標示。
paper-fact
哪些傳統假設沒有成立
在 inner-set homology 固定時,降低 train–test inter-set identity 幾乎不改 training、testing 或 independent accuracy;相反地,降低各 query set 內部 redundancy 讓 apparent accuracy 靠近 independent accuracy,減少 overfitting。
把 PSSM reference 分成 train/test/independent 各自一份也沒有降低 overfitting;真正重要的是 query-reference homology 與 reference set 本身的 redundancy。
原文定位: PDF pp. 3–8, Figs. 1–5
paper-fact
高估與低估可同時發生
多數 within/between dataset redundancy 會讓 accuracy 高估;尤其 independent query 與 reference set 同源過高時,『independent』名義不足以阻止 leakage。
但 reference set 內部 redundancy 可能降低多樣性、使 PSSM information entropy 與 accuracy 偏低。paper 認為整體而言 inflation effects 較強。
原文定位: PDF pp. 8–15, Figs. 5–8, Discussion
paper-fact
平均方式也是 evaluation design
arithmetic mean 讓 50-residue 與 500-residue protein 權重相同,可能 overweight small proteins;作者因此使用 residue-based micro Q 與 length-weighted SOV。
Q 計算逐 residue 命中,SOV 評估整段 secondary-structure segment;兩者趨勢可相同但回答層次不同,不能只挑最好看的數字。
原文定位: PDF p. 20, 'Computation of secondary structure prediction accuracy' and averaging formulas
研究設計與評估
資料與樣本
PDB/UniRef 2015 的多個 identity-reduced sets;TS115、CASP12/13、TS416;small layouts 與最多 5M-reference large-scale tests。
比較基準
- 不同 inner/inter homology layouts、shared versus isolated references、in-house ANN/DT/SVM 與 11 種 SSP programs
指標
- Q3
- 逐 residue 三態正確率;paper 主要以 residue micro-average 報告。
邊界: 不描述 segment continuity。 - SOV3
- secondary-structure segments 的 overlap quality,採 length weighting。
邊界: 與 Q3 數值尺度不同。 - independent-test gap
- training/testing 與 time-held-out accuracy 差距,作 overfitting 診斷。
邊界: gap 小仍不保證 deployment distribution matching。
論文報告的結果
inter-query-set homology 影響小;query sets 內 redundancy 與 query-reference homology 易高估;reference 內 redundancy 可低估;推薦所有 sets 內外約 30% nonredundancy 與多次抽樣。
PDF pp. 3–20, Figs. 1–9 and Materials and methods
teaching-model · not a reported experiment
教學例(不是論文實驗)
稽核一個聲稱 1% 提升的 SSP
教學模型:新方法 Q3 85%,baseline 84%,但資料切分只說 train/test <30%。
- 畫出 train、test、independent、reference 四節點,逐一詢問每個節點內與每條邊的 identity cutoff。
- 確認 independent set 是否時間外,且對 reference 與所有 training data 都去同源。
- 要求 multiple seeds 的 micro-Q、weighted SOV 與 confidence interval;若偏差可達 1%,不能宣稱模型本身勝出。
帶走什麼: 排行榜差距只有在資料關係圖透明時才有意義。
benchmark-methodology
相對 FAST 的證據邊界
雖不比較 FAST,卻是讀任何演算法 leaderboard 的重要警告:同源洩漏與評估設計可比模型差異更重要。
合法來源與取用
31 pages · SHA-256 85e31de3ef6ad2ec5460b6590f5144003dee37a4e2f7c77de7fbfaa578d7660d
合法開放全文;NYCU Pure 將其分類為 review article,但正式出版記錄為研究論文。
限制與防誤讀
- 結論針對 SSP 特徵流程,移植到結構 alignment 仍需重新驗證。
證據定位清單
- PDF pp. 3–8, Results and Figs. 1–5
- PDF pp. 11–16, Discussion and Figs. 7–9
- PDF pp. 16–20, Materials and methods
理解檢查
降低 train–test inter-set homology 就足夠嗎?
答案: 不夠。
還需控制各 set 內 redundancy、query-reference 及 independent-to-all relations。
同一 reference 給 train/test 一定 leakage 嗎?
答案: paper 的 controlled tests 顯示 reference isolation 本身不改 accuracy。
關鍵是 query 與 reference 的 homology,而非檔案是否物理分開。
為何用 micro-Q?
答案: 讓每個 residue 等權,避免小 protein 被 arithmetic mean 過度加權。
averaging choice 會改變看見的 performance。
讀完標準: 為任一 SSP benchmark 畫 dataset homology graph,標出所有 inner/inter cutoffs、release dates、sampling repeats 與 averaging rule。
本篇詞彙表
- inner-dataset homology
- 同一 dataset 內 sequences 的相似/冗餘程度。
- inter-dataset homology
- 兩個 datasets 之間 sequences 的相似程度。
- practical accuracy
- paper 用來指嚴格 independent test 上較接近新蛋白使用情境的 accuracy。
- micro-average
- 先加總所有 residue 的命中再除以總 residue 數。