Evidence-guided paper · 2021 · benchmark-methodology · full-text

資料集同源性如何扭曲二級結構預測評估

Teng-Ruei Chen; Chia-Hua Lo; Sheng-Hung Juan; Wei-Cheng Lo. The influence of dataset homology and a rigorous evaluation strategy on protein secondary structure prediction. PLOS ONE 16:e0254555 (2021).

30 秒理解

系統拆解 training、test、independent test 與 reference database 內外的同源冗餘,辨識哪些會造成成績高估或低估。

核心問題

training、testing、independent-test 與 PSSM reference datasets 之內/之間的 homology,哪些會讓 SSP accuracy 被高估或低估?

直覺

『train 與 test 低同源』不是充分條件;要把每個 dataset 內的 redundancy、query-reference leakage、independent set 與所有 development data 的關係分開控制。paper 的核心是把一團 homology 拆成可操縱的邊。

為什麼重要

模型差異可能小於資料切分偏差。這篇提供一個更嚴格的 evaluation blueprint:所有 sets 內外約 30% cutoff、真正時間外 independent tests、多次 random sampling,以及 length-weighted/micro averages。

閱讀前置

  • 理解 training/test/independent test 的角色
  • 理解 PSSM reference set 與 query set 不同
  • 知道 overfitting、data leakage 與 sequence identity cutoff

paper-specific guide · plain → technical → input → output → source

逐步方法導讀

  1. 01 · 逐邊操縱 homology

    固定其他關係,一次只改 train–test inter-set、各 query set inner-set、reference isolation 或 query–reference homology。

    論文語言: 小型 experiments 用 NrPdbx-2015、250 train/250 test、10,000 reference,通常 random sampling 重複 20 次;TS115/CASP12 作 independent tests。

    輸入: 具有明確 inner/inter identity cutoffs 的 dataset layouts

    輸出: 每條 homology edge 對 apparent/practical accuracy 的影響

    邊界: identity cutoff 是 heuristic proxy,不等於完整 evolutionary independence。

    PDF pp. 3–8 and 17–19, Results, Figs. 1–5, Materials and methods

  2. 02 · 用真正新蛋白挑戰模型

    把 training/testing 的 accuracy 視為 apparent,只有時間外、低同源 independent tests 的 accuracy 接近 practical performance。

    論文語言: TS115/CASP12/13 為 2016 後 structures;TS416 則用 2019–2020 PDB 且對 2018 前資料 <25% identity。

    輸入: 未被 model development 看過的 time-split query sets

    輸出: 較可信的 practical accuracy 與 overfitting gap

    邊界: 資料庫隨時間新增 homologs,曾經 independent 的 set 會逐漸失去隔離性。

    PDF pp. 3–6 and 18–19, Results Figs. 1–3, Independent test datasets

  3. 03 · 建立完整低冗餘 layout

    讓所有 query/reference sets 內外都高度 nonredundant,尤其 independent set 對所有其他資料要低同源。

    論文語言: 作者建議約 30% identity cutoff、reference 約 5M 足以飽和 accuracy、multiple random repeats,並用 length-weighted SOV 與 residue micro-Q。

    輸入: 原始 PDB/UniRef 與 homology-reduction tools

    輸出: 可稽核且較少 leakage 的 development/evaluation protocol

    邊界: paper 針對 PSSM-based SSP;轉移到 structure alignment 或 foundation models 需重新定義 leakage。

    PDF pp. 15–20, Fig. 9, 'Details of the proposed strategy' and accuracy methods

關鍵結果

多數資料集內外冗餘會高估成績;參考資料庫內冗餘反而可能低估,整體高估效應較強。

逐節證據導讀

論文事實、本站判讀與教學模型分開標示。

paper-fact

哪些傳統假設沒有成立

在 inner-set homology 固定時,降低 train–test inter-set identity 幾乎不改 training、testing 或 independent accuracy;相反地,降低各 query set 內部 redundancy 讓 apparent accuracy 靠近 independent accuracy,減少 overfitting。

把 PSSM reference 分成 train/test/independent 各自一份也沒有降低 overfitting;真正重要的是 query-reference homology 與 reference set 本身的 redundancy。

原文定位: PDF pp. 3–8, Figs. 1–5

paper-fact

高估與低估可同時發生

多數 within/between dataset redundancy 會讓 accuracy 高估;尤其 independent query 與 reference set 同源過高時,『independent』名義不足以阻止 leakage。

但 reference set 內部 redundancy 可能降低多樣性、使 PSSM information entropy 與 accuracy 偏低。paper 認為整體而言 inflation effects 較強。

原文定位: PDF pp. 8–15, Figs. 5–8, Discussion

paper-fact

平均方式也是 evaluation design

arithmetic mean 讓 50-residue 與 500-residue protein 權重相同,可能 overweight small proteins;作者因此使用 residue-based micro Q 與 length-weighted SOV。

Q 計算逐 residue 命中,SOV 評估整段 secondary-structure segment;兩者趨勢可相同但回答層次不同,不能只挑最好看的數字。

原文定位: PDF p. 20, 'Computation of secondary structure prediction accuracy' and averaging formulas

研究設計與評估

資料與樣本

PDB/UniRef 2015 的多個 identity-reduced sets;TS115、CASP12/13、TS416;small layouts 與最多 5M-reference large-scale tests。

比較基準

  • 不同 inner/inter homology layouts、shared versus isolated references、in-house ANN/DT/SVM 與 11 種 SSP programs

指標

Q3
逐 residue 三態正確率;paper 主要以 residue micro-average 報告。
邊界: 不描述 segment continuity。
SOV3
secondary-structure segments 的 overlap quality,採 length weighting。
邊界: 與 Q3 數值尺度不同。
independent-test gap
training/testing 與 time-held-out accuracy 差距,作 overfitting 診斷。
邊界: gap 小仍不保證 deployment distribution matching。

論文報告的結果

inter-query-set homology 影響小;query sets 內 redundancy 與 query-reference homology 易高估;reference 內 redundancy 可低估;推薦所有 sets 內外約 30% nonredundancy 與多次抽樣。

PDF pp. 3–20, Figs. 1–9 and Materials and methods

teaching-model · not a reported experiment

教學例(不是論文實驗)

稽核一個聲稱 1% 提升的 SSP

教學模型:新方法 Q3 85%,baseline 84%,但資料切分只說 train/test <30%。

  1. 畫出 train、test、independent、reference 四節點,逐一詢問每個節點內與每條邊的 identity cutoff。
  2. 確認 independent set 是否時間外,且對 reference 與所有 training data 都去同源。
  3. 要求 multiple seeds 的 micro-Q、weighted SOV 與 confidence interval;若偏差可達 1%,不能宣稱模型本身勝出。

帶走什麼: 排行榜差距只有在資料關係圖透明時才有意義。

benchmark-methodology

相對 FAST 的證據邊界

雖不比較 FAST,卻是讀任何演算法 leaderboard 的重要警告:同源洩漏與評估設計可比模型差異更重要。

合法來源與取用

31 pages · SHA-256 85e31de3ef6ad2ec5460b6590f5144003dee37a4e2f7c77de7fbfaa578d7660d

合法開放全文;NYCU Pure 將其分類為 review article,但正式出版記錄為研究論文。

Europe PMC open-access PDF

限制與防誤讀

  • 結論針對 SSP 特徵流程,移植到結構 alignment 仍需重新驗證。

證據定位清單

  1. PDF pp. 3–8, Results and Figs. 1–5
  2. PDF pp. 11–16, Discussion and Figs. 7–9
  3. PDF pp. 16–20, Materials and methods

理解檢查

  1. 降低 train–test inter-set homology 就足夠嗎?

    答案: 不夠。

    還需控制各 set 內 redundancy、query-reference 及 independent-to-all relations。

  2. 同一 reference 給 train/test 一定 leakage 嗎?

    答案: paper 的 controlled tests 顯示 reference isolation 本身不改 accuracy。

    關鍵是 query 與 reference 的 homology,而非檔案是否物理分開。

  3. 為何用 micro-Q?

    答案: 讓每個 residue 等權,避免小 protein 被 arithmetic mean 過度加權。

    averaging choice 會改變看見的 performance。

讀完標準: 為任一 SSP benchmark 畫 dataset homology graph,標出所有 inner/inter cutoffs、release dates、sampling repeats 與 averaging rule。

本篇詞彙表

inner-dataset homology
同一 dataset 內 sequences 的相似/冗餘程度。
inter-dataset homology
兩個 datasets 之間 sequences 的相似程度。
practical accuracy
paper 用來指嚴格 independent test 上較接近新蛋白使用情境的 accuracy。
micro-average
先加總所有 residue 的命中再除以總 residue 數。