Evidence-guided paper · 2012 · outside-scope · full-text
植物非特異性脂質轉運蛋白資料庫 nsLTPDB
Nai-Jyuan Wang; Chi-Ching Lee; Chao-Sheng Cheng; Wei-Cheng Lo; Ya-Fen Yang; Ming-Nan Chen; Ping-Chiang Lyu. Construction and analysis of a plant non-specific lipid transfer protein database (nsLTPDB). BMC Genomics 13(Suppl 1):S9 (2012).
30 秒理解
整合植物 nsLTP 序列、結構、分類與功能註解,讓分散資料可以系統比較。
核心問題
如何從分散且分類標準不一的植物 non-specific lipid transfer proteins(nsLTPs),建立可追溯的序列/結構資料庫,並以 8-cysteine motif 與序列證據提出較可用的家族分類與 signatures?
直覺
先用已知 nsLTPs 當種子在公開序列庫找候選,再把「有訊號肽、成熟序列帶 8-Cys 骨架、去除完全相同序列」當成逐層濾網。得到的 curated set 不只放進網頁,也用 cysteine 間距、multiple alignment 與 phylogenetic trees 分型,最後把保守位置轉成可掃描的 Prosite-style patterns。
為什麼重要
資料庫工作的價值在於把 record provenance、去冗餘規則、family classification、結構與文獻放在同一可查詢邊界,讓後續實驗能挑選 residues 與代表性 proteins。不過 motif 命中、phylogenetic separation 與先前 mutagenesis 的一致性都不是新 signature 的 prospective diagnostic accuracy。
閱讀前置
- 理解 precursor、signal peptide、mature protein 與 disulfide bond 的關係。
- 能閱讀 sequence motif/regular-expression notation,例如 C-Xn-C 與 residue classes。
- 知道 BLAST identity threshold、100% deduplication 與 manual curation 各自解決不同問題。
- 理解 multiple sequence alignment、UPGMA、neighbor joining 與分類標籤不是同一種證據。
paper-specific guide · plain → technical → input → output → source
逐步方法導讀
-
01 · 搜尋、篩選與去冗餘 nsLTP 候選
從已知蛋白找相似序列,先排除不具典型 cysteine 骨架的候選,再移除完全重複項並人工檢視。
論文語言: BLAST 2.2.17 以已知植物 nsLTPs 搜 SwissProt;植物 homologs 在 sequence identity >15% 時列為 candidates,沒有 8-Cys motif 者濾除,100% sequence-identity records 去冗餘,剩餘 candidates 再 manual examination。SignalP 3.0 用於預測與移除 signal peptide,以取得 putative mature sequences。
輸入: RefSeq/GenBank/SwissProt protein records、已知 nsLTP seed sequences 與 SignalP predictions。
輸出: 1,395 個 putative resource records,以及供後續分析的 595-sequence、100% identity 去冗餘集合。
邊界: >15% identity 是候選召回條件,8-Cys motif 與人工檢視才是後續篩選;595/1,395 不是有 gold-standard negatives 支持的 precision。
PDF pp. 2–3, Methods, SignalP, Data mining and Table 1
-
02 · 以 cysteine 間距與 phylogeny 分型
對成熟序列排齊,觀察八個保守 cysteines 之間的間距與其他保守位置,再檢查分型是否也會在樹上分開。
論文語言: ClustalW 2.0.12 建 MSA 並人工修整;PHYLIP 3.67 以 UPGMA 與 neighbor joining 重建 phylogenetic trees。作者將 Boutrot 的九型架構調整為五型,依 8-Cys motif 間距與 sequence similarity 分組,並為 Type I/II 形成新 Prosite-style patterns。
輸入: 595 條 putative mature nsLTP sequences 與其 8-Cys spacing/alignment columns。
輸出: 五個 nsLTP types、phylogenetic visualization,以及 Type I/II signatures。
邊界: 同一批 sequences 同時用於產生與描述分類;tree separation 是內部一致性,不是獨立 test set 的 classification accuracy。Type III–V 因案例較少而未建立新 patterns。
PDF pp. 3–5, Sequence alignment, Table 2, phylogenetic analysis and Figure 1
-
03 · 發布可查詢資源並連結既有實驗
把 sequence、structure、species、reference 與分類放進網頁,再檢查 signature 的保守位置是否對應先前已知會影響結構或 lipid binding 的 residues。
論文語言: historical nsLTPDB 介面分為 Homepage、Species Browsing、Structure Browsing、Related References 與 tools,並列出 PDB structures。Type I/II conserved positions再與作者群先前的 mungbean alanine scanning、rice nsLTP2 docking/stability experiments 對照。
輸入: curated annotations、PDB structures、literature records、新 patterns 與既有 mutagenesis/modeling 結果。
輸出: 可瀏覽資料庫與一組供後續 wet-lab 驗證的候選 conserved residues。
邊界: Figures 2–3 明確引用先前資料;本篇的 pattern–function 連結是回溯式 triangulation,不應敘述成新做的 prospective mutagenesis。
PDF pp. 5–7, Figures 2–5 and The nsLTPDB
關鍵結果
把植物 nsLTP 家族整理成可重用資料基礎,重點是資料整合與分類。
逐節證據導讀
論文事實、本站判讀與教學模型分開標示。
paper-fact
先釐清 1,395 與 595 是哪一層資料
Methods 表示 signal-peptide/8-Cys screening 後得到 1,395 putative nsLTP sequences,web interface 收錄這些 putative records;再去除 100% sequence-identical redundancies後,留下 595 條序列供 protein analysis 與 evolutionary study,來自 121 species。
PDF p.6 的 resource description 說資料庫當時有 1,395 putative sequences 與 32 PDB structures,但同頁 Conclusion 又說 database contains 595 nsLTPs。最穩健的讀法是把 1,395 視為 putative resource layer、595 視為 nonredundant curated analysis set,同時保留原文口徑衝突,不把兩數強行合併。
原文定位: PDF pp. 1–3, Abstract, Methods and Table 1; PDF p. 6, The nsLTPDB and Conclusions
paper-fact
8-Cys 是 family 骨架,不足以單獨分型
作者指出共同的 C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C motif 可辨識家族骨架,卻不能單獨完成分類,因此將 cysteine 間 flanking lengths、sequence similarity 與 phylogenetic separation 合併,提出五型分類;595 sequences 在所建樹中分成五群。
Prosite 20.8 有 1,331 patterns,但既有 PLANT_LTP pattern PS00597 只在收集資料中辨識到 86 cases,且漏掉多數 nsLTP sequences。作者因而從 alignment conservation 形成 Type I/II 新 patterns;這改善的是 family-specific coverage hypothesis,尚缺獨立 positives/negatives 的 sensitivity 與 specificity。
原文定位: PDF pp. 3–5, Table 2, phylogenetic analysis and Strategies for defining new Prosite-styled patterns
project-reading
保守位置是實驗優先序,不是功能定論
Type I pattern 中多個先前由 mungbean alanine scanning 指出的 cavity/lipid-transfer-related positions,其主要 residue classes 多為至少 86% 保守;rice Type II 的 Leu8、Phe36、Phe39、Tyr48、Val49 等位置對應 residue classes 至少 92%,position 45 為 75%。
這種一致性提高新保守位置作為 mutagenesis targets 的合理性,卻不能把 conservation percentage 當作 effect size 或 lipid-transfer prediction probability。下一步應在預先保留的 sequences 與新實驗中,分開驗證 pattern classification、fold stability、binding 與 transfer activity。
原文定位: PDF p. 5, Figure 1, The mungbean nsLTP1 and The rice nsLTP2; PDF p. 6, Figures 2–3
研究設計與評估
資料與樣本
候選來源涵蓋 NCBI RefSeq/GenBank、SwissProt 與 PDB;BLAST >15% identity、8-Cys motif、SignalP 3.0、100% identity deduplication 與 manual examination 形成 595-sequence/121-species analysis set。結構、文獻與 biological data 另整合至 historical web resource。
比較基準
- Boutrot 等人以 267/251-scale records 建立的九型分類,作為本文調整為五型架構的歷史對照。
- Prosite 20.8 的既有 PLANT_LTP signature PS00597,作為新 Type I/II patterns 的 coverage motivation。
- 作者群先前的 mungbean alanine scanning 與 rice nsLTP2 structural/docking results,作為保守位置的回溯式功能參照。
指標
- Curated nonredundant set size
- 在本文 retrieval、motif filtering、100% identity deduplication 與人工檢視流程後,用於分析的 595 sequences/121 species。
邊界: 是資料資源規模,不是 search precision、recall 或生物界完整覆蓋率;paper 對 1,395/595 的 database wording 也不完全一致。 - Signal-peptide prevalence
- 595-set 中由 SignalP 3.0 預測具有 7–49 amino-acid signal peptide 的 precursor 比例為 98%。
邊界: 這是 software prediction prevalence,不是以實驗定位驗證的 secretion sensitivity。 - Phylogenetic type separation
- 依本文五型 labels 分組的 595 sequences 在 UPGMA/neighbor-joining visualizations 中可分開。
邊界: labels 與 tree 都源自同一 sequence collection,沒有 held-out test set 或 confusion matrix。 - Conserved-residue frequency
- alignment column 中主要 amino acid 或 physicochemical class 的 occurrence percentage,用於挑選可能重要 residues。
邊界: 不是 mutation effect size、binding affinity 或功能機率;會受 sequence sampling 與 redundancy policy 影響。
論文報告的結果
本文建立 595-sequence、121-species 的 nonredundant analysis set,98% precursors 被預測含 7–49-aa signal peptide;依 8-Cys spacing 與 sequence evidence 提出五型分類,並為 Type I/II 建立新 signatures。resource description 另報 1,395 putative sequences 與 32 PDB structures,應與 595 curated set 分層閱讀。
PDF pp. 1–5, Abstract, Methods, Tables 1–2 and Figure 1; PDF p. 6, The nsLTPDB and Conclusions
teaching-model · not a reported experiment
教學例(不是論文實驗)
從一條候選 precursor 到可審核的 nsLTP record
本站教學模型:一條 112-aa 植物 protein 被 BLAST 找到,對 seed identity 為 24%,SignalP 預測前 23 aa 是 signal peptide;剩餘 mature sequence 含八個 cysteines。
- 保留原始 accession、資料庫版本與 BLAST evidence;24% identity 只讓它進 candidate pool。
- 移除預測 signal peptide 後,標出八個 cysteines,確認順序符合 C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C 骨架並記錄每個 n。
- 以 100% identity policy 查重;若不是 exact duplicate,再放入 MSA,依 spacing 與 alignment evidence 產生 provisional type。
- 另列 SignalP、type、signature match 與 PDB/literature links 的 evidence fields;若沒有實驗資料,function 欄標為未驗證。
帶走什麼: 一筆可重用 database record 應保留每一層推論與來源,不能把候選召回、motif 通過、type assignment 與實驗功能合成單一「已確認」。
outside-scope
相對 FAST 的證據邊界
資料庫與蛋白家族分析不構成 FAST alignment 比較。
合法來源與取用
9 pages · SHA-256 e96fbfd616e08023d6591c060e66f0c422c3c0137b0809b09f51ed9752c13b66
刊於期刊 supplement、源自 APBC 2012;合法開放全文。
限制與防誤讀
- 資料庫內容反映當時可取得的植物基因體與註解。
證據定位清單
- PDF pp. 1–3, Abstract, Methods and Table 1
- PDF p. 3, SignalP, Data mining and sequence-alignment methods
- PDF pp. 4–5, Table 2, phylogenetic analysis and Figure 1
- PDF pp. 5–6, prior mutagenesis/modeling comparisons and Figures 2–3
- PDF pp. 6–7, The nsLTPDB, Conclusions and Figure 5
理解檢查
BLAST identity >15% 是否足以確認 nsLTP?
答案: 不足;它只定義 candidate,後面仍有 8-Cys filtering、去冗餘與 manual examination。
低 threshold 用於召回候選,不能當作最終 family label。
為何不能直接說 nsLTPDB 只有 595 筆?
答案: Methods/resource description 另報 1,395 putative sequences,而 595 是 100% 去冗餘後的分析集合;結論用詞又混合兩層。
正確報告應同時給 count、curation state 與用途。
保守頻率 ≥92% 是否代表該 mutation 有 92% 機率破壞 binding?
答案: 不代表;那是 alignment occurrence frequency,不是 mutation-outcome probability。
功能影響仍需 fold、stability、binding 與 transfer assays 分開測量。
讀完標準: 設計 nsLTPDB 的可重現更新流程:鎖定 source database versions,保存 raw candidates、每一步 exclusion reason、100% identity clusters、mature-sequence coordinates 與 provisional types;再以獨立 positives/negatives 評估新 Type I/II signatures 的 sensitivity、specificity 與 family coverage。
本篇詞彙表
- 8-Cys motif
- plant nsLTP mature sequence 中八個高度保守 cysteines 及其間隔所形成的家族骨架。
- Putative record
- 由計算 evidence 支持、但不等於已由實驗完整確認的候選資料。
- Prosite-style pattern
- 以固定 residues、允許的 residue classes 與可變間隔描述 protein-family signature 的規則表示。
- 100% identity deduplication
- 將完全相同的 sequences 合併,避免同一序列重複計數;不會移除所有近緣偏差。