Evidence-guided paper · 2012 · outside-scope · full-text

植物非特異性脂質轉運蛋白資料庫 nsLTPDB

Nai-Jyuan Wang; Chi-Ching Lee; Chao-Sheng Cheng; Wei-Cheng Lo; Ya-Fen Yang; Ming-Nan Chen; Ping-Chiang Lyu. Construction and analysis of a plant non-specific lipid transfer protein database (nsLTPDB). BMC Genomics 13(Suppl 1):S9 (2012).

30 秒理解

整合植物 nsLTP 序列、結構、分類與功能註解,讓分散資料可以系統比較。

核心問題

如何從分散且分類標準不一的植物 non-specific lipid transfer proteins(nsLTPs),建立可追溯的序列/結構資料庫,並以 8-cysteine motif 與序列證據提出較可用的家族分類與 signatures?

直覺

先用已知 nsLTPs 當種子在公開序列庫找候選,再把「有訊號肽、成熟序列帶 8-Cys 骨架、去除完全相同序列」當成逐層濾網。得到的 curated set 不只放進網頁,也用 cysteine 間距、multiple alignment 與 phylogenetic trees 分型,最後把保守位置轉成可掃描的 Prosite-style patterns。

為什麼重要

資料庫工作的價值在於把 record provenance、去冗餘規則、family classification、結構與文獻放在同一可查詢邊界,讓後續實驗能挑選 residues 與代表性 proteins。不過 motif 命中、phylogenetic separation 與先前 mutagenesis 的一致性都不是新 signature 的 prospective diagnostic accuracy。

閱讀前置

  • 理解 precursor、signal peptide、mature protein 與 disulfide bond 的關係。
  • 能閱讀 sequence motif/regular-expression notation,例如 C-Xn-C 與 residue classes。
  • 知道 BLAST identity threshold、100% deduplication 與 manual curation 各自解決不同問題。
  • 理解 multiple sequence alignment、UPGMA、neighbor joining 與分類標籤不是同一種證據。

paper-specific guide · plain → technical → input → output → source

逐步方法導讀

  1. 01 · 搜尋、篩選與去冗餘 nsLTP 候選

    從已知蛋白找相似序列,先排除不具典型 cysteine 骨架的候選,再移除完全重複項並人工檢視。

    論文語言: BLAST 2.2.17 以已知植物 nsLTPs 搜 SwissProt;植物 homologs 在 sequence identity >15% 時列為 candidates,沒有 8-Cys motif 者濾除,100% sequence-identity records 去冗餘,剩餘 candidates 再 manual examination。SignalP 3.0 用於預測與移除 signal peptide,以取得 putative mature sequences。

    輸入: RefSeq/GenBank/SwissProt protein records、已知 nsLTP seed sequences 與 SignalP predictions。

    輸出: 1,395 個 putative resource records,以及供後續分析的 595-sequence、100% identity 去冗餘集合。

    邊界: >15% identity 是候選召回條件,8-Cys motif 與人工檢視才是後續篩選;595/1,395 不是有 gold-standard negatives 支持的 precision。

    PDF pp. 2–3, Methods, SignalP, Data mining and Table 1

  2. 02 · 以 cysteine 間距與 phylogeny 分型

    對成熟序列排齊,觀察八個保守 cysteines 之間的間距與其他保守位置,再檢查分型是否也會在樹上分開。

    論文語言: ClustalW 2.0.12 建 MSA 並人工修整;PHYLIP 3.67 以 UPGMA 與 neighbor joining 重建 phylogenetic trees。作者將 Boutrot 的九型架構調整為五型,依 8-Cys motif 間距與 sequence similarity 分組,並為 Type I/II 形成新 Prosite-style patterns。

    輸入: 595 條 putative mature nsLTP sequences 與其 8-Cys spacing/alignment columns。

    輸出: 五個 nsLTP types、phylogenetic visualization,以及 Type I/II signatures。

    邊界: 同一批 sequences 同時用於產生與描述分類;tree separation 是內部一致性,不是獨立 test set 的 classification accuracy。Type III–V 因案例較少而未建立新 patterns。

    PDF pp. 3–5, Sequence alignment, Table 2, phylogenetic analysis and Figure 1

  3. 03 · 發布可查詢資源並連結既有實驗

    把 sequence、structure、species、reference 與分類放進網頁,再檢查 signature 的保守位置是否對應先前已知會影響結構或 lipid binding 的 residues。

    論文語言: historical nsLTPDB 介面分為 Homepage、Species Browsing、Structure Browsing、Related References 與 tools,並列出 PDB structures。Type I/II conserved positions再與作者群先前的 mungbean alanine scanning、rice nsLTP2 docking/stability experiments 對照。

    輸入: curated annotations、PDB structures、literature records、新 patterns 與既有 mutagenesis/modeling 結果。

    輸出: 可瀏覽資料庫與一組供後續 wet-lab 驗證的候選 conserved residues。

    邊界: Figures 2–3 明確引用先前資料;本篇的 pattern–function 連結是回溯式 triangulation,不應敘述成新做的 prospective mutagenesis。

    PDF pp. 5–7, Figures 2–5 and The nsLTPDB

關鍵結果

把植物 nsLTP 家族整理成可重用資料基礎,重點是資料整合與分類。

逐節證據導讀

論文事實、本站判讀與教學模型分開標示。

paper-fact

先釐清 1,395 與 595 是哪一層資料

Methods 表示 signal-peptide/8-Cys screening 後得到 1,395 putative nsLTP sequences,web interface 收錄這些 putative records;再去除 100% sequence-identical redundancies後,留下 595 條序列供 protein analysis 與 evolutionary study,來自 121 species。

PDF p.6 的 resource description 說資料庫當時有 1,395 putative sequences 與 32 PDB structures,但同頁 Conclusion 又說 database contains 595 nsLTPs。最穩健的讀法是把 1,395 視為 putative resource layer、595 視為 nonredundant curated analysis set,同時保留原文口徑衝突,不把兩數強行合併。

原文定位: PDF pp. 1–3, Abstract, Methods and Table 1; PDF p. 6, The nsLTPDB and Conclusions

paper-fact

8-Cys 是 family 骨架,不足以單獨分型

作者指出共同的 C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C motif 可辨識家族骨架,卻不能單獨完成分類,因此將 cysteine 間 flanking lengths、sequence similarity 與 phylogenetic separation 合併,提出五型分類;595 sequences 在所建樹中分成五群。

Prosite 20.8 有 1,331 patterns,但既有 PLANT_LTP pattern PS00597 只在收集資料中辨識到 86 cases,且漏掉多數 nsLTP sequences。作者因而從 alignment conservation 形成 Type I/II 新 patterns;這改善的是 family-specific coverage hypothesis,尚缺獨立 positives/negatives 的 sensitivity 與 specificity。

原文定位: PDF pp. 3–5, Table 2, phylogenetic analysis and Strategies for defining new Prosite-styled patterns

project-reading

保守位置是實驗優先序,不是功能定論

Type I pattern 中多個先前由 mungbean alanine scanning 指出的 cavity/lipid-transfer-related positions,其主要 residue classes 多為至少 86% 保守;rice Type II 的 Leu8、Phe36、Phe39、Tyr48、Val49 等位置對應 residue classes 至少 92%,position 45 為 75%。

這種一致性提高新保守位置作為 mutagenesis targets 的合理性,卻不能把 conservation percentage 當作 effect size 或 lipid-transfer prediction probability。下一步應在預先保留的 sequences 與新實驗中,分開驗證 pattern classification、fold stability、binding 與 transfer activity。

原文定位: PDF p. 5, Figure 1, The mungbean nsLTP1 and The rice nsLTP2; PDF p. 6, Figures 2–3

研究設計與評估

資料與樣本

候選來源涵蓋 NCBI RefSeq/GenBank、SwissProt 與 PDB;BLAST >15% identity、8-Cys motif、SignalP 3.0、100% identity deduplication 與 manual examination 形成 595-sequence/121-species analysis set。結構、文獻與 biological data 另整合至 historical web resource。

比較基準

  • Boutrot 等人以 267/251-scale records 建立的九型分類,作為本文調整為五型架構的歷史對照。
  • Prosite 20.8 的既有 PLANT_LTP signature PS00597,作為新 Type I/II patterns 的 coverage motivation。
  • 作者群先前的 mungbean alanine scanning 與 rice nsLTP2 structural/docking results,作為保守位置的回溯式功能參照。

指標

Curated nonredundant set size
在本文 retrieval、motif filtering、100% identity deduplication 與人工檢視流程後,用於分析的 595 sequences/121 species。
邊界: 是資料資源規模,不是 search precision、recall 或生物界完整覆蓋率;paper 對 1,395/595 的 database wording 也不完全一致。
Signal-peptide prevalence
595-set 中由 SignalP 3.0 預測具有 7–49 amino-acid signal peptide 的 precursor 比例為 98%。
邊界: 這是 software prediction prevalence,不是以實驗定位驗證的 secretion sensitivity。
Phylogenetic type separation
依本文五型 labels 分組的 595 sequences 在 UPGMA/neighbor-joining visualizations 中可分開。
邊界: labels 與 tree 都源自同一 sequence collection,沒有 held-out test set 或 confusion matrix。
Conserved-residue frequency
alignment column 中主要 amino acid 或 physicochemical class 的 occurrence percentage,用於挑選可能重要 residues。
邊界: 不是 mutation effect size、binding affinity 或功能機率;會受 sequence sampling 與 redundancy policy 影響。

論文報告的結果

本文建立 595-sequence、121-species 的 nonredundant analysis set,98% precursors 被預測含 7–49-aa signal peptide;依 8-Cys spacing 與 sequence evidence 提出五型分類,並為 Type I/II 建立新 signatures。resource description 另報 1,395 putative sequences 與 32 PDB structures,應與 595 curated set 分層閱讀。

PDF pp. 1–5, Abstract, Methods, Tables 1–2 and Figure 1; PDF p. 6, The nsLTPDB and Conclusions

teaching-model · not a reported experiment

教學例(不是論文實驗)

從一條候選 precursor 到可審核的 nsLTP record

本站教學模型:一條 112-aa 植物 protein 被 BLAST 找到,對 seed identity 為 24%,SignalP 預測前 23 aa 是 signal peptide;剩餘 mature sequence 含八個 cysteines。

  1. 保留原始 accession、資料庫版本與 BLAST evidence;24% identity 只讓它進 candidate pool。
  2. 移除預測 signal peptide 後,標出八個 cysteines,確認順序符合 C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C 骨架並記錄每個 n。
  3. 以 100% identity policy 查重;若不是 exact duplicate,再放入 MSA,依 spacing 與 alignment evidence 產生 provisional type。
  4. 另列 SignalP、type、signature match 與 PDB/literature links 的 evidence fields;若沒有實驗資料,function 欄標為未驗證。

帶走什麼: 一筆可重用 database record 應保留每一層推論與來源,不能把候選召回、motif 通過、type assignment 與實驗功能合成單一「已確認」。

outside-scope

相對 FAST 的證據邊界

資料庫與蛋白家族分析不構成 FAST alignment 比較。

合法來源與取用

9 pages · SHA-256 e96fbfd616e08023d6591c060e66f0c422c3c0137b0809b09f51ed9752c13b66

刊於期刊 supplement、源自 APBC 2012;合法開放全文。

Europe PMC open-access PDF

限制與防誤讀

  • 資料庫內容反映當時可取得的植物基因體與註解。

證據定位清單

  1. PDF pp. 1–3, Abstract, Methods and Table 1
  2. PDF p. 3, SignalP, Data mining and sequence-alignment methods
  3. PDF pp. 4–5, Table 2, phylogenetic analysis and Figure 1
  4. PDF pp. 5–6, prior mutagenesis/modeling comparisons and Figures 2–3
  5. PDF pp. 6–7, The nsLTPDB, Conclusions and Figure 5

理解檢查

  1. BLAST identity >15% 是否足以確認 nsLTP?

    答案: 不足;它只定義 candidate,後面仍有 8-Cys filtering、去冗餘與 manual examination。

    低 threshold 用於召回候選,不能當作最終 family label。

  2. 為何不能直接說 nsLTPDB 只有 595 筆?

    答案: Methods/resource description 另報 1,395 putative sequences,而 595 是 100% 去冗餘後的分析集合;結論用詞又混合兩層。

    正確報告應同時給 count、curation state 與用途。

  3. 保守頻率 ≥92% 是否代表該 mutation 有 92% 機率破壞 binding?

    答案: 不代表;那是 alignment occurrence frequency,不是 mutation-outcome probability。

    功能影響仍需 fold、stability、binding 與 transfer assays 分開測量。

讀完標準: 設計 nsLTPDB 的可重現更新流程:鎖定 source database versions,保存 raw candidates、每一步 exclusion reason、100% identity clusters、mature-sequence coordinates 與 provisional types;再以獨立 positives/negatives 評估新 Type I/II signatures 的 sensitivity、specificity 與 family coverage。

本篇詞彙表

8-Cys motif
plant nsLTP mature sequence 中八個高度保守 cysteines 及其間隔所形成的家族骨架。
Putative record
由計算 evidence 支持、但不等於已由實驗完整確認的候選資料。
Prosite-style pattern
以固定 residues、允許的 residue classes 與可變間隔描述 protein-family signature 的規則表示。
100% identity deduplication
將完全相同的 sequences 合併,避免同一序列重複計數;不會移除所有近緣偏差。