Evidence-guided paper · 2012 · outside-scope · full-text

Construction and analysis of a plant non-specific lipid transfer protein database (nsLTPDB)

Nai-Jyuan Wang; Chi-Ching Lee; Chao-Sheng Cheng; Wei-Cheng Lo; Ya-Fen Yang; Ming-Nan Chen; Ping-Chiang Lyu. Construction and analysis of a plant non-specific lipid transfer protein database (nsLTPDB). BMC Genomics 13(Suppl 1):S9 (2012).

30-second read

nsLTPDB integrates plant nsLTP sequences, structures, classification, and functional annotations for systematic comparison.

Central question

How can dispersed plant non-specific lipid transfer proteins (nsLTPs), classified under inconsistent standards, be organized into a traceable sequence/structure database and used to derive a more practical family classification and signatures from the eight-cysteine motif and sequence evidence?

Intuition

Known nsLTPs first seed searches of public sequence repositories, after which signal peptides, an eight-cysteine scaffold in the mature sequence, and removal of completely identical sequences act as successive filters. The resulting curated set is not only published on a website but also typed through cysteine spacing, multiple alignment, and phylogenetic trees, with conserved positions translated into scannable Prosite-style patterns.

Why it matters

The value of the database is to bring record provenance, deduplication rules, family classification, structures, and literature into one searchable boundary, enabling later experiments to select residues and representative proteins. However, motif matches, phylogenetic separation, and agreement with prior mutagenesis are not prospective diagnostic accuracy for the new signatures.

Prerequisites

  • Understand the relationship among a precursor, signal peptide, mature protein, and disulfide bonds.
  • Be able to read sequence-motif/regular-expression notation such as C-Xn-C and residue classes.
  • Know that a BLAST identity threshold, 100% deduplication, and manual curation solve different problems.
  • Understand that multiple-sequence alignment, UPGMA, neighbor joining, and class labels are different forms of evidence.

paper-specific guide · plain → technical → input → output → source

Method walkthrough

  1. 01 · Mine, filter, and deduplicate nsLTP candidates

    Search for sequences similar to known proteins, remove candidates lacking the characteristic cysteine scaffold, then eliminate exact duplicates and inspect the remainder manually.

    Technical reading: BLAST 2.2.17 searches SwissProt with known plant nsLTPs; plant homologs with sequence identity >15% become candidates, those lacking the eight-cysteine motif are removed, records at 100% sequence identity are deduplicated, and remaining candidates are manually examined. SignalP 3.0 predicts and removes signal peptides to obtain putative mature sequences.

    Input: RefSeq/GenBank/SwissProt protein records, known nsLTP seed sequences, and SignalP predictions.

    Output: A resource layer of 1,395 putative records and a 595-sequence set deduplicated at 100% identity for downstream analysis.

    Boundary: The >15% identity threshold is a candidate-retrieval condition, while the eight-cysteine motif and manual review provide subsequent filtering; 595/1,395 is not precision supported by gold-standard negatives.

    PDF pp. 2–3, Methods, SignalP, Data mining and Table 1

  2. 02 · Classify by cysteine spacing and phylogeny

    Align mature sequences, examine spacing among the eight conserved cysteines and other conserved positions, then check whether the proposed types also separate in trees.

    Technical reading: ClustalW 2.0.12 builds an MSA that is manually refined; PHYLIP 3.67 reconstructs phylogenetic trees with UPGMA and neighbor joining. The authors modify Boutrot's nine-type framework into five types, grouping by eight-cysteine-motif spacing and sequence similarity and deriving new Prosite-style patterns for Types I and II.

    Input: The 595 putative mature nsLTP sequences and their eight-cysteine spacing/alignment columns.

    Output: Five nsLTP types, phylogenetic visualizations, and Type I/II signatures.

    Boundary: The same sequences generate and describe the classification; tree separation is internal consistency, not classification accuracy on an independent test set. New patterns were not derived for Types III–V because their case counts were small.

    PDF pp. 3–5, Sequence alignment, Table 2, phylogenetic analysis and Figure 1

  3. 03 · Publish a queryable resource and triangulate with prior experiments

    Publish sequences, structures, species, references, and classifications on a website, then check whether conserved signature positions correspond to residues previously known to affect structure or lipid binding.

    Technical reading: The historical nsLTPDB interface contains Homepage, Species Browsing, Structure Browsing, Related References, and tools, and lists PDB structures. Conserved Type I/II positions are then compared with the group's earlier mungbean alanine-scanning and rice nsLTP2 docking/stability experiments.

    Input: Curated annotations, PDB structures, literature records, new patterns, and prior mutagenesis/modeling results.

    Output: A browsable database and a set of candidate conserved residues for subsequent wet-lab validation.

    Boundary: Figures 2–3 explicitly refer to prior data; the pattern–function connection in this paper is retrospective triangulation and should not be described as newly performed prospective mutagenesis.

    PDF pp. 5–7, Figures 2–5 and The nsLTPDB

Key result

The contribution is a reusable data foundation for plant nsLTP family integration and classification.

Evidence-guided deep reading

Paper facts, project readings, and teaching models are labelled separately.

paper-fact

First distinguish the 1,395- and 595-record layers

The Methods report 1,395 putative nsLTP sequences after signal-peptide/eight-cysteine screening, with the web interface collecting these putative records; after removing sequences identical at 100%, 595 sequences from 121 species remain for protein analysis and evolutionary study.

On PDF p.6, the resource description says the database then held 1,395 putative sequences and 32 PDB structures, while the Conclusion on the same page says the database contains 595 nsLTPs. The most defensible reading treats 1,395 as the putative resource layer and 595 as the nonredundant curated analysis set, while preserving the paper's wording conflict rather than forcing the counts into one.

Source locator: PDF pp. 1–3, Abstract, Methods and Table 1; PDF p. 6, The nsLTPDB and Conclusions

paper-fact

The eight-cysteine motif is a family scaffold, not sufficient by itself for typing

The authors state that the shared C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C motif identifies the family scaffold but cannot classify it alone, so they combine cysteine-flanking lengths, sequence similarity, and phylogenetic separation into a five-type classification; the 595 sequences separate into five groups in the constructed trees.

Prosite 20.8 contains 1,331 patterns, but the existing PLANT_LTP pattern PS00597 recognizes only 86 cases in the collection and misses most nsLTP sequences. The authors therefore derive new Type I/II patterns from alignment conservation; this improves a hypothesis about family-specific coverage but still lacks sensitivity and specificity on independent positives and negatives.

Source locator: PDF pp. 3–5, Table 2, phylogenetic analysis and Strategies for defining new Prosite-styled patterns

project-reading

Conserved positions prioritize experiments; they do not settle function

In the Type I pattern, several cavity/lipid-transfer-related positions previously identified through mungbean alanine scanning have major residue classes conserved at mostly at least 86%; in rice Type II, residue classes corresponding to positions such as Leu8, Phe36, Phe39, Tyr48, and Val49 are at least 92%, while position 45 is 75%.

This agreement strengthens the rationale for using newly conserved positions as mutagenesis targets, but conservation percentage is neither an effect size nor a probability of lipid-transfer function. The next step should separately validate pattern classification, fold stability, binding, and transfer activity on held-out sequences and new experiments.

Source locator: PDF p. 5, Figure 1, The mungbean nsLTP1 and The rice nsLTP2; PDF p. 6, Figures 2–3

Study design and evaluation

Data and samples

Candidate sources include NCBI RefSeq/GenBank, SwissProt, and PDB; BLAST >15% identity, the eight-cysteine motif, SignalP 3.0, 100%-identity deduplication, and manual examination produce a 595-sequence/121-species analysis set. Structures, literature, and biological data are additionally integrated into the historical web resource.

Baselines

  • Boutrot and colleagues' nine-type classification based on a substantially smaller collection as the historical contrast for the paper's five-type framework.
  • The existing PLANT_LTP signature PS00597 in Prosite 20.8 as the coverage motivation for the new Type I/II patterns.
  • The group's prior mungbean alanine-scanning and rice nsLTP2 structural/docking results as retrospective functional references for conserved positions.

Metrics

Curated nonredundant set size
The 595 sequences from 121 species used for analysis after the paper's retrieval, motif filtering, 100%-identity deduplication, and manual-review workflow.
Boundary: This is resource scale, not search precision, recall, or complete biological coverage; the paper's database wording around 1,395 versus 595 is also not fully consistent.
Signal-peptide prevalence
SignalP 3.0 predicts that 98% of precursors in the 595-sequence set have a signal peptide 7–49 amino acids long.
Boundary: This is software-predicted prevalence, not experimentally validated sensitivity for secretion.
Phylogenetic type separation
The 595 sequences grouped by the paper's five-type labels separate in UPGMA/neighbor-joining visualizations.
Boundary: The labels and trees come from the same sequence collection, with no held-out test set or confusion matrix.
Conserved-residue frequency
The occurrence percentage of the major amino acid or physicochemical class in an alignment column, used to prioritize potentially important residues.
Boundary: It is not mutation effect size, binding affinity, or functional probability and is affected by sequence sampling and redundancy policy.

Reported result

The study builds a nonredundant analysis set of 595 sequences from 121 species; 98% of precursors are predicted to contain signal peptides 7–49 amino acids long. It proposes a five-type classification from eight-cysteine spacing and sequence evidence and derives new signatures for Types I and II. The resource description separately reports 1,395 putative sequences and 32 PDB structures, which should be read as a different layer from the 595-record curated set.

PDF pp. 1–5, Abstract, Methods, Tables 1–2 and Figure 1; PDF p. 6, The nsLTPDB and Conclusions

teaching-model · not a reported experiment

Teaching example (not a reported experiment)

From one candidate precursor to an auditable nsLTP record

Project teaching model: BLAST retrieves a 112-aa plant protein with 24% identity to a seed; SignalP predicts the first 23 aa as a signal peptide, and the remaining mature sequence contains eight cysteines.

  1. Preserve the original accession, database version, and BLAST evidence; 24% identity only admits it to the candidate pool.
  2. After removing the predicted signal peptide, mark all eight cysteines, confirm the C-Xn-C-Xn-CC-Xn-CXC-Xn-C-Xn-C scaffold, and record every n.
  3. Deduplicate under a 100%-identity policy; if it is not an exact duplicate, add it to the MSA and assign a provisional type from spacing and alignment evidence.
  4. Store SignalP, type, signature match, and PDB/literature links as separate evidence fields; if no experimental data exist, mark function as unvalidated.

Takeaway: A reusable database record preserves each inference layer and its source; candidate retrieval, motif passage, type assignment, and experimental function must not collapse into a single “confirmed” label.

outside-scope

Evidence boundary versus FAST

Database and protein-family analysis do not constitute a FAST alignment comparison.

Lawful source and access

9 pages · SHA-256 e96fbfd616e08023d6591c060e66f0c422c3c0137b0809b09f51ed9752c13b66

Published in a journal supplement from APBC 2012; lawful open full text.

Europe PMC open-access PDF

Limits and misreadings

  • Coverage reflects plant genomes and annotations available at the time.

Source locator map

  1. PDF pp. 1–3, Abstract, Methods and Table 1
  2. PDF p. 3, SignalP, Data mining and sequence-alignment methods
  3. PDF pp. 4–5, Table 2, phylogenetic analysis and Figure 1
  4. PDF pp. 5–6, prior mutagenesis/modeling comparisons and Figures 2–3
  5. PDF pp. 6–7, The nsLTPDB, Conclusions and Figure 5

Check understanding

  1. Is BLAST identity >15% sufficient to confirm an nsLTP?

    Answer: No; it only defines a candidate, followed by eight-cysteine filtering, deduplication, and manual examination.

    A low threshold retrieves candidates and cannot serve as the final family label.

  2. Why can we not simply say nsLTPDB has only 595 records?

    Answer: The Methods/resource description also reports 1,395 putative sequences, while 595 is the 100%-deduplicated analysis set; the Conclusion additionally mixes the two layers in its wording.

    Correct reporting pairs each count with its curation state and purpose.

  3. Does conservation ≥92% mean a mutation has a 92% probability of disrupting binding?

    Answer: No; it is an alignment occurrence frequency, not a mutation-outcome probability.

    Functional effects still require separate measurements of fold, stability, binding, and transfer activity.

Completion task: Design a reproducible nsLTPDB update workflow: pin source-database versions; preserve raw candidates, exclusion reasons at each step, 100%-identity clusters, mature-sequence coordinates, and provisional types; then evaluate the new Type I/II signatures on independent positives and negatives for sensitivity, specificity, and family coverage.

Paper-specific glossary

Eight-cysteine motif
The family scaffold formed by eight highly conserved cysteines and their spacing in mature plant nsLTP sequences.
Putative record
A candidate record supported by computational evidence but not equivalent to complete experimental confirmation.
Prosite-style pattern
A rule representation of a protein-family signature using fixed residues, allowed residue classes, and variable gaps.
100%-identity deduplication
Merging exactly identical sequences to avoid repeated counting; it does not remove all close-homology bias.