Evidence-guided paper · 2012 · historical-ecosystem · full-text
Evolutionary information hidden in a single protein structure
Chien-Hua Shih; Chih-Min Chang; Yeong-Shin Lin; Wei-Cheng Lo; Jenn-Kang Hwang. Evolutionary information hidden in a single protein structure. Proteins 80:1647–1657 (2012).
30-second read
Reciprocal weighted contact number and related packing-density descriptors estimate residue-conservation profiles from a single Cα structure.
Central question
Given only the Cα coordinates of one protein structure, and not its multiple-sequence alignment, can we recover a residue-level evolutionary-conservation profile?
Intuition
Think of a protein as a crowded city: a residue in the core with many nearby neighbors is more likely to disrupt the whole structure when changed, so it often experiences stronger evolutionary constraint. The weighted contact number (WCN) quantifies this crowding with inverse-squared Cα distances; taking its reciprocal aligns the profile direction with sequence-conservation scores in which lower values mean greater conservation.
Why it matters
The study demonstrates a residue-level quantitative relationship between structural packing and sequence evolution and offers a structural prior when homologous sequences are scarce. However, it validates profile correlation; it does not prove that one structure fully replaces multiple-sequence alignment or directly identifies functional sites.
Prerequisites
- Understand protein residues, Cα atoms, and distances in three-dimensional coordinates.
- Know the basics of multiple-sequence alignment, phylogeny, and residue-conservation scores.
- Be able to interpret a Pearson correlation coefficient and know that correlation is neither causation nor pointwise accuracy.
- Understand that protein cores, surfaces, oligomeric interfaces, and domain packing alter local contact environments.
paper-specific guide · plain → technical → input → output → source
Method walkthrough
-
01 · Derive a packing profile from one structure
For each residue, summarize how close all other residues are into a crowding value, then convert the values into a comparable residue-by-residue curve.
Technical reading: Using Cα distances, define WCN_i = Σ_{j≠i} 1/r_ij²; normalize the WCN profile to z-scores and use reciprocal WCN (rWCN) for comparison with the conservation profile.
Input: Cα coordinates for every residue in one protein structure and the chosen chain/subunit/domain scope.
Output: A normalized residue-level rWCN packing profile.
Boundary: WCN is a structural-packing descriptor without amino-acid identity, sequence-homology information, or an explicit energy model; its values change with the included subunits/domains.
PDF pp. 9–10, Methods, Equation 1
-
02 · Build an independent sequence-conservation reference
Separately estimate conservation at each position from homologous sequences and a phylogenetic tree, providing a reference against which to validate the structural profile.
Technical reading: Following the ConSurf protocol, search SwissProt with PSI-BLAST, remove redundant/closely related sequences with CD-HIT, construct an MSA with MUSCLE, build a neighbor-joining tree in Rate4Site, estimate position-specific scores with the empirical Bayesian method, smooth over a five-residue window, and z-score normalize.
Input: The query sequence, SwissProt homologs, and phylogenetic relationships derived from the MSA.
Output: A normalized conservation profile derived from sequence and evolutionary information.
Boundary: This sequence profile is a benchmark reference, not an input to the single-structure method; insufficient homolog diversity changes the reference itself.
PDF p. 10, Methods, Sequence conservation profiles
-
03 · Compare profiles and test structural context
Correlate the two curves, then recompute packing with a single chain, the full complex, or the complete multidomain chain to see whether interfaces restore or distort the signal.
Technical reading: The main benchmark computes Pearson r for 554 nonhomologous enzymes; multichain and multidomain cases compare rWCN–conservation correlations under different coordinate scopes and contrast them with a contact-number/sequence-entropy baseline.
Input: Paired rWCN and conservation profiles plus alternative biological-assembly/domain scopes.
Output: Per-protein Pearson r values, the fraction above a threshold, and cases sensitive to structural scope.
Boundary: A higher correlation only indicates that the profiles covary; by itself it does not prove subunit coevolution, functional mechanism, or residue-level classification correctness.
PDF pp. 2–3, Figure 1; pp. 5–6, Figures 4–7; pp. 8–9, baseline comparison
Key result
For 74% of 554 enzymes, structure-derived and sequence-conservation profiles correlated above 0.5, supporting structure-encoded evolutionary constraints.
Evidence-guided deep reading
Paper facts, project readings, and teaching models are labelled separately.
paper-fact
The main relationship between packing and conservation profiles
Across 554 nonhomologous enzymes, the mean Pearson correlation between rWCN and conservation profiles is 0.57; 408/554 proteins (74%) have a coefficient greater than 0.5. This is a per-protein comparison of residue-level profiles.
For 1ONR:A, 1KZH:A, and 1FGH:A in Figure 2, the profile minima overlap and marked catalytic residues lie near those minima. This supports the reading that highly packed regions are often more conserved, but it is not a sensitivity/specificity evaluation of a catalytic-site predictor.
Source locator: PDF pp. 1–2, Abstract, Figure 1 and Figure 2
paper-fact
Structural scope is itself an assumption
For tetrameric 1ARZ, chain A correlates at 0.18 when WCN uses only that chain and rises to 0.72 when all four chains are included. Conversely, 1RO7:A falls from 0.58 for one chain to 0.17 for the tetramer, showing that including the entire crystal assembly does not guarantee improvement.
In the multidomain single-chain analysis, 388/478 domains (81%) correlate better when WCN includes the complete chain, but the authors explicitly state that further study is required. Opposite cases in 1GOG:A and 1DGK:N likewise make assembly/domain choice a testable model setting rather than a fixed truth.
Source locator: PDF pp. 3–6, Figures 4–7; PDF p. 9, Figure 8
project-reading
It establishes correlated profiles, not the obsolescence of MSA
On the same 554-protein set, the Jernigan–Lustig baseline has only 7/554 proteins (1%) above r=0.5 and a mean r of 0.29, versus 408/554 and 0.57 for this study. Replacing the baseline's CN with WCN raises the count above 0.5 to 147/554, indicating that both the packing descriptor and the conservation-estimation procedure matter.
These comparisons are neither an external blind test of a residue classifier nor proof that packing is the sole cause of conservation. A better use is to treat rWCN as a structural prior and triangulate it with MSA, functional experiments, and the correct biological assembly.
Source locator: PDF pp. 8–9, Discussion and comparison with the Jernigan–Lustig method
Study design and evaluation
Data and samples
The main dataset contains 554 enzymes selected from Catalytic Site Atlas 2.2.11 with pairwise sequence identity below 25% and X-ray structures with fewer than five missing residues. The structural side uses only Cα atoms; the sequence side builds a conservation reference from SwissProt homologs through a ConSurf-like pipeline.
Baselines
- The Jernigan–Lustig method using BLASTP pairwise-alignment sequence entropy and a 9 Å cutoff contact number (CN), plus a variant in which CN is replaced by WCN.
- WCN computed for the same structure using a single subunit/domain versus the complete assembly/chain as a structural-scope contrast.
Metrics
- Pearson profile correlation
- Measures linear covariation between residue-level rWCN and conservation profiles within one protein, ranging from -1 to 1.
Boundary: It is not residue-prediction accuracy, causality, or functional-site precision and is affected by smoothing, sequence diversity, and coordinate scope. - Fraction with r > 0.5
- The proportion of the 554 proteins whose individual profile correlation exceeds the authors' descriptive threshold of 0.5.
Boundary: The value 0.5 is not an ROC-derived optimal classification cutoff and does not mean all residues are predicted correctly. - Context-induced correlation change
- Compares how much Pearson r for the same chain/domain rises or falls when different subunits or domains are included.
Boundary: It is an assembly-sensitivity signal; an increase alone does not establish coevolution or obligate interaction.
Reported result
The main dataset yields a mean Pearson r of 0.57, with 408/554 proteins (74%) above 0.5; on the same 554 proteins, the Jernigan–Lustig baseline averages r=0.29, with 7/554 (1%) above 0.5. The result supports quantifiable evolutionary signal in the packing profile of a single Cα structure but does not establish equivalence to MSA.
PDF pp. 1–2, Abstract and Figure 1; PDF pp. 8–10, baseline comparison, Methods and Dataset
teaching-model · not a reported experiment
Teaching example (not a reported experiment)
Build an rWCN conservation prior from a toy structure
Project teaching model: a small protein has Cα coordinates for six residues; residues 3 and 4 lie in a compact core, while residues 1 and 6 sit at extended termini.
- For each residue, compute distances to the other five Cα atoms, convert each contribution to 1/r², and sum them as WCN.
- Convert WCN to a reciprocal profile and normalize it; core residues have higher WCN and therefore form lower rWCN valleys.
- Label low-rWCN positions as more likely structurally constrained rather than proven functional sites, then compare with a homolog-derived conservation profile using Pearson r.
- If residue 1 contacts another subunit, recompute with the biological assembly and report the correlation change separately.
Takeaway: rWCN provides a testable structural prior; coordinate scope and the sequence reference are both part of the result.
historical-ecosystem
Evidence boundary versus FAST
WCN later becomes a SARST2 feature, but this paper is not an alignment benchmark.
Lawful source and access
11 pages · SHA-256 a7ca3434e69aca156dfb415c38c5b105986b0fb6efb0cc4409474d6e2bc116e4
Lawful institutional-repository full text, not a paywall bypass.
NYCU institutional-repository PDF
Limits and misreadings
- Correlation does not imply full replacement of multiple-sequence alignment.
- Oligomeric interfaces and conformational state alter packing signals.
Source locator map
- PDF pp. 1–2, Abstract, Results and Figures 1–2
- PDF pp. 3–6, Evolutionary coupling between subunits and Figures 4–7
- PDF pp. 8–9, Discussion and Jernigan–Lustig baseline comparison
- PDF pp. 9–10, Methods, Equation 1 and Sequence conservation profiles
- PDF p. 10, Dataset
Check understanding
What does a higher WCN usually represent?
Answer: Denser Cα packing around that residue.
Short distances receive larger 1/r² weights, so many close neighbors increase WCN.
Can 408/554 with r>0.5 be read as 74% of residues predicted correctly?
Answer: No; 74% is a fraction of proteins, not residue accuracy.
Each protein first receives one profile-level Pearson r, after which proteins above 0.5 are counted.
Why do 1ARZ and 1RO7 warn against always including every chain?
Answer: The full assembly raises the correlation for 1ARZ but lowers it for 1RO7.
Whether an interface carries evolutionary constraint depends on biological assembly and function, not automatically on crystal contacts.
Completion task: Choose a protein with a known biological assembly and compute rWCN for both one chain and the full assembly. Build one ConSurf-like reference, report profile Pearson r, coordinate missingness, homolog diversity, and assembly choice, and explain whether the scope difference supports interface constraint.
Paper-specific glossary
- Weighted contact number (WCN)
- A residue-packing descriptor formed by summing inverse-squared distances to all other Cα atoms.
- Reciprocal WCN (rWCN)
- A reciprocal form of WCN in which highly packed regions have lower values, matching the direction of the paper's conservation scores.
- Conservation profile
- A curve of evolutionary-conservation scores indexed by sequence position.
- Coordinate scope
- The choice to include a single chain, complete complex, one domain, or full chain when computing contacts.