See the Biology, Not Just a Score

A BRCA1 case study in how CodeXome resolves a real variant of uncertain significance using primate evolutionary evidence, not a predicted score.

Most inherited diseases are still poorly understood. Thousands of conditions trace back to changes in DNA that alter how a protein functions, and researchers trying to find the right targets for cures and drug development are often working without good tools to interpret what they're looking at. CodeXome approaches that gap with a different kind of evidence. Not a predicted pathogenicity score, but a direct record of what evolution has already tested in the primate lineage.

Let's Explore BRCA Together

Cornerstone Genomics first validated the underlying premise against ClinGen, one of the most rigorously vetted clinical variant panels, curated by expert committees under the 2015 ACMG guideline for sequence variant interpretation. Comparing 2,231 SNPs across 46 ClinGen-reviewed genes against the CodeXome primate database, the pattern held cleanly: benign mutations recur across primate lineages; pathogenic and likely-pathogenic mutations essentially don't. They turn out to be unique to human patients.

The same test, run specifically on BRCA1 and BRCA2 (2,551 SNPs, using clinical assertions from the ENIGMA consortium via BRCA Exchange), produced the identical split.

Figure 1. BRCA1/BRCA2 variants cross-referenced against CodeXome's primate database, by ENIGMA clinical category (2,551 SNPs). Pathogenic variants are almost entirely absent from the primate record; benign and likely-benign categories overlap it substantially.

That gives researchers a fast first pass: strip out the variants that recur naturally across primates, and rank what's left as diagnostic markers by functional impact.

Why Is This Residue Constrained?

Line up BRCA1 across humans, chimpanzees, and the other great apes, and the sequence is nearly featureless: clear boxes matching the human reference at almost every position. That similarity is expected given how essential BRCA1 is, but it's also a dead end: no insight into how the gene or its protein tolerates change is possible until the comparison widens to include the full primate order, not just our nearest relatives.

CodeXome gene profile view showing BRCA1 aligned across humans and the great apes, nearly identical across the row
Figure 2. BRCA1 aligned across humans and the other great apes in the CodeXome platform. Clear boxes are identical to the human reference; colored boxes mark amino acid changes in other primate genera, of which there are almost none here.

Widen the comparison, and position 323 becomes visible as a genuinely informative spot in the gene: multiple amino acid changes are tolerated there across the full primate lineage, including the specific substitution at the center of this case study, Gly323Glu.

CodeXome alignment view across the full primate order, position 323 highlighted, showing a narrow set of tolerated amino acids
Figure 3. The same region aligned across the full primate order. Position 323 (highlighted column, arrow) tolerates multiple amino acid states, including Gly323Glu.

Here's What Evolution Actually Shows

Gly323Glu sits in ClinVar as one of BRCA1's many VUS and conflicting records. CodeXome resolves a meaningful share of exactly that kind of ambiguity: across the gene, nearly 30% of BRCA1's VUS and conflicting ClinVar records turn out to be likely benign once evolutionary recurrence is factored in, because ClinVar variants, including VUS, that are shared with primates are, empirically, not pathogenic.

CodeXome gene-wide view showing ClinVar clinical variants alongside a CodeXome annotation track reclassifying a large share of them as likely benign
Figure 4. ClinVar's clinical variant calls for BRCA1 (top row) against CodeXome's own annotation track (highlighted, below). Roughly 30% of BRCA1's VUS and conflicting records resolve to likely benign.

Gly323Glu is one specific example of what that reclassification looks like up close. Reconstructing the amino acid history at this position shows that Glutamate (Glu) is the ancestral state, shared with the outgroup tree shrew (Tupaia) and retained today in New World monkeys, while Glycine (Gly) arose later, in the common ancestor of humans, gibbons, and Old World monkeys. Gly323Glu, the variant flagged in ClinVar, is a reversion to that ancestral amino acid in humans.

Primate phylogenetic tree with arrows marking where glutamate (ancestral) and glycine (derived) states occur at BRCA1 position 323
Figure 5. The evolution of Gly323Glu, revealed. Glutamate is ancestral, retained in New World monkeys and the tree shrew outgroup (red arrows); glycine arose in the common ancestor of humans, gibbons, and Old World monkeys (green arrow).

CodeXome classifies it accordingly: likely benign, inserted as its own annotation track directly beneath ClinVar's own conflicting record for the same variant.

CodeXome annotation tooltip for Gly323Glu showing it reclassified as likely benign, with the CodeXome track inserted beneath the ClinVar track
Figure 6. Gly323Glu reclassified as likely benign by CodeXome. Each green circle in the CodeXome track marks a VUS or conflicting ClinVar record resolved the same way.

How Evolution Helps Prioritize Variants

The BRCA1 case study is really a demonstration of a general workflow. A researcher uploads variant data to the CodeXome cloud platform, where it's cross-referenced against the CodeXome primate databases. Variants that recur naturally in primates are filtered out as benign, in minutes rather than the weeks, months, or years this normally takes. What's left is prioritized using an evolutionary score, then run against integrated UniProt, ClinVar, and gnomAD data, all mapped to GRCh38 coordinates, for an instant, all-in-one read on variant function correlated with predicted protein structure.

CodeXome platform view showing UniProt domain and site data, ClinVar clinical variants, and gnomAD population data all integrated and mapped to GRCh38 coordinates for BRCA1
Figure 7. UniProt, ClinVar, and gnomAD integrated in one view, mapped to GRCh38, an all-in-one read on variant function correlated with predicted protein structure across the entire gene.

That integration is what turns millions of years of evolutionary time into something actionable: it defines where and how a gene and its protein tolerate change naturally, which variants are worth a second look, and where to focus next: predicting protein alterations from natural selection, identifying key functional motifs for intervention, and ultimately informing drug design and development.

Because the underlying evidence is gene evolution rather than human population frequency, it's also agnostic to human population structure, which matters directly for resolving variant function in patients from under-represented groups.

None of this replaces a researcher's judgment. What it does is remove a genuinely large amount of noise before that judgment is needed, filtering out the natural, benign background of a gene in minutes, and leaving a shorter, better-ranked list of what's actually worth investigating. That's the shift this case study is meant to show: from a predicted score to an observed history, one residue at a time.

Ready to add evolutionary evidence?
Try the Gene Previewer for an instant look at the data, or request full platform access for your research.