Split out of the #124 discussion, which it does not belong inside.
#124 asks which genes are excluded from the self proteome — a
membership question, answered by oncoref, and settled. This is a different
question: is this self peptide actually observed on healthy tissue, and at
which allele?
Observed healthy-tissue presentation is a distinct evidence axis that sequence similarity alone cannot express. It establishes presentation under the reported assay conditions, not TCR recognition. Preserve confirmed versus inferred allele attribution, tissue, source and assay provenance; absence from the supplied observations is not proof of absence in healthy tissue.
The tiers
There are three, and they are not topiary's invention — tsarina already draws
exactly these lines:
| Tier |
Evidence |
| Observed, allele confirmed |
Seen in MS data on a cell line expressing a single HLA allele (--mono-allelic-only) |
| Observed, allele predicted |
Seen in MS data, but the allele was assigned by NetMHCpan / MHCflurry from a multi-allelic sample |
| Predicted only |
No MS observation; presentation inferred |
tsarina.ms_evidence.cta_healthy_tissue_ms_hits already returns this at
peptide × tissue × allele granularity, with allele (recorded
mhc_restriction) versus allele_set (candidate set for multiallelic
samples) — which is the top two tiers at row level — plus a vital_organ
flag.
What topiary should and should not do
Not depend on tsarina. tsarina/scoring.py does
from topiary import TopiaryPredictor, so tsarina sits above topiary; a
dependency the other way is a cycle and a layering inversion. Same reasoning
that keeps the CTA list out of topiary in #124.
Carry the tier, don't own the data. The injection points already exist:
SelfProteome.from_peptides takes caller-supplied peptides, so a consumer
can build a proteome from observed immunopeptidome hits rather than from a
translated transcriptome.
field_provenance (5.32.0) is the same idea one level over —
strength-of-evidence attached per fact, with a closed validated vocabulary
(measured / approximated / synthesized). The MS tiers want the same
treatment: a small closed vocabulary, validated, carried rather than
inferred.
- The tier then needs to reach the ranking DSL, so a safety filter can say
"reject if the nearest self peptide was observed at this allele" rather than
only "reject if edit distance is small".
Open questions
- Where the tier lives: a
self_nearest_evidence column alongside the other
self_nearest_* axes, or provenance on the proteome's peptides that
nearest() propagates.
- Whether the vocabulary is the three tiers above or something more general,
given field_provenance already exists and is close in spirit but not the
same axis (how a value was obtained vs how strongly a presentation is
evidenced).
- Whether "observed on healthy tissue at all" and "observed at this allele"
are one ordered scale or two independent flags. They are not obviously
totally ordered: an observation at a different allele is weaker evidence for
this allele but strong evidence the peptide is presented somewhere.
Not blocking anything. Filing because the distinction is real and #124 as
written does not cover it.
Vaxrank feedback (2026-10-02): preserve multiple self candidates and all origins; provide exact window-contained self ligands separately from similar intended-target matches. Same-allele, other-allele and unresolved observations should remain independent facts rather than an invented total safety ordering. Primary study: https://www.nature.com/articles/s41541-023-00713-y .
Split out of the #124 discussion, which it does not belong inside.
#124 asks which genes are excluded from the self proteome — a
membership question, answered by oncoref, and settled. This is a different
question: is this self peptide actually observed on healthy tissue, and at
which allele?
Observed healthy-tissue presentation is a distinct evidence axis that sequence similarity alone cannot express. It establishes presentation under the reported assay conditions, not TCR recognition. Preserve confirmed versus inferred allele attribution, tissue, source and assay provenance; absence from the supplied observations is not proof of absence in healthy tissue.
The tiers
There are three, and they are not topiary's invention — tsarina already draws
exactly these lines:
--mono-allelic-only)tsarina.ms_evidence.cta_healthy_tissue_ms_hitsalready returns this atpeptide × tissue × allele granularity, with
allele(recordedmhc_restriction) versusallele_set(candidate set for multiallelicsamples) — which is the top two tiers at row level — plus a
vital_organflag.
What topiary should and should not do
Not depend on tsarina.
tsarina/scoring.pydoesfrom topiary import TopiaryPredictor, so tsarina sits above topiary; adependency the other way is a cycle and a layering inversion. Same reasoning
that keeps the CTA list out of topiary in #124.
Carry the tier, don't own the data. The injection points already exist:
SelfProteome.from_peptidestakes caller-supplied peptides, so a consumercan build a proteome from observed immunopeptidome hits rather than from a
translated transcriptome.
field_provenance(5.32.0) is the same idea one level over —strength-of-evidence attached per fact, with a closed validated vocabulary
(
measured/approximated/synthesized). The MS tiers want the sametreatment: a small closed vocabulary, validated, carried rather than
inferred.
"reject if the nearest self peptide was observed at this allele" rather than
only "reject if edit distance is small".
Open questions
self_nearest_evidencecolumn alongside the otherself_nearest_*axes, or provenance on the proteome's peptides thatnearest()propagates.given
field_provenancealready exists and is close in spirit but not thesame axis (how a value was obtained vs how strongly a presentation is
evidenced).
are one ordered scale or two independent flags. They are not obviously
totally ordered: an observation at a different allele is weaker evidence for
this allele but strong evidence the peptide is presented somewhere.
Not blocking anything. Filing because the distinction is real and #124 as
written does not cover it.
Vaxrank feedback (2026-10-02): preserve multiple self candidates and all origins; provide exact window-contained self ligands separately from similar intended-target matches. Same-allele, other-allele and unresolved observations should remain independent facts rather than an invented total safety ordering. Primary study: https://www.nature.com/articles/s41541-023-00713-y .