Skip to content

Carry a self-presentation evidence tier on nearest-self, without owning the data #224

Description

@iskandr

Split out of the #124 discussion, which it does not belong inside.

#124 asks which genes are excluded from the self proteome — a
membership question, answered by oncoref, and settled. This is a different
question: is this self peptide actually observed on healthy tissue, and at
which allele?

Observed healthy-tissue presentation is a distinct evidence axis that sequence similarity alone cannot express. It establishes presentation under the reported assay conditions, not TCR recognition. Preserve confirmed versus inferred allele attribution, tissue, source and assay provenance; absence from the supplied observations is not proof of absence in healthy tissue.

The tiers

There are three, and they are not topiary's invention — tsarina already draws
exactly these lines:

Tier Evidence
Observed, allele confirmed Seen in MS data on a cell line expressing a single HLA allele (--mono-allelic-only)
Observed, allele predicted Seen in MS data, but the allele was assigned by NetMHCpan / MHCflurry from a multi-allelic sample
Predicted only No MS observation; presentation inferred

tsarina.ms_evidence.cta_healthy_tissue_ms_hits already returns this at
peptide × tissue × allele granularity, with allele (recorded
mhc_restriction) versus allele_set (candidate set for multiallelic
samples) — which is the top two tiers at row level — plus a vital_organ
flag.

What topiary should and should not do

Not depend on tsarina. tsarina/scoring.py does
from topiary import TopiaryPredictor, so tsarina sits above topiary; a
dependency the other way is a cycle and a layering inversion. Same reasoning
that keeps the CTA list out of topiary in #124.

Carry the tier, don't own the data. The injection points already exist:

  • SelfProteome.from_peptides takes caller-supplied peptides, so a consumer
    can build a proteome from observed immunopeptidome hits rather than from a
    translated transcriptome.
  • field_provenance (5.32.0) is the same idea one level over —
    strength-of-evidence attached per fact, with a closed validated vocabulary
    (measured / approximated / synthesized). The MS tiers want the same
    treatment: a small closed vocabulary, validated, carried rather than
    inferred.
  • The tier then needs to reach the ranking DSL, so a safety filter can say
    "reject if the nearest self peptide was observed at this allele" rather than
    only "reject if edit distance is small".

Open questions

  • Where the tier lives: a self_nearest_evidence column alongside the other
    self_nearest_* axes, or provenance on the proteome's peptides that
    nearest() propagates.
  • Whether the vocabulary is the three tiers above or something more general,
    given field_provenance already exists and is close in spirit but not the
    same axis (how a value was obtained vs how strongly a presentation is
    evidenced).
  • Whether "observed on healthy tissue at all" and "observed at this allele"
    are one ordered scale or two independent flags. They are not obviously
    totally ordered: an observation at a different allele is weaker evidence for
    this allele but strong evidence the peptide is presented somewhere.

Not blocking anything. Filing because the distinction is real and #124 as
written does not cover it.

Vaxrank feedback (2026-10-02): preserve multiple self candidates and all origins; provide exact window-contained self ligands separately from similar intended-target matches. Same-allele, other-allele and unresolved observations should remain independent facts rather than an invented total safety ordering. Primary study: https://www.nature.com/articles/s41541-023-00713-y .

Activity

  1. iskandr commented on Oct 2, 2026

    @iskandr
    ContributorAuthor

    Topiary 5.92.0 / PR #456 now carries caller-supplied presentation records through the all-match evidence API. It preserves evidence IDs, source/tissue/assay facts, observed versus predicted evidence, confirmed/predicted/unknown allele assignment, optional gene attribution and unresolved genotypes. Independent observation flags avoid inventing a total ordering; unreported evidence remains unknown.

    Named criteria can consume the joined evidence columns with explicit unknown handling. A composed consumer test saves/reloads the full evidence and reproduces the decisions exactly. CTA/reference ownership remains oncoref; there is no tsarina dependency or automatic data acquisition.

    The legacy single-nearest TopiaryPredictor columns do not automatically acquire these supplied observations. Any nearest-view adapter or acquisition integration still needs an explicit contract; this bounded change uses match_self_peptides / SelfProteome.match_candidates and self_matches_in_windows. Presentation does not establish TCR recognition.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions