Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ jobs:
python-version: "3.12"
- name: Install the released downstream consumer
run: |
bash scripts/pip_install.sh pytest 'vaxrank==3.35.0' -e .
bash scripts/pip_install.sh pytest 'vaxrank==3.36.0' -e .
python -m pip check
- name: Verify table features through DSL scoring and vaccine construction
run: python -m pytest scripts/check_vaxrank_candidates.py -q
Expand Down
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,19 @@ between published tags; older pre-5.0 notes are retained below.

For current interfaces, see the [consumer guide](docs/consumer-guide.md).

## 5.92.0

- Add all-candidate self matching within an explicit same-length Hamming radius
and exact matching across complete vaccine windows. Retain every origin,
including shared CTA/non-CTA sequences, with caller-resolved exclusions (#455).
- Preserve supplied presentation observations, allele attribution, conflicting
predictions and same-allele coverage separately from search completeness.
Typed CSV/TSV retains evidence and search identity; missing evidence remains
unknown. Neither presentation nor similarity establishes TCR recognition.
- Test against Vaxrank 3.36.0's published `builtin:openvax-v1` bundle. A synthetic
consumer workflow changes the selected window using self evidence while
retaining both intended epitopes and alleles; Topiary adds no selection rule.

## 5.91.0

- Score explicitly identified peptide occurrences with their original flanks,
Expand Down
2 changes: 1 addition & 1 deletion RELEASING.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ CI installs a pinned published Vaxrank release and runs the real candidate
scoring and vaccine-construction workflow. To repeat it locally:

```bash
"$PYTHON" -m pip install 'vaxrank==3.35.0'
"$PYTHON" -m pip install 'vaxrank==3.36.0'
"$PYTHON" -m pip check
"$PYTHON" -m pytest scripts/check_vaxrank_candidates.py -q
```
Expand Down
12 changes: 12 additions & 0 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,6 +70,18 @@ symmetric conservative transformation, and O/U/X/* or unrecognized characters
receive the symmetric worst-case canonical distance (15). Both functions
return fresh read-only int8 arrays, so callers never share mutable matrix state.

## Self-match evidence

`match_self_peptides(reference, peptides, *, max_mismatches=0, alleles=None,
excluded_gene_ids=(), observations=None, predictions=None)` returns every
same-length Hamming match within the radius, with all source origins and
supplied presentation/prediction records. `reference.match_candidates(...)`
delegates to it. `self_matches_in_windows(windows, reference, *, peptide_lengths,
...)` searches exact peptides at every requested window position. Results use
ordinary typed CSV/TSV, with explicit coverage and evidence identity.
See [self evidence](self_proteome.md#all-match-evidence) for the input schemas,
policy composition and interpretation of missing results.

## TopiaryPredictor

| Parameter | Type | Description |
Expand Down
31 changes: 15 additions & 16 deletions docs/combined-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,9 +54,11 @@ version-qualified expression for this simple case.

`SelectionPolicy` saves the exact candidate filter, score, model/version
selections, ranking direction, duplicate policy and strata under a stable name.
It has no implicit scientific recipe. Topiary does not currently ship an
official `openvax-v1` preset; that name can identify a reviewed definition you
freeze, and later definitions can use new names.
It has no implicit scientific recipe. **Vaxrank owns the published
`builtin:openvax-v1` bundle**, shipped with experimental overlays in
[Vaxrank 3.36.0 / PR #565](https://github.com/openvax/vaxrank/pull/565).
Topiary evaluates its policy subtree and tests against that released bundle;
it does not maintain a second definition. Give changed policies new names.

```python
from topiary import (
Expand Down Expand Up @@ -115,7 +117,7 @@ site-level evidence remains [#288](https://github.com/openvax/topiary/issues/288
### Composed consumer configuration

Vaxrank owns YAML composition and window/construct configuration. Its repeated
`--config openvax-v1.yaml --config overrides.yaml` workflow merges mappings
`--config builtin:openvax-v1 --config overrides.yaml` workflow merges mappings
left to right and replaces lists/scalars. The Topiary integration boundary is
the **already-composed policy subtree**:

Expand All @@ -142,16 +144,13 @@ the old expression. Model selections are mappings; version selections are lists
of `{kind, method, version}` records, so an override replaces that entire list.
Topiary does not merge files or parse unrelated consumer settings.

Vaxrank adoption of this subtree remains
[Vaxrank #497](https://github.com/openvax/vaxrank/issues/497); the example above
describes the consumer boundary, not a newly supported Vaxrank configuration
key. Its current occurrence-level scorer can consume `policy.score_by`,
`policy.filter_by`, and the selection mappings through the existing public
`EvalContext` / `apply_filter` APIs, retaining its explicit grouping and
per-occurrence `alleles` lookup. Vaxrank's post-score minimum gate, missing-score
fill, window rules, RNA weighting and construct settings must also be frozen in
its bundle. A cutoff inside a score expression must not become a destructive
pre-filter when capturing that baseline.
Vaxrank 3.36.0 consumes this subtree in its shipped configuration bundle.
Its post-score minimum gate, missing-score fill, window rules, RNA weighting
and construct settings belong to that bundle too. The consumer fixture in
`scripts/check_vaxrank_candidates.py` loads `builtin:openvax-v1` through
Vaxrank's public config loader and compares its numeric scores with
`evaluate_selection_policy`. A cutoff inside a score expression must not become
a destructive pre-filter when replaying that baseline.

The representative ranking above is one view. `evaluate_selection_policy`
retains every input row and evaluates explicit occurrence identities before
Expand All @@ -175,8 +174,8 @@ evaluated.evidence.to_tsv("complete-evaluation.tsv")
replayed = replay_selection_policy(read_tsv("complete-evaluation.tsv"))
```

This example freezes a scoring convention; it does not establish an official
`openvax-v1` recipe or a calibrated probability of immunogenicity. The cutoff
This example illustrates a scoring convention; use Vaxrank's published bundle
for `openvax-v1`. Neither expression is a calibrated probability of immunogenicity. The cutoff
remains inside the score, with a separate inclusive minimum-score gate.
`raw_score` preserves missingness even when `score_fill=0` supplies an effective
zero. Pre-filtered groups are not scored and cannot be restored by filling.
Expand Down
144 changes: 115 additions & 29 deletions docs/self_proteome.md
Original file line number Diff line number Diff line change
@@ -1,21 +1,107 @@
# Cross-reactivity: `self_nearest_*` via `SelfProteome`
# Self-sequence and presentation evidence

Given a mutant peptide, find its closest match in a reference proteome
of healthy self peptides. The distance (and the source gene of the
match) is a cross-reactivity risk signal: mutant neoantigens that look
a lot like a real self-peptide may trigger T-cell cross-reactivity
against healthy tissue.
`SelfProteome` searches an explicit reference corpus. Sequence similarity,
predicted binding and observed presentation are separate facts; none establishes
TCR recognition. Peptide-loaded or minigene recognition can also differ from
recognition of native protein processing and presentation ([primary study](https://www.nature.com/articles/s41541-023-00713-y)).

This page covers the `SelfProteome` class and the `self_nearest_*`
columns it adds to `TopiaryPredictor` output.
Use [all-match evidence](#all-match-evidence) for complete vaccine-window searches
or multiple similar candidates. The existing `nearest()` and `self_nearest_*`
predictor columns remain a single-nearest-sequence view.

> **Status.** `SelfProteome` currently exposes the sequence-nearest
> axis: substitutions plus 1aa indel neighbors against a scoped self
> proteome (`include="all"`, `"non_cta"`, `"protected_tissues"`, or a
> callable). Binding-aware axes (`self_mimic_*`, `self_strongest_nearby_*`)
> and the structured full-candidate column (`self_nearest_candidates`)
> require MHC prediction on candidate peptides and are tracked under
> [#124](https://github.com/openvax/topiary/issues/124).
> remain [#412](https://github.com/openvax/topiary/issues/412).
> All-match evidence accepts supplied observations and predictions; it does not
> run models, rank by binding or infer TCR recognition.

## All-match evidence

```python
from topiary import SelfProteome, match_self_peptides, self_matches_in_windows

# Synthetic reference. In production, retain an unfiltered corpus so a shared
# sequence keeps both CTA and non-CTA origins. oncoref owns CTA membership.
reference = SelfProteome.from_peptides(
{"CTA": "GILGFVFTL", "healthy": "GILGFVFTL", "near": "SIINFEKM"},
peptide_lengths=[8, 9],
)
windows = {"full": "GILGFVFTLGGGSIINFEKLELAGIGILT", "trimmed": "SIINFEKLELAGIGILT"}
exact = self_matches_in_windows(
windows, reference, peptide_lengths=[8, 9],
alleles={name: ["A0201", "B0702"] for name in windows},
excluded_gene_ids={"CTA"}, # illustrative caller-resolved exclusion
)
similar = match_self_peptides(reference, ["SIINFEKL"], max_mismatches=1, alleles=["A0201"])
exact.to_tsv("window-self-evidence.tsv")
```

`reference.match_candidates(...)` delegates to `match_self_peptides`. Queries
retain order and duplicates through `query_index`; every matched sequence and
gene/transcript/reference-offset origin has a row. `self_match_id` identifies the
reference occurrence across queries. Window output also carries `window_id`,
`window_sequence` and every zero-based `peptide_offset`. The exact adapter checks
all requested lengths and offsets, not only mutation-overlapping peptides.

Exclusions flag `self_in_scope=False`; they never delete origins. A reference
already filtered with `include="non_cta"` cannot recover removed CTA origins.
For an Ensembl corpus use `include="all"` and resolve the exclusion set with
`cta_gene_ids(source="oncoref")` for the appropriate species/tier. A sequence
match alone is not evidence that the peptide was presented.

Both functions accept `observations=` and `predictions=` as tables or record
lists. They preserve multiple records in typed list/dict columns:

| Input | Required fields | Interpretation |
|---|---|---|
| Observations | `evidence_id`, `peptide`, `source`, `evidence_kind`, `allele_assignment` | `evidence_kind` is `observed` or `predicted`; allele assignment is `confirmed`, `predicted` or `unknown`. Confirmed/predicted assignments require an explicit `allele`. |
| Predictions | `peptide`, `allele`, `kind`, `value`, `prediction_method_name`, `predictor_version` | Preserve supplied model facts, including unknown versions, missing values and conflicts; no averaging or model choice. |

Observation records may carry tissue, assay, source version and other
JSON-compatible provenance. An optional `gene_id` restricts attribution to that
origin; its absence does not identify the producing gene. An unresolved
`allele_set` remains a genotype, not confirmed restriction. Alleles are parsed
with mhcgnomes, including non-human alleles.

`self_observations` and `self_predictions` retain the records.
`self_same_allele_predictions` contains only finite numeric `value` measurements
at the query allele with per-allele MHC scope. Haplotype presenter labels do not
become per-allele evidence. This reports any supplied measurement, not complete
coverage of a desired model panel. Other-allele records remain in
`self_predictions`. Observation flags distinguish confirmed and inferred
same-allele assignments; unreported observations remain unknown.

Coverage is explicit:

| Field | Meaning |
|---|---|
| `self_search_status` | `matched`, `no_match_in_scope`, or `unassessed`. |
| `self_search_complete`, `self_search_reason` | Whether the requested canonical, same-length search is complete; unavailable lengths and unsupported query/reference sequences have explicit reasons. A match can coexist with incomplete coverage. |
| `self_observation_status` | Observation input absent, a record reported, or no record reported for this match/origin. |
| `self_prediction_status` | Prediction input absent, query allele unknown, same-allele measurement reported, or not reported. |

`no_match_in_scope` is negative evidence only within the declared reference,
lengths and Hamming radius. It does not mean no presentable peptide or no risk.
Indels, substitutions outside the radius and unsearched reference data are not
assessed. Comparisons use chunks of 65,536 reference rows; exact queries reuse
the reference index. Metadata under `extra['self_search']` records the reference
identity, parameters, coverage counts and hashes of normalized supplied evidence.
`extra['self_window_coverage']` also records requests too short for a peptide.

CSV/TSV preserves these fields and nested records. This is an evidence table,
not a fabricated prediction table: join it to candidate measurements at the
explicit query/allele identity before using a prediction policy. Preserve all
matches and their IDs when there are several per candidate. The composed test
in `tests/test_consumer_workflows.py` checks named-criterion unknown handling and
exact replay after joining and saving the evidence.

Vaxrank owns window selection. `scripts/check_vaxrank_candidates.py` consumes its
released `builtin:openvax-v1` bundle and shows that self evidence changes a window
choice while retaining both intended epitopes and their alleles. That synthetic
fixture uses 100% target-score retention and checks every required target; it
does not establish that a general aggregate-score threshold protects each target.

## Basic usage

Expand All @@ -24,7 +110,7 @@ from topiary import SelfProteome, TopiaryPredictor
from mhctools import NetMHCpan

ref = SelfProteome.from_ensembl(species="human")
# Default include="non_cta" strips CTAs via pirlygenes.
# Default include="non_cta" strips CTAs using oncoref membership via pirlygenes.

predictor = TopiaryPredictor(
models=NetMHCpan,
Expand Down Expand Up @@ -64,8 +150,11 @@ Three construction modes:
|---|---|---|
| `"all"` | Whole proteome, no filter | — |
| `"non_cta"` (default for human Ensembl) | Remove CTA genes | `cta_source="pirlygenes"` default; set / callable accepted |
| `"protected_tissues"` | Keep only genes expressed in named tissues | `tissues=[…]`, `tissue_source="hpa"`/`"gtex"`, `min_expression=…` |
| callable | Arbitrary `gene → bool` filter | — |
| `"protected_tissues"` | Keep only genes expressed in named tissues | `tissues=[…]`, `min_tissue_ntpm=…` for human HPA data; or explicit `tissue_gene_ids={…}` for any species |
| callable | Arbitrary `gene_id → bool` filter | — |

Tissue expression selects the reference genes; it is not evidence of peptide
presentation. Supply presentation observations separately when available.

**Human users** get zero-config `include="non_cta"` via pirlygenes:

Expand All @@ -74,7 +163,7 @@ ref = SelfProteome.from_ensembl(species="human")
```

**Non-human users** must either use `include="all"` or supply their own
CTA source, because pirlygenes is human-only today:
CTA source, because oncoref's CTA membership is human-only today:

```python
ref = SelfProteome.from_ensembl(species="mouse", release=102, include="all")
Expand All @@ -89,8 +178,7 @@ ref = SelfProteome.from_ensembl(
```

A non-human `include="non_cta"` call without `cta_source=` raises at
construction — silent unfiltered results would be a misleading
cross-reactivity signal.
construction — silent unfiltered results would misstate the reference scope.

## Non-Ensembl sources

Expand All @@ -115,12 +203,12 @@ ref = SelfProteome.from_peptides(
## Reference version

Every row of the output carries a `self_nearest_reference_version`
string. Two runs produce interchangeable `self_nearest_*` values iff the
strings match. Its form is
string. Matching strings identify the same indexed reference; comparing results
also requires the same queries, metric and search settings. Its form is
`{source}-{species}[-{release}]+include-{scope}+sha256:{digest}`:

```
ensembl-human-115+include-non_cta+cta-pirlygenes-6.0.4+sha256:3f2a9c1e7b40
ensembl-human-115+include-non_cta+cta-oncoref-VERSION+sha256:3f2a9c1e7b40
ensembl-mouse-102+include-all+sha256:91d0c4e2a8b7
fasta-fasta+include-callable-keep_named+sha256:0be51d2c9f63
peptides-synthetic+include-all+sha256:bd8a4e854d1c
Expand Down Expand Up @@ -153,12 +241,10 @@ when enabled.

- Construction: one pass over the reference proteome extracts every
L-mer for each configured length, dedupes per length, and encodes
into a `(M, L) int8` array. For human non-CTA Ensembl × length 9,
expect ~200k rows.
- Lookup: per query, the full reference array is compared in one SIMD
operation. Chunked to bound peak memory. Typical throughput for
~200k reference × ~10k queries is seconds.

Seed-and-extend indexing and 1aa indel candidates are queued in
[#124](https://github.com/openvax/topiary/issues/124) — benchmark
decides the default algorithm.
into a `(M, L) int8` array. Corpus size depends on the release, scope and lengths.
- Lookup: vectorized comparisons are chunked to bound temporary memory.
Runtime depends on the corpus and query count; exact matching reuses a hash index.

The existing `nearest()` path checks one-residue indel neighbors. The new
all-match API is limited to same-length Hamming candidates; broader candidate
search and binding-ranked axes remain [#412](https://github.com/openvax/topiary/issues/412).
Loading
Loading