Skip to content

Score exact peptide occurrences with retained source context - #451

Merged
iskandr merged 4 commits into
masterfrom
feat/contextual-peptide-rescoring
Oct 2, 2026
Merged

iskandr merged 4 commits into
masterfrom
feat/contextual-peptide-rescoring

Conversation

@iskandr

@iskandr iskandr commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Vaxrank needs to rescore a reported peptide at its original occurrence without scanning additional windows or collapsing different flanks. Add public predict_peptide_occurrences and TopiaryPredictor.predict_from_peptide_occurrences, retaining the existing compound ID/peptide/offset identity (and sample scope when supplied) and configured mhctools models. Identical inputs share inference within sample/flank/genotype context; each original observation and its evidence remains separate.

Reuse that implementation from additive rescore_candidates, batching compatible peptides while keeping original measurements, feature names and the selected candidate universe intact. Validate model-kind/allele coverage, retain allele-free and haplotype scope, distinguish unknown flanks from known termini, and support explicitly supplied comparators with their own context. Haplotype comparator scores follow the genotype even when the deconvolved presenter changes. Policy evaluation and replay consume the result directly; no scientific ranking defaults change.

The shared public prediction_mhc_scope also fixes the existing fragment WT join: a changed haplotype presenter no longer erases its supplied comparator score. Paired fragment/occurrence tests cover per-allele, haplotype and allele-free scope.

Pair-driven tests also exposed the existing empty named-peptide schema bug. Empty and reported-all-missing batches now keep the normal prediction columns.

Validation: lint and the full local one-worker ./test.sh pass: 5,025 tests, no skips or unexpected warnings. The 99 focused checks include all 20 published Vaxrank 3.35.0 consumer cases. Repeated source IDs retain distinct peptide windows through actual Vaxrank scoring; enabling flanks changes the scores. Composed LENS/pVACseq/ProteinFragment workflows demonstrate changed selection from fresh context-sensitive scores, unchanged original scores/read support, and exact CSV/TSV policy replay without inference. Mixed-kind fixture output is exercised under both pandas string modes (#454). All PR and master CI jobs passed. Merged as c668ef4c6b863d956171e95ffc81c22199abc125; clean-master ./deploy.sh passed 5,025 tests and published 5.91.0. Both published artifact SHA256 digests match the local build, and v5.91.0 points to that merge commit.

Fixes #367. Fixes #450. Fixes #453. Fixes #454.

Historical sparse-table prediction sources with unknown model versions remain #368. Vaxrank CLI/config adoption remains openvax/vaxrank#497; this PR changes only Topiary.

@coveralls

coveralls commented Oct 2, 2026 •

Copy link
Copy Markdown

Coverage Status

coverage: 93.261% (-0.06%) from 93.324% — feat/contextual-peptide-rescoring into master

@iskandr
iskandr marked this pull request as ready for review October 2, 2026 04:38
@iskandr
iskandr merged commit c668ef4 into master Oct 2, 2026
12 checks passed
@iskandr
iskandr deleted the feat/contextual-peptide-rescoring branch October 2, 2026 05:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants