Score exact peptide occurrences with retained source context - #451
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Vaxrank needs to rescore a reported peptide at its original occurrence without scanning additional windows or collapsing different flanks. Add public
predict_peptide_occurrencesandTopiaryPredictor.predict_from_peptide_occurrences, retaining the existing compound ID/peptide/offset identity (and sample scope when supplied) and configured mhctools models. Identical inputs share inference within sample/flank/genotype context; each original observation and its evidence remains separate.Reuse that implementation from additive
rescore_candidates, batching compatible peptides while keeping original measurements, feature names and the selected candidate universe intact. Validate model-kind/allele coverage, retain allele-free and haplotype scope, distinguish unknown flanks from known termini, and support explicitly supplied comparators with their own context. Haplotype comparator scores follow the genotype even when the deconvolved presenter changes. Policy evaluation and replay consume the result directly; no scientific ranking defaults change.The shared public
prediction_mhc_scopealso fixes the existing fragment WT join: a changed haplotype presenter no longer erases its supplied comparator score. Paired fragment/occurrence tests cover per-allele, haplotype and allele-free scope.Pair-driven tests also exposed the existing empty named-peptide schema bug. Empty and reported-all-missing batches now keep the normal prediction columns.
Validation: lint and the full local one-worker
./test.shpass: 5,025 tests, no skips or unexpected warnings. The 99 focused checks include all 20 published Vaxrank 3.35.0 consumer cases. Repeated source IDs retain distinct peptide windows through actual Vaxrank scoring; enabling flanks changes the scores. Composed LENS/pVACseq/ProteinFragment workflows demonstrate changed selection from fresh context-sensitive scores, unchanged original scores/read support, and exact CSV/TSV policy replay without inference. Mixed-kind fixture output is exercised under both pandas string modes (#454). All PR and master CI jobs passed. Merged asc668ef4c6b863d956171e95ffc81c22199abc125; clean-master./deploy.shpassed 5,025 tests and published 5.91.0. Both published artifact SHA256 digests match the local build, andv5.91.0points to that merge commit.Fixes #367. Fixes #450. Fixes #453. Fixes #454.
Historical sparse-table prediction sources with unknown model versions remain #368. Vaxrank CLI/config adoption remains openvax/vaxrank#497; this PR changes only Topiary.