Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ jobs:
python-version: "3.12"
- name: Install the released downstream consumer
run: |
bash scripts/pip_install.sh pytest 'vaxrank==3.32.0' -e .
bash scripts/pip_install.sh pytest 'vaxrank==3.35.0' -e .
python -m pip check
- name: Verify table features through DSL scoring and vaccine construction
run: python -m pytest scripts/check_vaxrank_candidates.py -q
Expand Down
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,18 @@ between published tags; older pre-5.0 notes are retained below.

For current interfaces, see the [consumer guide](docs/consumer-guide.md).

## 5.90.0

- Evaluate saved policies at explicit occurrence/allele identity before selecting
representatives. Retain all evidence, concrete genotype declarations,
source-local model choices, score fill/gates and exact replay context (#444).
- Add reusable named DSL criteria, explicit composition and ordered tie-breaks,
with pass/fail/unknown/not-applicable/not-evaluated audit records. Schema 2
preserves expanded definitions; schema-1 definitions keep their digests (#445).
- Fix context derivation with genotype callbacks and changed mappings (#447).
- Verify released Vaxrank 3.35.0 scoring, window selection, peptide/mRNA
construction and native dataset reload against the shared policy evaluator.

## 5.89.0

- Save and replay named `SelectionPolicy` definitions with complete DSL expressions,
Expand Down
2 changes: 1 addition & 1 deletion RELEASING.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ CI installs a pinned published Vaxrank release and runs the real candidate
scoring and vaccine-construction workflow. To repeat it locally:

```bash
"$PYTHON" -m pip install 'vaxrank==3.32.0'
"$PYTHON" -m pip install 'vaxrank==3.35.0'
"$PYTHON" -m pip check
"$PYTHON" -m pytest scripts/check_vaxrank_candidates.py -q
```
Expand Down
22 changes: 22 additions & 0 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -663,3 +663,25 @@ available_properties()
```

Groups: `"core"`, `"manufacturability"`, `"immunogenicity"`. See [Peptide Properties](properties.md) for details.

### Occurrence policy and criterion evaluation

- `evaluate_selection_policy(result, policy, group_keys=..., alleles=...,
source_contexts=...)` returns `PolicyEvaluation` with complete `evidence`,
`occurrences`, `selected` and criterion `audit` views.
- `replay_selection_policy(evidence)` repeats evaluation from the embedded
definition and materialized runtime context without prediction.
- `select_policy_representatives(evaluation, candidate_keys=..., strata=...)`
selects actual occurrences and links alternatives without adding support.
- `describe_evaluation_context(context)` materializes grouping, model choices
and per-occurrence genotype declarations for replay.
- `candidate_identifier(sample, peptide, allele)` shares combined/projected
pMHC identity, using mhcgnomes for allele canonicalization.
- `SelectionCriterion`, `RankingTerm`, and `resolve_selection_expression`
define and expand named DSL criteria with role/cycle validation.
- `evaluate_selection_criteria(policy, context, references=...)` exposes named
measurements and reasons; `evaluate_filter(frame, expression, context=...)`
exposes group decisions without discarding rows.

See [the complete source-policy workflow](combined-sources.md#named-criteria-and-audit-decisions)
for schema compatibility, source-local defaults, unknown handling and persistence.
127 changes: 119 additions & 8 deletions docs/combined-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,14 +153,125 @@ fill, window rules, RNA weighting and construct settings must also be frozen in
its bundle. A cutoff inside a score expression must not become a destructive
pre-filter when capturing that baseline.

The representative ranking above is one view, not the complete construction
interface. Shared occurrence-level evaluation with retained alternatives is
tracked in [#444](https://github.com/openvax/topiary/issues/444). Optional named
criteria, references and pass/fail/unknown audit records are tracked in
[#445](https://github.com/openvax/topiary/issues/445); schema 1 currently accepts
direct DSL expressions and rejects unknown criteria/registry fields. These
layers will reuse the same DSL rather than require a second evaluator in
Vaxrank.
The representative ranking above is one view. `evaluate_selection_policy`
retains every input row and evaluates explicit occurrence identities before
representative selection or vaccine-window construction:

```python
from topiary import (
SelectionPolicy, evaluate_selection_policy, replay_selection_policy,
select_policy_representatives, read_tsv,
)

policy = SelectionPolicy(
"example-baseline", "(affinity.value < 5000) * affinity.value.logistic_normalized(350, 150)",
score_fill=0.0, min_score=1e-5, duplicates="best",
)
evaluated = evaluate_selection_policy(combined, policy)
occurrences = evaluated.occurrences # includes excluded and unscorable alternatives
eligible = evaluated.selected # no candidate collapse
representatives = select_policy_representatives(evaluated)
evaluated.evidence.to_tsv("complete-evaluation.tsv")
replayed = replay_selection_policy(read_tsv("complete-evaluation.tsv"))
```

This example freezes a scoring convention; it does not establish an official
`openvax-v1` recipe or a calibrated probability of immunogenicity. The cutoff
remains inside the score, with a separate inclusive minimum-score gate.
`raw_score` preserves missingness even when `score_fill=0` supplies an effective
zero. Pre-filtered groups are not scored and cannot be restored by filling.
With no minimum gate, missing scores remain eligible but unranked, matching the
existing candidate-ranker convention.

For direct Vaxrank frames, pass
`group_keys=["prediction_id", "peptide", "peptide_offset", "allele"]` and the
consumer's per-occurrence `alleles` mapping/callback. The callback is evaluated
once per peptide identity and saved as concrete declarations, never Python code.
Projected allele groups carry `supporting_rows` links, not duplicated RNA counts.
`evidence_rows` links the group's original observations. These positions refer
to `evaluated.evidence.long_df`; retain that complete long-form result when
saving an evaluation. Input columns and metadata remain available there.

Use `source_contexts={label: {...}}` when sources have independent model defaults.
Every source label must be named; each mapping may replace `default_methods`,
`default_versions`, `kind_support`, and `alleles`. Filtering and scoring share
those choices. Definitions remain separate from runtime contexts, which are
stored alongside decisions in `extra["policy_evaluation"]`. Replay uses the
complete evidence and stored contexts and never invokes prediction. Model
defaults resolve ambiguity in the existing DSL; they are not a requirement that
all input measurements come from the selected model.

`select_policy_representatives` chooses an actual eligible occurrence, without
summing support, and records all alternative occurrence IDs, including excluded
alternatives. It records the choice and its runtime keys in the evaluation
evidence metadata; replay repeats that explicit choice and `audit` links
criteria to the selected occurrence and rationale. Direct consumers supply
their `candidate_keys` and `strata` if
combined-source candidate columns are absent. Missing values sort last; stable
input order breaks exact ties. Topiary owns these generic decisions; Vaxrank
still owns window geometry, source admission and construct assembly.

### Named criteria and audit decisions

Criteria reuse existing DSL expressions through an explicit namespace. This
example is synthetic policy content, not a recommended processing weight:

```python
from topiary import SelectionCriterion, RankingTerm

binding = SelectionCriterion("binding", "affinity.value < 500", "eligibility")
processing = SelectionCriterion(
"processing", "peptide_view(proteasome_cleavage.score)", "score",
)
policy = SelectionPolicy(
"example-processing", 'criterion("processing")',
filter_by='criterion("binding")', criteria=(binding, processing),
ranking_by=(RankingTerm("n_rna_alt", ascending=False),),
unknown="exclude", duplicates="best",
)
evaluated = evaluate_selection_policy(combined, policy)
audit = evaluated.audit
```

Eligibility criteria compose with explicit `&`, `|`, and `~`; score terms
combine through explicit arithmetic; `ranking_by` lists ordered expressions
and directions after the primary score. YAML overrides do none of this
implicitly. Numeric references can be compared to form predicates, and predicates can
participate in explicit score arithmetic. A bare numeric reference is not an
eligibility predicate; a bare predicate is not a named numeric score term.
An input column called `binding` remains a column; only `criterion("binding")`
means the named criterion. Unknown/cyclic references, duplicate names, wrong
roles and non-boolean eligibility outputs raise. Optional `applies_to` is an
eligibility expression; false applicability is recorded as `not_applicable`,
and its reference remains unknown rather than becoming an observed failure.

Audit rows retain occurrence identity, raw row links, criterion name and role,
value, status and reason:

| Status | Meaning |
| --- | --- |
| `pass` / `fail` | An applicable predicate evaluated true / false |
| `unknown` | Missing column/input, missing/ambiguous/conflicting model evidence, unknown applicability, or an out-of-domain calculation |
| `value` | An observed numeric term, including zero |
| `not_applicable` | The applicability predicate was observed false |
| `not_evaluated` | Unreferenced, or scoring skipped after a pre-filter |

Named predicates use three-valued boolean logic. `unknown="exclude"`,
`"include"`, or `"error"` controls the decision on an unknown final eligibility
result, while audit values remain unknown. Direct-expression policies without
criteria keep historical DSL comparison behavior. A vector expression with an
ambiguous/conflicting model is unknown for that evaluation context; separate
source contexts prevent one source's model selection from being imposed on
another. Unexpected programming/DSL errors still raise.

Schema 2 persists criteria, their complete expansions, references, ordered
terms, fill/gate settings and unknown handling in the policy digest. Original
schema-1 definitions retain their original serialization and digest. Definitions
are self-contained; a mutable registry cannot change a saved recipe. Export
`evaluated.evidence`, not a filtered ranking, to reproduce rejected alternatives.
The Vaxrank consumer test carries these records through native dataset save/load
and verifies frozen scores, selected windows, peptide constructs and mRNA
constructs, plus a changed criterion that changes the selected window.

## Identities and source evidence

Expand Down
118 changes: 118 additions & 0 deletions scripts/check_vaxrank_candidates.py
Original file line number Diff line number Diff line change
Expand Up @@ -100,3 +100,121 @@ def test_candidate_features_reach_vaxrank_scoring_and_vaccine_construction(antig
vaccines.sort(key=lambda v: v.target_epitope_score, reverse=True)
assert vaccines[0].antigen.amino_acids == ("SIINFEKL" if policy == "original" else "GILGFVFTL")
assert len(vaccines) == 2 # two pipelines did not become four vaccine targets


def test_occurrence_policy_matches_vaxrank_sources_windows_constructs_and_native_reload(tmp_path):
"""Real Vaxrank consumer operations over retained heterogeneous evidence."""
from dataclasses import replace
import json
from pathlib import Path
from varcode import Variant
from topiary import (
CachedPredictor, ProteinFragment, TopiaryPredictor, TopiaryResult,
SelectionCriterion, evaluate_selection_policy, replay_selection_policy,
read_lens, read_pvacseq,
)
from vaxrank.epitope_dataset import EpitopeDataset
from vaxrank.epitope_dsl import (
default_score_expr, genotype_lookup, score_predictions,
prediction_group_columns, resolve_default_methods, resolve_default_versions,
epitopes_for_ranking,
)
from vaxrank.core_logic import vaccine_peptides_from_epitopes
from vaxrank.mutant_protein_fragment import MutantProteinFragment
from vaxrank.peptide import assemble_peptide_constructs, PeptideConstructConfig
from vaxrank.mrna import assemble_mrna_constructs, RNAConstructConfig
from vaxrank.native_serialization import to_native_json

root = Path(__file__).resolve().parents[1] / "tests/data"
translation = json.loads((root / "osteosarc_shared/translation-v1.json").read_text())
prediction = json.loads((root / "osteosarc_shared/prediction-contract-v1.json").read_text())
fragment = ProteinFragment.from_dict(translation["fragment"])
# This versioned fixture is reconstructed from VCF/BAM in the full suite.
# Its numeric binding predictions are deliberately synthetic and cached.
direct = TopiaryPredictor(models=CachedPredictor(pd.DataFrame(prediction["rows"])),
only_novel_epitopes=True).predict_from_fragments([fragment])
# Carry the original translated context alongside the measurement table.
direct["source_sequence"] = fragment.sequence
combined = combine_sources({
"direct": TopiaryResult(direct),
"normalized": source(),
"lens": read_lens(root / "lens/sample_v1_4.tsv"),
"pvacseq": read_pvacseq(root / "pvacseq/mhc_i_all_epitopes.tsv"),
}, sample_name="fixture-patient")
dataset = EpitopeDataset.from_topiary(combined)
frame = dataset.scoring_frame()
keys = prediction_group_columns(frame)
cfg = EpitopeConfig()
policy = SelectionPolicy("frozen-vaxrank", default_score_expr(cfg), score_fill=0.,
min_score=cfg.min_epitope_score, duplicates="best")
contexts, expected = {}, []
for label, part in frame.groupby("source_label", sort=False):
contexts[label] = dict(
default_methods=resolve_default_methods(cfg, part),
default_versions=resolve_default_versions(cfg, part),
alleles=genotype_lookup(dataset.epitopes, keys))
expected.append(score_predictions(dataset.epitopes, cfg, topiary_df=part))
evaluation = evaluate_selection_policy(TopiaryResult(frame, metadata=combined.metadata), policy,
group_keys=keys, source_contexts=contexts)
original_scores = pd.concat(expected).sort_index()
actual_scores = evaluation.occurrences.set_index(keys).score.sort_index()
pd.testing.assert_series_equal(actual_scores, original_scores, check_names=False, check_exact=True)

def transfer_scores(result):
# Match score_predictions: pre-filtered groups have no score entry.
retained = result.occurrences.loc[result.occurrences.filter_retained]
records = retained.set_index(keys).score.to_dict()
return [replace(epitope, per_allele_scores={
allele: score for (*identity, allele), score in records.items()
if tuple(identity) == epitope.prediction_group_key}) for epitope in dataset.epitopes]

source_ids = set(frame.loc[frame.source_label.eq("direct"), "prediction_id"])
variant = Variant("12", 5494381, "A", "G")
native_fragment = MutantProteinFragment(
variant=variant, gene_name=fragment.gene, amino_acids=fragment.sequence,
mutant_amino_acid_start_offset=10, mutant_amino_acid_end_offset=11,
supporting_reference_transcripts=[], n_overlapping_reads=14, n_alt_reads=9,
n_ref_reads=5, n_alt_reads_supporting_protein_sequence=9)

def construct(epitopes):
selected = [epitope for epitope in epitopes_for_ranking(epitopes, cfg)
if epitope.prediction_id in source_ids]
windows = vaccine_peptides_from_epitopes(variant, native_fragment, selected, vaccine_peptide_length=11)
assert windows
ranked = [(variant, windows)]
peptide = assemble_peptide_constructs(ranked, PeptideConstructConfig(
min_antigen_length_aa=9, max_antigen_length_aa=11))
mrna = assemble_mrna_constructs(ranked, RNAConstructConfig(
signal_peptide="", include_mitd=False, poly_a_length=0, optimize_linkers=False,
min_antigen_length_aa=9, max_antigen_length_aa=11))
assert peptide and mrna
return [window.amino_acids for window in windows], peptide, to_native_json(mrna)

old_epitopes = []
for label, part in frame.groupby("source_label", sort=False):
ids = set(part.prediction_id)
old_epitopes.extend(attach_per_allele_scores(
[e for e in dataset.epitopes if e.prediction_id in ids], cfg, topiary_df=part))
original = construct(old_epitopes)
assert construct(transfer_scores(evaluation)) == original

criterion = SelectionCriterion("late_occurrence", "peptide_offset >= 10", "eligibility")
changed_policy = replace(policy, name="late-context", criteria=(criterion,),
filter_by='criterion("late_occurrence")')
changed = evaluate_selection_policy(TopiaryResult(frame, metadata=combined.metadata), changed_policy,
group_keys=keys, source_contexts=contexts)
changed_epitopes = transfer_scores(changed)
assert construct(changed_epitopes)[0] != original[0]
excluded = changed.audit.loc[changed.audit.status.eq("fail")]
assert not excluded.empty
assert excluded.criterion.eq("late_occurrence").all()
assert excluded.reason.eq("predicate_false").all()
saved = EpitopeDataset(result=changed.evidence, epitopes=tuple(changed_epitopes), config=cfg,
selection={"audit": changed.audit.to_json(orient="records")})
path = tmp_path / "native-vaxrank.tsv"
saved.save(path)
restored = EpitopeDataset.load(path)
replay = replay_selection_policy(restored.result)
assert restored.selection == saved.selection
pd.testing.assert_frame_equal(replay.occurrences, changed.occurrences, check_exact=True)
assert construct(restored.epitopes) == construct(changed_epitopes)
Loading
Loading