Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,16 @@ between published tags; older pre-5.0 notes are retained below.

For current interfaces, see the [consumer guide](docs/consumer-guide.md).

## 5.89.0

- Save and replay named `SelectionPolicy` definitions with complete DSL expressions,
method/version selections, ranking settings, and a content digest. Ranked
exports retain the policy, derivation and separate execution provenance.
Public mapping IO accepts composed consumer configuration and freezes defaults
without running predictors or changing scientific defaults (#442, #443).
- Preserve binary floating-point measurements and annotations through typed
CSV/TSV reload, keeping threshold decisions and scores reproducible (#441).

## 5.88.1

- Prefer the current predictor DataFrame API for cache fallbacks, retaining
Expand Down
11 changes: 11 additions & 0 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -431,6 +431,17 @@ always `Column("x") == "y"` (see below).

### Context options

For portable policy definitions, `SelectionPolicy` stores explicit filter/score
strings and model/version selections. `resolve_selection_policy(mapping)` fills
authoring defaults after a consumer has composed its overrides;
`SelectionPolicy.from_dict(mapping)` strictly reloads a complete definition.
`to_dict()` and `sha256` expose the effective settings and content identity.
`read_selection_policy(path)` / `write_selection_policy(policy, path)` provide
JSON IO; writes refuse existing files. `rank_with_policy(result, policy,
provenance=None)` wraps `rank_candidates` and returns a `TopiaryResult` carrying
the effective definition and separate provenance/execution metadata. See the
[policy workflow and consumer boundary](combined-sources.md#named-selection-policies).

`apply_filter`, `apply_sort` and `evaluate_scores` share five keyword-only
options, all forwarded to `EvalContext`:

Expand Down
112 changes: 112 additions & 0 deletions docs/combined-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,118 @@ Unknown predictor names and versions remain unknown. The ordinary expression
`affinity.value` continues to work; adding provenance does not require a
version-qualified expression for this simple case.

## Named selection policies

`SelectionPolicy` saves the exact candidate filter, score, model/version
selections, ranking direction, duplicate policy and strata under a stable name.
It has no implicit scientific recipe. Topiary does not currently ship an
official `openvax-v1` preset; that name can identify a reviewed definition you
freeze, and later definitions can use new names.

```python
from topiary import (
SelectionPolicy, read_selection_policy, write_selection_policy,
rank_with_policy, read_tsv,
)

# Illustrative settings, not a calibrated biological model.
policy = SelectionPolicy(
name="example-v1", score_by="1 / affinity.value",
filter_by="n_rna_alt >= 5", duplicates="best",
)
write_selection_policy(policy, "example-v1.json") # refuses to overwrite
combined.to_tsv("all-evidence.tsv")
ranked = rank_with_policy(combined, policy)
ranked.to_tsv("ranked.tsv")

replayed = rank_with_policy(
read_tsv("all-evidence.tsv"), read_selection_policy("example-v1.json"),
)
```

The immutable policy has `to_dict()` / `from_dict()` for complete, strict,
schema-versioned mappings and `sha256` for the canonical definition, including
its name. Every default is saved explicitly. Changing an expression or model
selection changes the digest even if the name remains `openvax-v1`. Runtime
versions and input facts live separately in
`ranked.extra['selection_policy']['execution']`; authoring history can be passed
as `provenance=` and is stored alongside the definition, outside its digest.
Record actual base digests and ordered config-file hashes there when available.

`rank_with_policy` delegates to `rank_candidates`: it filters first, then scores
surviving evidence and chooses one representative per candidate/stratum. Unknown
filter results exclude groups under the existing DSL's evidence-retention
rules. Missing scores remain unranked; zero scores remain zero. It performs no
predictor calls, automatic method/version preference selection, or score filling.
Explicit model selections follow the same DSL rules as direct calls. To choose
input-dependent defaults, call `resolve_default_methods` and
`resolve_default_versions` explicitly and include their results in the effective
policy before saving. Unstated historical versions remain unknown. Distinct
source-local selections need separate effective policies/invocations; do not
apply one source's choice globally.

CSV/TSV exports retain the full definition, digest, derivation and execution
record in long and wide form. Numeric measurements and annotations round-trip
exactly as binary floats. Preserve the **full evidence** as well: a filtered
ranking cannot restore excluded observations or discarded alternatives. Replay
also needs the recorded Topiary version when exact execution semantics matter.

Proteasome cleavage and whole-peptide half-life can already be named explicitly,
for example `peptide_view(proteasome_cleavage.score)` and
`peptide_view(serum_half_life.value)`. Saving a policy does not enable these
predictors or add processing weights to existing defaults. Typed extracellular
site-level evidence remains [#288](https://github.com/openvax/topiary/issues/288).

### Composed consumer configuration

Vaxrank owns YAML composition and window/construct configuration. Its repeated
`--config openvax-v1.yaml --config overrides.yaml` workflow merges mappings
left to right and replaces lists/scalars. The Topiary integration boundary is
the **already-composed policy subtree**:

```python
from topiary import resolve_selection_policy

# `merged` comes from the consumer's config loader, after all overrides.
policy = resolve_selection_policy(merged["selection_policy"])
saved_settings = {
"selection_policy": policy.to_dict(),
"selection_policy_sha256": policy.sha256,
"selection_policy_provenance": config_provenance,
"vaccine_settings": effective_vaccine_settings,
}
# Reload complete definitions strictly; never apply newer authoring defaults.
policy = SelectionPolicy.from_dict(saved_settings["selection_policy"])
```

`resolve_selection_policy` accepts a JSON- or YAML-decoded mapping with required
`name` and `score_by`. It fills omitted optional settings only after composition.
An omitted filter in an override inherits the base filter; an explicit YAML
`filter_by: null` removes it. Replacing a filter expression does not AND it with
the old expression. Model selections are mappings; version selections are lists
of `{kind, method, version}` records, so an override replaces that entire list.
Topiary does not merge files or parse unrelated consumer settings.

Vaxrank adoption of this subtree remains
[Vaxrank #497](https://github.com/openvax/vaxrank/issues/497); the example above
describes the consumer boundary, not a newly supported Vaxrank configuration
key. Its current occurrence-level scorer can consume `policy.score_by`,
`policy.filter_by`, and the selection mappings through the existing public
`EvalContext` / `apply_filter` APIs, retaining its explicit grouping and
per-occurrence `alleles` lookup. Vaxrank's post-score minimum gate, missing-score
fill, window rules, RNA weighting and construct settings must also be frozen in
its bundle. A cutoff inside a score expression must not become a destructive
pre-filter when capturing that baseline.

The representative ranking above is one view, not the complete construction
interface. Shared occurrence-level evaluation with retained alternatives is
tracked in [#444](https://github.com/openvax/topiary/issues/444). Optional named
criteria, references and pass/fail/unknown audit records are tracked in
[#445](https://github.com/openvax/topiary/issues/445); schema 1 currently accepts
direct DSL expressions and rejects unknown criteria/registry fields. These
layers will reuse the same DSL rather than require a second evaluator in
Vaxrank.

## Identities and source evidence

| Column | Meaning |
Expand Down
5 changes: 5 additions & 0 deletions docs/ranking.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,11 @@ df = apply_sort(df, [Presentation.score, Affinity.score])

`TopiaryPredictor(filter_by=..., sort_by=[...])` applies them automatically during prediction.

For reusable named configurations and exact save/replay, see
[selection policies](combined-sources.md#named-selection-policies). A policy
stores explicit expressions and settings without changing the DSL or enabling
new predictors.

## Group identity

Expressions evaluate per *group*, not per row: one peptide-allele group can span several rows (one per predictor and kind). `apply_filter` keeps or drops whole groups, `apply_sort` orders groups, and `evaluate_scores` gives every row its group's score.
Expand Down
22 changes: 17 additions & 5 deletions scripts/check_vaxrank_candidates.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,10 @@
import pandas as pd
from mhctools import Prediction

from topiary import combine_sources, evaluate_scores, join_annotations, parse, rank_candidates, rescore_candidates
from topiary import (
SelectionPolicy, combine_sources, evaluate_scores, join_annotations, parse,
rank_with_policy, read_selection_policy, write_selection_policy, rescore_candidates,
)
from tests.test_candidate_tables import Model, source
from vaxrank.candidate_epitope import candidate_epitopes_from_rows
from vaxrank.epitope_config import EpitopeConfig
Expand All @@ -21,7 +24,7 @@

@pytest.mark.parametrize("antigen_kind", ["mutation", "fusion", "splice", "CTA", "ERV", "viral"])
@pytest.mark.parametrize("policy", ["original", "rescored", "rna_overlay"])
def test_candidate_features_reach_vaxrank_scoring_and_vaccine_construction(antigen_kind, policy):
def test_candidate_features_reach_vaxrank_scoring_and_vaccine_construction(antigen_kind, policy, tmp_path):
proteins = ["MAAASIINFEKL", "MAAAGILGFVFTL"]
identity = dict(protein_sequence=proteins, event_id=["event-1", "event-2"])
combined = combine_sources({
Expand All @@ -39,8 +42,16 @@ def test_candidate_features_reach_vaxrank_scoring_and_vaccine_construction(antig
combined = join_annotations(combined, annotations, on=keys, prefix="rna",
provenance={"source": "rna_only", "unit": "TPM"})
expression = "rna_transcript_expression / affinity.value"
ranked = rank_candidates(combined, expression, duplicates="best")
selected_ids = set(ranked.source_observation_id)
saved_policy = SelectionPolicy(
name="example-" + policy, score_by=expression,
filter_by="n_rna_alt >= 5", duplicates="best",
)
policy_path = tmp_path / "selection.json"
write_selection_policy(saved_policy, policy_path)
restored_policy = read_selection_policy(policy_path)
assert restored_policy.sha256 == saved_policy.sha256
ranked = rank_with_policy(combined, restored_policy)
selected_ids = set(ranked.df.source_observation_id)
frame = combined.long_df[combined.long_df.source_observation_id.isin(selected_ids)].copy()
# Vaxrank's current public scoring interface names its provenance key
# prediction_id. Retain originals in the combined result; this is a
Expand All @@ -59,7 +70,8 @@ def test_candidate_features_reach_vaxrank_scoring_and_vaccine_construction(antig
overlaps_targetable=True, patient_alleles=[row.allele],
))
epitopes = candidate_epitopes_from_rows(rows)
cfg = EpitopeConfig(score_expr=expression, filter_expr="n_rna_alt >= 5", min_epitope_score=0.)
cfg = EpitopeConfig(score_expr=restored_policy.score_by,
filter_expr=restored_policy.filter_by, min_epitope_score=0.)
scored = attach_per_allele_scores(epitopes, cfg, topiary_df=frame)
expected = dict(zip(frame.source_observation_id, evaluate_scores(frame, parse(expression))))
vaccines = []
Expand Down
69 changes: 69 additions & 0 deletions tests/test_consumer_workflows.py
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,75 @@
from .test_twin_conformance import DSL_MEASUREMENT_TWINS


@pytest.mark.parametrize("form", ["long", "wide"])
def test_saved_profiles_replay_exact_scores_and_change_selection(tmp_path, form):
from topiary import (
SelectionPolicy, combine_sources, rank_with_policy,
read_selection_policy, write_selection_policy,
)
from .test_candidate_tables import source
from .test_twin_conformance import DELIMITED_IO_TWINS

boundary = 0.12345678901234567
annotations = [boundary, np.nextafter(boundary, np.inf)]
binding = source(values=(50., 100.), review_score=annotations)
processing = source(
values=(0.1, 0.9), score=[0.1, 0.9], review_score=annotations,
kind="proteasome_cleavage", prediction_method_name="cleavage_fixture", allele="",
)
combined = combine_sources({"input": TopiaryResult(pd.concat(
[binding.df, processing.df], ignore_index=True))}, sample_name="patient")
before = combined.df.copy(deep=True)
# Synthetic policies demonstrate behavior, not a calibrated scientific recipe.
profiles = [
SelectionPolicy("example-v1", "1 / affinity.value"),
SelectionPolicy("example-v2", "peptide_view(proteasome_cleavage.score)"),
SelectionPolicy("example-v3", "1 / affinity.value", filter_by=f"review_score <= {boundary!r}"),
]
paths = []
for profile in profiles:
path = tmp_path / (profile.name + ".json")
write_selection_policy(profile, path)
paths.append(path)
expected = [rank_with_policy(combined, profile) for profile in profiles]
assert [r.df.peptide.tolist() for r in expected] == [
["SIINFEKL", "GILGFVFTL"], ["GILGFVFTL", "SIINFEKL"], ["SIINFEKL"],
]
pd.testing.assert_frame_equal(combined.df, before)
assert "selection_policy" not in combined.extra
columns = ["candidate_id", "candidate_score", "candidate_rank", "ranking_status", "candidate_observations"]
for suffix, writer, method, reader in DELIMITED_IO_TWINS:
evidence_path = tmp_path / ("evidence." + suffix)
writer(combined if form == "long" else combined.to_wide(), evidence_path)
restored = reader(evidence_path)
for path, original in zip(paths, expected):
profile = read_selection_policy(path)
ranked = rank_with_policy(restored, profile)
pd.testing.assert_frame_equal(ranked.df[columns], original.df[columns], check_exact=True)
output_path = tmp_path / (profile.name + "." + suffix)
method(ranked if form == "long" else ranked.to_wide(), output_path)
saved = reader(output_path).to_long()
# Wide IO records the models actually present in this narrower
# view; the full policy and original-source provenance survive.
for key in ("selection_policy", "combined_sources"):
assert saved.extra[key] == ranked.extra[key]
assert saved.topiary_version == ranked.topiary_version
assert saved.filter_by_str == profile.filter_by
assert saved.sort_by_str == profile.score_by
record = saved.extra["selection_policy"]
assert record["definition"] == profile.to_dict()
assert record["sha256"] == profile.sha256
assert record["execution"]["topiary_version"] == ranked.topiary_version
# Loading the definition from the ranked export is sufficient to
# replay against the full evidence; filtered exports cannot recover it.
embedded = SelectionPolicy.from_dict(saved.extra["selection_policy"]["definition"])
reranked = rank_with_policy(restored, embedded)
pd.testing.assert_frame_equal(reranked.df[columns], original.df[columns], check_exact=True)
actual = saved.df.set_index("candidate_id").candidate_score.sort_index()
wanted = original.df.set_index("candidate_id").candidate_score.sort_index()
pd.testing.assert_series_equal(actual, wanted, check_exact=True)


def test_sv_interest_api_and_cli_retain_and_rank_the_same_nominations(tmp_path):
import json
from .test_sv_interest import catalogue, export
Expand Down
20 changes: 20 additions & 0 deletions tests/test_io.py
Original file line number Diff line number Diff line change
Expand Up @@ -658,6 +658,26 @@ def test_legacy_comment_metadata_keeps_builtins_and_custom_values(tmp_path):
# ---------------------------------------------------------------------------


@pytest.mark.parametrize("form", ["long", "wide"])
def test_binary_floats_round_trip_exactly_through_both_writer_doors(tmp_path, form):
from .test_candidate_tables import source

boundary = 0.12345678901234567
values = [boundary, np.nextafter(boundary, np.inf)]
result = source(values=values, review_score=values, score=values)
if form == "wide":
result = result.to_wide()
for suffix, writer, method, reader in DELIMITED_IO_TWINS:
for write in (writer, method):
path = tmp_path / ("precise." + suffix)
write(result, path)
restored = reader(path).to_long().df.set_index("peptide")
for column in ("value", "score", "review_score"):
actual = restored.loc[["SIINFEKL", "GILGFVFTL"], column].tolist()
assert [x.hex() for x in actual] == [x.hex() for x in values]
assert [x <= boundary for x in actual] == [True, False]


class TestReadWriteTSV:
def test_long_form_roundtrip(self, tmp_path):
df = _sample_long_df()
Expand Down
Loading
Loading