Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,16 @@ between published tags; older pre-5.0 notes are retained below.

For current interfaces, see the [consumer guide](docs/consumer-guide.md).

## 5.88.0

- Import native Exacto peptide variants, primary structures and translations,
retaining original records, DNA/RNA links, alternative ORFs and unknown
specificity. Validate recovered sequence/novelty geometry against a pinned
public corpus; preserve transcript-scoped read membership when supplied (#365).
- Combine native observations with reported LENS/pVACseq candidates without
prediction, preserving evidence through CSV/TSV and explicit rescoring.
Add optional `read_exacto_fragments` for separate new-window scanning.

## 5.87.0

- Link normalized events, nucleotide ORF hypotheses, full protein products and
Expand Down
10 changes: 10 additions & 0 deletions docs/api.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# API Reference

## Native Exacto input

`read_exacto(path, *, sample_name, schema=None, primary_structures=None,
reference_name=None, tag=None, transcript_read_support=None, library_id=None,
read_set_id=None)` reads native sequence/evidence tables without prediction.
`read_exacto_fragments(path, **kwargs)` uses the same validation and returns
optional fragments for explicit new-window scanning. `EXACTO_SCHEMAS` and
`EXACTO_SCHEMA_COMMIT` describe the tested producer contract. See [Reading
Exacto](exacto.md) for supported files, evidence scope and composition examples.

## Fragment identity and RNA outcomes

A fragment record is identified by `(sample_name, fragment_id)`: the ID
Expand Down
5 changes: 2 additions & 3 deletions docs/combined-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,8 @@ ranked.to_csv("ranked-candidates.tsv", sep="\t", index=False)
```

Inputs may be `TopiaryResult` objects or normalized DataFrames. Existing readers
normalize LENS and pVACseq. Native Exacto parsing is separate work in
[#365](https://github.com/openvax/topiary/issues/365); an already normalized
ORF/RNA DataFrame is supported now. Aggregated reports contribute only the
normalize LENS and pVACseq; [native Exacto tables](exacto.md) use `read_exacto`.
Already normalized ORF/RNA DataFrames are also supported. Aggregated reports contribute only the
candidates they actually report.

A single input needs no model name, version or per-row source metadata. For
Expand Down
124 changes: 124 additions & 0 deletions docs/exacto.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# Reading native Exacto output

`read_exacto` imports native Exacto TSV files into `TopiaryResult` without
running Exacto, reconstructing reads or executing a prediction model. Supply
`sample_name`: these files do not identify the patient/sample themselves.

```python
from topiary import read_exacto

exacto = read_exacto(
"sample_exacto_peptide_variants.tsv",
sample_name="patient-01",
primary_structures="sample_exacto_primary_structures.tsv",
tag="exacto-run-01",
)
```

## Supported native schemas

The tested producer is [Exacto 0.4.6, commit
307c086](https://github.com/pirl-unc/exacto/tree/307c08670d5e706734bddf393bcebc84db497f9f).
Topiary identifies required column sets rather than assuming that a filename
or the producer's package version establishes a schema. `EXACTO_SCHEMAS`
exports these column sets; `schema=` can select one explicitly. Input may be a
path, gzip-compressed path or open text stream. Extra columns and original
records remain in `result.extra["exacto"]`, including companion tables.

| Schema | Native output | Normalized observations |
| --- | --- | --- |
| `peptide_variants` | `call-peptide-vars` TSV | Reported peptide sequences and parent IDs; optional primary structures supply exact occurrences and context |
| `primary_structures` | `translate-structs` TSV | One translated ORF hypothesis per native peptide ID, with producer-annotated changed intervals |
| `translations` | `translate-seqs` TSV/TSV.gz | Every reported RNA/ORF translation, including partial alternatives |

`transcript_read_support` is a companion table, supplied with the argument of
that name. It is accepted with primary structures, or peptide variants plus
primary structures. Provide `library_id` and `read_set_id`; stream inputs also
require a `tag` identifying the Exacto run. Local transcript model IDs are
scoped by this tag, or by the input path. Use the same run tag when importing
different files from one run. Distinct read names supply a transcript-level
read count with explicit membership. Missing models remain unknown. These
counts are not variant support or ORF-specific abundance; repeated peptide
rows do not create independent RNA observations.

The native corpus in `tests/data/exacto` pins source paths and SHA-256 digests:
228 peptide records, six primary structures and 77 translations. Unsupported
headers, primary record types, sequence-bearing event rows, invalid codons or
inconsistent peptide/parent geometry raise `ValueError`. This reader does not
import arbitrary Exacto tables or reconstruct genomic events from local IDs.

## What the evidence means

Native DNA and RNA call IDs, transcript model IDs, reference transcript IDs and
all source records are retained. Local numeric IDs do not establish agreement
with another caller's genomic events. Translation ORF bounds are converted from
zero-based inclusive coordinates to zero-based half-open coordinates; original
values remain in metadata. Translation IDs include both RNA and peptide IDs,
since the producer reuses peptide IDs across RNA records.

With primary structures, peptide coordinates and flanks are checked against
the reconstructed translation. Primary indices count base and event records,
so they are not treated as amino-acid coordinates. Changed amino-acid intervals
come from Exacto's explicit `amino_acid_change` labels. Without primary
structures, context and novelty intervals remain unknown and the reported
peptide is still available.

`sequence_source="predicted_from_observed_rna"` distinguishes inferred protein
sequence from a protein-expression measurement. A complete `protein_sequence`
requires a start codon and terminal stop. Partial translations remain in
`protein_hypothesis_sequence`. Native files do not establish matched-normal
specificity or comparator protein sequences: these stay unknown. A changed
sequence is not automatically tumor-specific.

## Combine with reported candidates

```python
from topiary import (
combine_sources, evidence_views, melt_pvacseq_algorithms, rank_candidates,
read_lens, read_pvacseq, reconcile_evidence,
)

combined = reconcile_evidence(combine_sources({
"exacto": exacto,
"lens": read_lens("lens.tsv"),
"pvacseq": melt_pvacseq_algorithms(read_pvacseq("all_epitopes.tsv")),
}, sample_name="patient-01"))
ranked = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True, duplicates="best",
)
by_source = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True, duplicates="best",
strata=["source_label", "candidate_mhc_class"],
)
combined.to_tsv("combined-evidence.tsv")
links = evidence_views(combined)["candidate_occurrences"]
```

Exacto's native files contain no HLA assignments or pMHC predictions. Therefore
its rows retain null `allele`, `value` and `candidate_id`; ORF-only rows also
have no peptide. The ranking universe comes from already reported pMHCs.
`candidate_occurrences` links same-sample, identical peptide occurrences to
those queries, with `reported_candidate=False` for a sequence match. This does
not transfer expression, scores or eligibility to a different observation.
All source records and native metadata survive long/wide CSV and TSV.

Choose compatible original measurements explicitly. [Combined-source
ranking](combined-sources.md) explains conflicts, missing scores and additive
`rescore_candidates` calls; rescoring alone does not invent HLA combinations.

## Optional new windows

`read_exacto_fragments(path, **kwargs)` accepts the same inputs and produces
`ProteinFragment` records with the recovered sequence, novelty intervals and
native provenance. It calls no model. Empty native tables return an empty list.

Pass these fragments to `TopiaryPredictor.predict_from_fragments` to explicitly
scan new peptide windows. This changes the candidate universe and is separate
from rescoring reported candidates. Select lengths/alleles deliberately and
use `overlaps_target` to filter recovered changed intervals. Missing intervals
remain unknown. Fragment conversion preserves each native observation, so
several reported peptides from one parent may carry the same parent sequence.

The general reader uses no Varcode fusion reconstruction: that adapter handles
selected DNA-SV-linked linear RNA models, whereas these native translated
sequences can be imported directly without genomic anchors or raw reads.
8 changes: 4 additions & 4 deletions docs/fragments.md
Original file line number Diff line number Diff line change
Expand Up @@ -467,7 +467,7 @@ Prefix is sanitized to `[A-Za-z0-9._:-]`; runs of other characters collapse to `

- **Indel / frameshift `wt_peptide`** — `wt_peptide` is only populated when the baseline is the same length as the mutant sequence (substitution-compatible). Length-changing edits yield `None`. Whether to define a remapped comparator remains an open design question ([#411](https://github.com/openvax/topiary/issues/411)).
- **Binding-aware cross-reactivity axes** (`self_mimic_*`, `self_strongest_nearby_*`, `self_nearest_candidates`) — the sequence-level `self_nearest_*` columns are computed by `SelfProteome` / `predict_self_nearest`; the axes that also require MHC prediction on each candidate are not implemented ([#412](https://github.com/openvax/topiary/issues/412)).
- **Dedicated file loaders** (`read_isovar_fragments`,
`read_exacto_fragments`) — separate work. Existing in-memory results already
compose through `fragment_from_isovar_result` and
`fragments_from_dataframe(read_pvacseq(...).df)`.
- **Dedicated Isovar fragment-file loader** (`read_isovar_fragments`) — separate
work. Existing in-memory results compose through `fragment_from_isovar_result`.
[Native Exacto](exacto.md) now has `read_exacto_fragments`; pVACseq composes
through `fragments_from_dataframe(read_pvacseq(...).df)`.
1 change: 1 addition & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ nav:
- Protein Fragments: fragments.md
- Cached Predictions: cached.md
- Reading pVACseq: pvacseq.md
- Reading Exacto: exacto.md
- Combining Sources: combined-sources.md
- Consumer Guide: consumer-guide.md
- Cache Miss Reports: cache-miss-reporting.md
Expand Down
Loading
Loading