Reconcile ORF hypotheses and scoped RNA observations - #436
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Combined tables now distinguish biological events, nucleotide ORF hypotheses, full protein products and peptide occurrences. Public reconciliation and relational views preserve alternate translations, source disagreements and the supporting hypothesis of a ranked window. Allele-free occurrences link to existing same-sample peptide-HLA queries without transferring scores or inventing candidates.
Cross-caller ORF identity requires explicit reference, coding sequence and transcript path/bounds. Scoped RNA observations retain units and provenance; count union requires known evidence membership in the same sample/library/read-set/entity scope. No TPM summing, assembly liftover or implicit variant harmonization.
Also distinguishes known terminal flanks from missing flanks (#435), and fixes pandas 3 empty-source masks, genotype filtering and warning-class compatibility (#437).
Validation: composed long/wide CSV/TSV, ranking, rescoring and RNA-union workflows; pandas 2/3 checks; lint and all PR/master CI green, including published Vaxrank integration. Clean-master release gate: 4,820 passed, zero failures/skips, 35 non-pandas warnings, with one local test worker. Published wheel/sdist SHA-256 hashes verified against PyPI and tag verified at the release commit.
Released as Topiary 5.87.0. Closes #370, #435 and #437. Native Exacto ingestion is the separate #438.