Skip to content

WT fragment join loses haplotype scores when the deconvolved presenter changes #453

Description

@iskandr

The existing fragment WT join matches the deconvolved allele as if it defined a haplotype score. When MT and WT select different presenters within the same configured genotype, the WT score is silently missing.

A deterministic presentation fixture with genotype HLA-A02:01/HLA-B07:02 gives MT SIINFEKL score 0.2 with presenter A02:01 and WT GIINFEKL score 0.8 with presenter B07:02. predict_from_fragments(..., predict_wt=True) produces wt_value=NaN, although both scores were returned. The explicit occurrence path in #451 correctly retains 0.8.

The join in _maybe_predict_wt_peptides includes allele unconditionally. Comparator matching must use declared MHC scope: the allele for per-allele scores, the full configured genotype for haplotype scores, and no allele key for allele-free scores. Keep the primary presenter label unchanged and do not invent WT context.

Reproducer on 5.90.0:

import pandas as pd
from topiary import TopiaryPredictor, ProteinFragment
class Haplotype:
    alleles = ['HLA-A*02:01', 'HLA-B*07:02']
    default_peptide_lengths = [8]
    uses_flanking_sequences = True
    def kind_support(self):
        return {'pMHC_presentation': {'mhc_dependence': 'haplotype', 'mhc_class': 'I'}}
    def predict_dataframe(self, peptides, n_flanks=None, c_flanks=None):
        return pd.DataFrame([dict(peptide=p, allele=self.alleles[0 if p.startswith('S') else 1],
            kind='pMHC_presentation', score=.2 if p.startswith('S') else .8,
            value=.2 if p.startswith('S') else .8, percentile_rank=None,
            n_flank=(n_flanks or [''] * len(peptides))[i], c_flank=(c_flanks or [''] * len(peptides))[i],
            prediction_method_name='haplotype_fixture', predictor_version='1') for i, p in enumerate(peptides)])
    def predict_proteins_dataframe(self, inputs):
        return pd.concat([self.predict_dataframe([sequence], n_flanks=[''], c_flanks=['']).assign(
            source_sequence_name=name, peptide_offset=0) for name, sequence in inputs.items()], ignore_index=True)
predictor = TopiaryPredictor(models=Haplotype(), predict_wt=True, only_novel_epitopes=False)
fragment = predictor.predict_from_fragments([ProteinFragment(
    fragment_id='source', sequence='SIINFEKL', reference_sequence='GIINFEKL')])
print(fragment[['peptide', 'allele', 'allele_set', 'value', 'wt_value']])

Acceptance: run fragment and exact-occurrence comparator paths through the same per-allele, haplotype and allele-free fixture battery. A changed haplotype presenter must retain the WT score, while distinct genotypes never match. Related #367 / PR #451.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions