Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 23 additions & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,12 @@ jobs:
fail-fast: true
matrix:
python-version: ["3.9", "3.10", "3.11"]
polars-version: ["latest"]
include:
- python-version: "3.11"
polars-version: "1.31.0"
- python-version: "3.11"
polars-version: "1.32.0"

steps:
- name: Checkout repository
Expand All @@ -30,12 +36,28 @@ jobs:
python -m pip install ruff pytest pytest-cov coveralls
pip install -r requirements.txt
pip install .
- name: Install compatibility Polars version
if: matrix.polars-version != 'latest'
run: python -m pip install 'polars==${{ matrix.polars-version }}'
- name: Run default linting script
run: |
./lint.sh
- name: Run unit tests
run: |
./test.sh
./test.sh -W error::DeprecationWarning
- name: Publish coverage to Coveralls
if: matrix.python-version == '3.11'
uses: coverallsapp/github-action@v2.2.3
with:
flag-name: polars-${{ matrix.polars-version }}
parallel: true

coverage:
needs: build
if: ${{ always() }}
runs-on: ubuntu-latest
steps:
- name: Finish combined coverage report
uses: coverallsapp/github-action@v2.2.3
with:
parallel-finished: true
2 changes: 1 addition & 1 deletion gtfparse/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@
)
from .write_gtf import write_gtf

__version__ = "2.8.0"
__version__ = "2.8.1"

__all__ = [
"GENCODE_BIOTYPE_ALIASES",
Expand Down
7 changes: 4 additions & 3 deletions gtfparse/read_gtf.py
Original file line number Diff line number Diff line change
Expand Up @@ -110,9 +110,10 @@
def parse_with_polars_lazy(
filepath_or_buffer, split_attributes=True, features=None, fix_quotes_columns=["attribute"]
):
# use a global string cache so that all strings get intern'd into
# a single numbering system
polars.enable_string_cache()
# Categories shares categorical mappings on modern Polars. Older releases
# need the global string cache for categoricals from separate reads to match.
if not hasattr(polars, "Categories"):
polars.enable_string_cache()
kwargs = {
"has_header": False,
"separator": "\t",
Expand Down
59 changes: 59 additions & 0 deletions tasks/todo.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# Issue #77: Polars string-cache compatibility

## Specification

`parse_with_polars_lazy` currently enables Polars' process-wide string cache on
every read. Modern Polars maintains categorical mappings through `Categories`;
the old setter is a no-op deprecated in 1.41.0. Skip the setter when the
`Categories` API exists, retaining shared categorical mappings on older supported
Polars releases. Keep parsing results, categorical dtypes, and interoperability
between separately parsed frames unchanged. Do not suppress warnings.

Add behavioral regression coverage for repeated pandas/Polars reads and joins of
categoricals from different files. Run CI with deprecations as errors, covering
modern Polars and representative legacy releases. Bump 2.8.0 to 2.8.1.

Separately verify Ensembl and Ensembl Genomes validators and matching/stale
`If-Range` responses using small ranges; report the implications for datacache
#80 without changing datacache in this PR.

## Plan

- [x] Read issue, repository guidance, dependency bounds, and current tests.
- [x] Check plan before implementation: use API capability detection rather than
assuming all Polars 1.x versions share the modern categorical behavior.
- [x] Reproduce #77 with deprecations as errors.
- [x] Implement compatibility guard and regression tests; bump patch version.
- [x] Add CI coverage and verify modern/legacy Polars behavior.
- [x] Run `./lint.sh` and `./test.sh`; review the diff and record results.
- [ ] Aggregate coverage from modern and legacy Polars jobs; confirm Coveralls.
- [ ] Open and merge a PR after checks pass.
- [ ] Run `./deploy.sh` from clean master and verify PyPI publication.
- [ ] Review relevant open issues and identify the next dependency/urgency group.
- [x] Report live Ensembl header and `If-Range` evidence for datacache #80.

## Review

The original code fails with Polars 1.44.2 when DeprecationWarning is an error.
With the capability guard, all 89 tests pass on 1.44.2 (95% coverage), and
`./lint.sh` passes. Polars 1.31.0 passes 88 tests (96% coverage) with the one
modern-only test skipped. The legacy run used pandas 3.0.6, also confirming that
the categorical conversion remains functional there.
Polars 1.32.0, the first release with Categories, also passes all 89 tests
(95% coverage). All three full-suite runs treated deprecations as errors.

Replan after the first CI run: all five matrix jobs passed, but Coveralls only
received the latest-Polars run and reported the intentionally unused legacy
setter as uncovered. Combine the three Python 3.11 coverage reports using
Coveralls parallel uploads and a finalization job. Preserve the coverage gate.
Tracked as https://github.com/openvax/gtfparse/issues/79 and fixed in PR #78.

On 2026-09-28, the Ensembl release-115 human toplevel DNA endpoint returned
strong ETag `"37d18890-63953fb016ef5"` and size 936478864. Ensembl Genomes
plants release-58 Arabidopsis DNA returned `"22c606f-607d8c61f60a5"` and size
36462703. Both returned `206` with a 16-byte body and the correct Content-Range
for `Range: bytes=1024-1039` plus a matching If-Range. A deliberately stale
If-Range returned `200`, so the client must restart instead of appending. The
probe capped response size to avoid downloading the full genomes. This supports
the strong-ETag path in openvax/datacache#80; an ETag is a representation
validator, not an independently verified SHA-256 digest.
51 changes: 51 additions & 0 deletions tests/test_polars_compatibility.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
from io import StringIO

import polars
import pytest

from gtfparse import read_gtf
from gtfparse.read_gtf import parse_with_polars_lazy

GTF_TEXT = (
'1\tensembl\tgene\t10\t20\t.\t+\t.\tgene_id "g1";\n'
'2\thavana\texon\t30\t40\t.\t-\t.\tgene_id "g2";\n'
)

pytestmark = pytest.mark.filterwarnings("error::DeprecationWarning")


@pytest.mark.parametrize("result_type", ["pandas", "polars"])
def test_repeated_reads_preserve_categoricals(result_type):
for _ in range(2):
df = read_gtf(StringIO(GTF_TEXT), result_type=result_type, features={"exon"})
assert list(df["seqname"]) == ["2"]
assert list(df["feature"]) == ["exon"]
assert list(df["gene_id"]) == ["g2"]
if result_type == "polars":
assert df["seqname"].dtype == polars.Categorical
else:
assert df["seqname"].dtype.name == "category"


def test_categoricals_from_separate_lazy_reads_can_join():
left = parse_with_polars_lazy(StringIO(GTF_TEXT))
reversed_text = "\n".join(reversed(GTF_TEXT.splitlines())) + "\n"
right = parse_with_polars_lazy(StringIO(reversed_text))
keys = ["seqname", "source", "feature", "strand"]
joined = left.join(right, on=keys).sort("start").collect()

assert joined.height == 2
assert joined["start"].to_list() == [10, 30]
assert joined["start_right"].to_list() == [10, 30]
for key in keys:
assert joined[key].dtype == polars.Categorical


@pytest.mark.skipif(not hasattr(polars, "Categories"), reason="Requires modern Polars")
def test_modern_reads_do_not_toggle_global_string_cache(monkeypatch):
def fail_on_cache_toggle():
pytest.fail("Reading a GTF must not toggle the global string cache on modern Polars")

monkeypatch.setattr(polars, "enable_string_cache", fail_on_cache_toggle)
monkeypatch.setattr(polars, "disable_string_cache", fail_on_cache_toggle)
assert read_gtf(StringIO(GTF_TEXT)).height == 2
Loading