Skip to content

fix(indexer): match a "Series N: Title" book on its own title - #2813

Merged
vavallee merged 7 commits into
mainfrom
fix/series-prefixed-title-relevance
Oct 2, 2026
Merged

vavallee merged 7 commits into
mainfrom
fix/series-prefixed-title-relevance

Conversation

@vavallee

@vavallee vavallee commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Summary

Reported on Discord: Feral Mage 4: The Mining Company Contract by Chase Kilgore never matched the MyAnonamouse release The Mining Company Contract by Chase Kilgore [ENG / EPUB] [VIP]. The search debug panel showed it dropped at the relevance stage.

The filter tries two readings of a title: the whole name, which needs "feral" and "mage", and `primaryTitle`, the part before the colon, which here is the series and position rather than the book. A release naming the book itself satisfies neither. Reproduced on main with the exact release.

Stacked on #2812 so this lands in the single filter implementation rather than being written into two copies. Retarget to `main` once #2812 merges, and merge #2812 without deleting its branch until this is retargeted, since deleting a base branch auto-closes the PR on top.

Change

`newznab.SeriesPositionTitle` returns the part after the colon when the part before it ends in a series position, and `""` otherwise:

Title Returns
Feral Mage 4: The Mining Company Contract The Mining Company Contract
Cradle Book 8: Wintersteel Wintersteel
Discworld #3: Equal Rites Equal Rites
12 Rules for Life: An Antidote to Chaos nothing, no position
1984: A Novel nothing, a bare number is a title
Rocky IV: The Novel nothing, roman numerals are ambiguous with words

The relevance filter tries it as a third reading, only after the full and primary readings fail, and with the same guards as the others:

  • At least two significant words, so "Wintersteel" can't turn every Will Wight release containing that word into a match.
  • The phrase gate, so "The Mining Company Breach of Contract" is still dropped.
  • The conflicting author check, so "The Mining Company Contract by Someone Else" is still dropped.
  • The volume number guard from fix(indexer): match "Volume" in a title against "Vol" in a release name #2921, the same as the full and primary readings (added when updating onto main).

The query tier gate in the newznab client gets the same reading, so a tier whose results carry only the book's title isn't taken for a canned feed and walked past.

Tests

  • `TestSeriesPositionTitle`: the shape rules above.
  • `TestFilterRelevant_SeriesPositionTitle`: the reported release and its variants kept; wrong author, different book, split phrase and the Pamela Palmer release from the same screenshot dropped. Runs through both filter wrappers.
  • `TestSeriesPositionTitle_ReachesProductionEntrypoints`: the reported release through `SearchBookWithDebug` and `SearchBookWithOutcomes` against a fake indexer.
  • `TestTitleHasRelevantResult_SeriesPositionTitle`: the tier gate.
  • Fail before: with the helper present but not wired in, every positive case, the entry point test and the gate test fail.
  • Mutations, each caught by a named test: dropping the two word floor, dropping the phrase gate on this reading, and dropping the conflicting author check on it.

Update onto main (2026-10-02)

Merged the updated #2812 branch, which itself now carries origin/main (#2863, #2921), and again after #2812 gained its zero padding fix (cc9d92b), which the series reading picks up through the shared matchers. One conflict, in the drop reason: it takes #2812's new wording and still counts a series reading keyword hit. Fragment renamed to changelog.d/2813-series-position-title-relevance.md with (#2813) in the lead. Fail before re-checked on the merged tree: with SeriesPositionTitle unwired in the filter and the tier gate, every positive case in TestFilterRelevant_SeriesPositionTitle (both wrappers), TestSeriesPositionTitle_ReachesProductionEntrypoints and TestTitleHasRelevantResult_SeriesPositionTitle fail; restored, all pass. Still stacked on #2812: retarget to main after #2812 merges, and do not delete #2812's branch before then.

Not in this PR

The reporter's second point, that there's no way to grab a release the filter dropped, is #2334. The debug panel already lists what was dropped and why; making those grabbable is a separate decision.

Checklist

  • Tests added or updated
  • Doc-update gate cleared (no user-facing setting or doc describes the readings)
  • Changelog fragment, crediting the reporter

Test plan

  • `go build ./... && go vet ./...`, plus `GOOS=windows` and `GOOS=darwin` builds
  • `go test ./internal/indexer/... ./internal/scheduler/ ./internal/api/`

🤖 Generated with Claude Code

https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9

@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.11765% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/indexer/newznab/words.go 86.66% 1 Missing and 1 partial ⚠️

📢 Thoughts on this report? Let us know!

… grab

filterRelevant and filterRelevantDebug were two copies of the relevance
filter. #2502 added its guards, conflictingTitleAuthor and the title
identity phrase gate, to filterRelevant only, and pinned them with tests
that call filterRelevant directly.

Production never runs filterRelevant. The scheduler calls
SearchBookWithOutcomes, which projects SearchBookWithDebug, and the
interactive search calls SearchBookWithDebug; both run
filterRelevantDebug. filterRelevant is reached only through the
scheduler's fallback for searchers that cannot report outcomes, which in
practice is test stubs. So every release #2502 was written to reject was
still kept, and automatic search could still grab it.

filterRelevantDetailed is now the one implementation and returns the
kept results plus a reason for each drop. filterRelevant and
filterRelevantDebug are wrappers over it. The drop record now
distinguishes a conflicting author and a split title phrase from a plain
keyword miss.

The #2502 tables run through both wrappers, and a new test drives
SearchBook, SearchBookWithDebug and SearchBookWithOutcomes against one
fake indexer serving the wrong releases beside the right ones. On main
both production entry points keep all four wrong releases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>
@vavallee
vavallee force-pushed the fix/relevance-filter-single-implementation branch from 90c12f5 to c74573f Compare September 27, 2026 02:01
A book stored as "Feral Mage 4: The Mining Company Contract" was
matched on two readings: the whole name, which needs "feral" and "mage",
and primaryTitle's part before the colon, which is the series and
position rather than the book. A release named "The Mining Company
Contract by Chase Kilgore" satisfies neither, so both automatic and
interactive search dropped it.

newznab.SeriesPositionTitle returns the part after the colon when the
part before it ends in a position: digits, optionally after "#" or
before a period, with at least one word ahead. An ordinary "Title:
Subtitle" has no such number and is untouched; roman numerals and a
bare leading number do not count.

The relevance filter tries that reading only after the full and primary
readings fail, requires at least two significant words, and applies the
same phrase gate and conflicting author check as the others. The query
tier gate in the newznab client gets the same reading, so a tier whose
results carry only the book's title is not taken for a canned feed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>
vavallee and others added 2 commits October 2, 2026 01:42
main gained #2863's hand copied author guard and #2921's volume number
guard in filterRelevantDebug. Both are folded into filterRelevantDetailed,
which both wrappers now share, so the production path also runs the
title identity gate it never had.

#2921 pinned the divergence in TestFilterRelevantVolumeMarkerSpelling
with separate plain and debug expectations. They collapse to one
expectation per case. Two cases change for production search: a same
spelling "Volume 16" release for a "Volume 17" title is now dropped
(intended), and a zero padded "Vol 07" release for a "Volume 7" title is
now dropped too (a known gap in the identity gate, marked in the test).

The drop reason for a keyword match that fails the identity gate or the
volume guard now reads "title words appear, but with other words or
numbers among them", which covers both. The changelog fragment is
renamed with the PR number and rewritten: the author half of the fix
reached production in v1.39.0 through #2863.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>
…ixed-title-relevance

Brings in the updated base, which itself now carries origin/main
(#2863's author guard, #2921's volume marker equivalence and volume
number guard).

The series position reading now also runs volumeNumberAgrees, so it
applies the same guards as the full and primary readings. The drop reason
takes the base branch's wording, which now covers a differing number as
well as inserted words. The changelog fragment is renamed with the PR
number and the lead gains (#2813).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>
@vavallee
vavallee marked this pull request as ready for review October 2, 2026 04:48
vavallee and others added 2 commits October 2, 2026 01:58
…matching

Putting the title identity gate on the production path (previous commits)
started dropping correct releases that pad a position: "Vol 07" for
"Volume 7", "Shield Hero 07" for "Shield Hero 7", "Book 07" for
"Book 7". The gate compared numbers as literal words.

keywordPattern now renders an all digit keyword as 0* followed by the
number without its padding, so the phrase, keyword and in order matchers
and the identity gate (ContainsPhrase) all accept "7", "07" and "007" as
one number, in either direction. The callers' word boundaries keep "7"
from matching "17" or "70". volumeNumberAgrees already ignored padding
through sameVolumeNumber.

conflictingTitleAuthor compares whole normalized strings, so it would
have read "Book 07 - Other Author" as a different title and skipped the
attribution. foldVolumeMarkers becomes foldTitleTokens and also strips
padding, so that release is still caught.

A digit and a spelled out number ("7" and "Seven") stay different; that
case is pinned as a known gap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>
…ixed-title-relevance

Brings in #2812's zero padding fix: a number matches the same number with
any zero padding in the phrase, keyword and identity checks, and the
attribution guard folds padding too. The series position reading goes
through the same matchers, so it gets the rule without changes here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fJcCVbNnKsmj2MAAMwWj9
Signed-off-by: vavallee <vavallee@protonmail.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Logic and tests look solid. A couple of small points:

Coverage gap in words.go: Codecov flags two lines. The len(head) < 2 branch of SeriesPositionTitle (line ~224) has no test case — something like {"A: Title", ""} (one word before the colon) would cover it. The pos == "" branch (when the last pre-colon word is just "#" or ".") is also missing — {"Series #: Title", ""} would cover it. Neither is a correctness problem since both paths correctly return "", but the gap means a future regression there goes undetected.

seriesOK skips the elided-apostrophe fallback: fullOK and primaryOK both run through identityOK(), which retries with the elided form when ContainsPhrase fails. The new seriesOK path (internal/indexer/searcher.go:991-993) uses ContainsPhrase directly, without that elided retry. In practice a series subtitle with an internal apostrophe is uncommon, and this is already a third-level fallback, so it isn't breaking — just noting the asymmetry.

Everything else is clean: conflictingTitleAuthor is correctly guarded on seriesTitle != "", the two-word floor and phrase gate are applied consistently, and the tier-gate in client.go gets the same reading. The "fail before / mutations" table in the PR body makes the test discipline easy to verify.

— 🤖 Bindery triage bot (automated). Reply to correct me; a human will see it.

@vavallee
vavallee changed the base branch from fix/relevance-filter-single-implementation to main October 2, 2026 05:13
Signed-off-by: vavallee <vavallee@protonmail.com>

# Conflicts:
#	internal/indexer/searcher.go
@vavallee
vavallee merged commit d7ac8dc into main Oct 2, 2026
31 of 34 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant