Skip to content

Release 0.17.0: JVM chunk context, parse retry, shared-state readers - #24

Merged
sblattj merged 49 commits into
mainfrom
release/0.17.0
Sep 25, 2026
Merged

sblattj merged 49 commits into
mainfrom
release/0.17.0

Conversation

@sblattj

@sblattj sblattj commented Sep 25, 2026

Copy link
Copy Markdown
Owner

Release 0.17.0: Java and Kotlin chunk context (#20), a bounded retry for unusable model replies (#21), and code that reads the state a change writes (#22, parts 2 and 3).

The user-facing account is CHANGELOG.md [0.17.0]; the maintainer's map is HANDOFF.md.

What ships

Behaviour change at the defaults

Verification

Issues

Closes #20. Closes #21.

#22 stays open: part 1 (a bounded follow-up call for more context) is deferred to a later release. It needs a decision on the project's single-shot rule.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459

Adds the `llm_parse_retries` key (int, default 1, >= 0, no ceiling)
across all four config surfaces: config._DEFAULTS/_INT_KEYS/_RANGES,
the config.py module docstring, .env.example, and docs/env-vars.md
(with its derived key/name/group counts). Nothing reads the key yet;
a later task wires it into the reply-retry loop in reviewer.py.

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add heuristics.toggle_pinned_off_findings(files), a pure, deterministic
check modelled on release_shape_findings. When a PR adds a toggle whose
default is a literal true and, in the same PR, a suite-wide test-setup
file (conftest.py, setupTests.*, jest.setup.*, vitest.setup.*) that sets
a matching name to "false", "0" or "off", it returns one warning at the
toggle's line with confidence 1.0 and the deterministic-check suffix. The
green suite never runs the shipped default, which is why issue #22's
fixture passed its own tests while breaking two unchanged readers.

Both sides must be added lines, so a pre-existing toggle or pin never
fires. A pin names a toggle when its lowercase name tokens (split on
_ . -) end with the toggle's: ASSISTANT_PROGRESS_NOTES names
progress_notes. The body quotes the toggle call so line alignment keeps
the anchor on the toggle line, and it carries no hedge wording.

The check is not wired into the review yet; tests cover the #22 fixture
lines, every supported toggle and pin form, and the negatives (no pin,
a true pin, a pin in a plain test, either side on a context line,
unrelated names, a False default).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
review_chunk and review_systemic take a keyword-only parse_retries (library
default 0) and thread it into _invoke_and_parse, which now runs one shared
retry budget over the same request: an empty reply, an unparseable one, a
non-object, and, at parse_retries >= 1 only, an object without a findings
list are sent again while fewer than N retries have run. An empty reply
keeps its single retry at N=0, so a unit makes at most 1+max(N,1) calls.
Budget stops (finish_reason length/max_tokens) and raised calls are never
retried.

A reply still unusable when the budget is spent fails with its own error,
worded as before; at N >= 1 an object without findings fails with
"worker review JSON has no findings list". At N=0 the behaviour, log lines,
meta keys and trace files are unchanged, and {} stays a clean review.

Every call's usage folds through _fold_retry_usage. When a retry ran at
N >= 1, meta and meta.json gain parse_retries and first_error after the
eight base keys, each retry logs one "parse retry <k> of <N>" warning, and
each discarded reply is traced as <label>.attempt<K>.response.json beside
the four files of the used attempt. _write_trace_files keeps its
positional signature and gains a keyword-only attempts argument; a later
write of the same label removes attempt files it did not make.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ency context (#20)

Add prxref/jvm_gradle.py, a stdlib leaf module (it imports nothing from
prxref) that reduces a build.gradle or build.gradle.kts, plus its version
catalog, to declared coordinates for the JVM dependency-version block.

- String notation "g:a:v" and "g:a" in either quote style, anywhere outside a
  comment, so every configuration call and calls wrapped across lines count.
  Comments are stripped by a string-aware scan, so a commented-out
  declaration never becomes a pin and a glob such as '**/*.java' never opens
  one. A classifier or @ext is dropped; an interpolated version stays literal.
- platform(), enforcedPlatform() and mavenBom are BOM owners, as a string or
  as a catalog accessor (platform(libs.spring.boot.bom)).
- group = "g" / group "g" on its own line sets the own group; exclude
  group: and exclude(group = ...) never do.
- gradle/libs.versions.toml (or libs.versions.toml beside the build file) via
  tomllib: [versions], and [libraries] in string, module and group/name form,
  with a string, rich, version.ref or { ref = } version. Every catalog
  library counts as declared. The catalog is probed only when the build file
  mentions libs. outside a comment: its own directory, then the root; the
  first non-empty read wins.
- All reads go through an injected read(path); broken TOML, an unparseable
  file or a raising read contributes nothing and never raises. Map notation
  is out of scope.

Tests: tests/test_jvm_gradle.py (39).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add src/prxref/jvm_maven.py, a stdlib-only leaf module that resolves one
pom.xml and its in-repository parents into groupId:artifactId coordinates
with a version, or with the owner that manages the version.

- Safety: text holding <!DOCTYPE or <!ENTITY (any case) or larger than
  MAX_POM_BYTES (512 KiB of UTF-8, equal to chunk_context.MAX_FILE_BYTES)
  is refused before xml.etree ever parses it. Namespaces are stripped. A
  parse error, a non-<project> root or a failing read contributes nothing,
  and resolve_pom never raises.
- Parents: <relativePath> (default ../pom.xml, a directory gets pom.xml
  appended) resolved against the child's directory; an empty
  <relativePath/>, a path outside the repository, an unreadable or
  mismatched artifactId, MAX_POM_PARENTS = 5 or a revisit stops the walk.
- Properties: ${name} from <properties> merged child-over-parent plus
  project.version, project.groupId and project.parent.version, in at most
  MAX_PROPERTY_PASSES = 5 passes; unresolved placeholders stay literal and
  a pass that would grow a value past MAX_INTERPOLATED_CHARS is abandoned.
- Dependencies and dependencyManagement are both read. A managed
  type=pom scope=import entry is a BOM owner. A versionless dependency
  takes the nearest managed version, else the first owner in Maven's own
  precedence (external parent, then imported BOMs nearest-first); with
  neither it renders no line. Lines are g:a@v or
  g:a@(managed by bg:ba@bv).

All I/O goes through an injected read(path) -> str | None. Covered by
tests/test_jvm_maven.py (50 tests, counting readers for the walk and a
parse spy for the refusals).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Brings the 0.16.0 trace-event docs fix (507a405) into the 0.17 line.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…rts (#20)

Adds src/prxref/jvm_lang.py, a stdlib-only leaf module that imports nothing
from prxref, so chunk_context can use it without the repo_context import
cycle. It holds:

- jvm_language(path): "java" for .java, "kotlin" for .kt and .kts.
- Java definition regexes in the order a caller tries them: types (the
  repo_context type pattern, copied byte for byte), methods, fields and
  constants, then enum constants. A statement-keyword lookahead at the type
  position keeps return/new/throw/else/assert lines out of the method and
  field regexes.
- Kotlin regexes for class (data, sealed, enum, value, annotation), interface,
  fun interface, object, typealias, fun (with extension receivers) and
  val/var, with modifier and annotation prefixes.
- The Java keyword and JDK name sets (equal to repo_context's) and a Kotlin
  keyword set.
- parse_import/parse_imports returning JvmImport(name, alias, static,
  wildcard) for any package, including static, wildcard and Kotlin `as`.
- annotation_start: the lookback to at most 2 directly preceding annotation
  lines, where a definition entry then starts.

tests/test_jvm_lang.py pins the equality with repo_context, the leaf-import
property, every regex against the hard negatives, and Spring and Kotlin
snippets end to end.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add prxref.repo_readers, a pure module that finds the unchanged code
reading state a PR's added lines write to. shared_state_keys takes the
keys from subscript stores (recv[k] = v) and write calls
(recv.append/add/insert/save/put/push/store/write/extend/update/setdefault),
resolving a bare alias such as data = run.root().data to its attribute.
A key the added lines bind to a NEW value (a literal, a constructor or
any call) is skipped as a local; a plain name or attribute chain on the
right side (self.table = table) binds existing state and keeps the key,
which the #22 fixture needs for both of its keys.

reader_matches finds .key, ["key"] and key.<method>( reads that are not
stores, and excerpts each from the enclosing definition down to the read
in at most 8 lines, eliding the middle of a longer span. reader_candidates
ranks same-language non-diff listing files by shared directory depth, and
reader_entries walks them under MAX_READER_SCAN = 24 read attempts and
MAX_READER_ENTRIES = 6 entries, both read at call time, returning reader /
shared-state ContextEntry values. Wiring into repo_unit and registering
the kind and reason in repo_context are separate changes.

On the #22 fixture the progress.py chunk yields the keys table and data,
and entries for history.py model_history (table.recent) and
state_store.py load (f["data"]).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
judge_case gains a keyword-only parse_retries (N, library default 0).
When parse_judge_response raises JudgeParseError, for any reason (empty,
not JSON, not an object, no grades list, or a grades row that is not an
object, has no human_id or has an unknown grade), the same request is
sent again while fewer than N retries have run, so a case makes at most
1 + N calls. A raised call and a cache hit are never retried. When every
attempt is rejected, the case fails with the last attempt's reason,
worded as before.

Every attempt's usage is folded into the outcome's unit through
reviewer._fold_retry_usage before costs.run_cost, so tokens, time and
cost cover all of them, and a retry that raises keeps the billed earlier
attempts. llm_calls is 1 plus the retries made, and JudgeOutcome gains
parse_retries (last field, default 0). Only the graded reply is cached.
With a trace on at N >= 1, each discarded reply lands in
judge.attempt<K>.response.json, and once a retry ran the meta adds
parse_retries and first_error. At N=0 nothing changes from 0.16.0.

tests/test_judge_parse_retry.py covers each of the parser's seven
rejection kinds, retries running out, N=0, raised calls, usage and cost,
cache hits and the trace files.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add prxref.jvm_deps.dependency_lines(path, added, read), which turns the
imports on a .java, .kt or .kts file's added lines into sorted, deduplicated
g:a@v and g:a@(managed by ...) lines, the same shape
chunk_context.dependency_versions returns for other languages. It imports
only the stdlib and the jvm_lang, jvm_maven and jvm_gradle leaf modules, so
chunk_context can import it without a cycle. Nothing calls it yet.

- Imports rooted at java, jdk, sun or kotlin are skipped (javax is kept);
  with nothing left, or on a non-JVM path, read is never called.
- The walk climbs from the file's directory to the repository root, trying
  pom.xml, build.gradle.kts, then build.gradle at each level. The first
  non-empty read wins and ends the walk even if it is unparseable, so a
  broken manifest contributes nothing. Every read goes through one per-call
  cache, so no path is read twice (a parent relativePath pointing back at a
  probed miss is not re-read), and a raising or non-text read is a miss.
- Imports under the project's own group (pom groupId, falling back to the
  parent's, or Gradle group) are skipped by segment prefix.
- Matching: a dependency is a candidate when its groupId has at least 2
  segments and prefixes the import, or shares at least 3 leading segments
  with it. The artifactId tokens (split on -._, minus groupId segments)
  that equal an import segment score the candidates; the best score wins
  and ties list every tied candidate, so a databind import picks
  jackson-databind over jackson-core and jackson-annotations.
- Maven renders through MavenDependency.line(), keeping jvm_maven's owner
  precedence. A versionless Gradle dependency is owned by the platform whose
  group shares the most leading segments with it, first declared on a tie,
  rendered through the same MavenDependency/MavenCoordinate code.
- Never raises; any failure yields [].

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
At PRXREF_REPO_CONTEXT=repo, build_unit_context gains a fourth source
after the resolver: repo_readers.reader_entries over the chunk's files,
with the same guarded reader (so excluded paths are never read and the
readers get only the chunk reads the resolver left), every PR diff path
as diff_paths (a diff file is never a candidate), and the exclude
predicate. It runs only with a reader and a file listing, never at diff
or off, and reads repo_readers.MAX_READER_ENTRIES and MAX_READER_SCAN at
call time, so zeroing MAX_READER_ENTRIES switches readers off.

The entries have kind "reader" and reason "shared-state", both appended
last to KINDS and REASONS, so the budget cuts them first. They render in
their own last prompt block, "### Code elsewhere that reads state this
chunk writes": UnitContext gains a trailing reader_lines field, _admit
routes reader entries there (the omitted marker follows the first entry
left out, readers included), render_context_blocks takes reader_lines
last, and the orchestrator's _context_blocks passes it. With no reader
lines every prompt renders byte-identically to before.

docs/env-vars.md's PRXREF_REPO_CONTEXT row names the readers. The two
tests that pin REASONS and KINDS are re-pinned, and the admission-rank
test now includes a shared-state entry. New tests cover the levels, the
listing, exclusion, diff-file candidates, the shared chunk cap, the
runtime switch, the budget, record rows, rendering and the #22 shape.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…cord its total (#21)

orchestrate_review gains a keyword-only llm_parse_retries (library default
0, which behaves as 0.16.0) and hands it to every review_chunk call, the
timeout retry's included, and to review_systemic as parse_retries=. The
CLI passes cfg["llm_parse_retries"], so a CLI run retries once by default.

_invoke_chunk, _run_worker and _run_sweep each rebuilt the unit result
from named keys and dropped the reviewer's parse_retries and first_error.
All three now carry both when the meta has them (_retry_meta) and neither
otherwise; a raised call and a crashed worker carry neither.

The run record gains parse_retries: None when N is not an int of 1 or
more, otherwise the sum over every chunk and the sweep (0 when nothing
retried, including exits before any review unit). It is stamped after the
sweep and before the total-failure exit, and every exit carries it through
_run_record. --format json emits it after repo_context, and the README key
list documents it. The timeout retry still replaces the first run's
counts, the known gap in docs/llm.md.

Test doubles for review_chunk/review_systemic with explicit keyword lists
accept parse_retries=0, and the pinned run-record and JSON key lists add
the new key. New tests in tests/test_parse_retry_orchestrator.py.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…pes-only (#20)

repo_context's Java type regex, keyword set and JDK names are now the same
objects as prxref.jvm_lang's JAVA_TYPE_RE, JAVA_KEYWORDS and JDK_NAMES, so
there is one implementation; repo_contracts keeps importing _JAVA_KEYWORDS
unchanged. The java/javax-only Java import regex stays, because
jvm_lang.parse_import does not reproduce it exactly (it also accepts a
semicolon-less import and rejects two imports on one line).

Kotlin is answered here, before chunk_context is asked:
- language_of maps .kt and .kts, in any case, to "kotlin";
- definition_regexes("kotlin") is (jvm_lang.KOTLIN_TYPE_RE,), types only,
  like Java: class of any flavour, interface, named object, typealias;
- referenced_names for Kotlin uses the Kotlin keywords, drops JDK names and
  every segment (and the as-alias) of a kotlin.*, java.* or javax.* import
  on the added lines, and keeps only type-like names. kotlinx.* is kept.

The module docstring now says chunk_context handles Java and Kotlin itself
for the same-file context, and that this module keeps a types-only view for
repository context with the shared facts in prxref.jvm_lang.

A Kotlin file is therefore a same-language reader candidate for a Kotlin
chunk. New tests in tests/test_repo_context_kotlin.py.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
eval run allowlists llm_parse_retries into run.json's config (seventeen
keys, placed after llm_max_tokens with the other LLM settings). eval
score passes the loaded PRXREF_LLM_PARSE_RETRIES to judge_case as
parse_retries, so the judge re-sends a rejected reply as the review does,
and score.json's judge block totals the retries as parse_retries, beside
llm_calls. score.md's judge cost line names the retries only when there
were any, so a scoring run without one renders byte-identically.

The review's own retries are not totalled in score.json: a record's
parse_retries is null at PRXREF_LLM_PARSE_RETRIES=0 on a healthy review,
which the metrics' total/missing shape would misreport as missing, and
each case's record.json already carries its count.

docs/evals.md documents the key, the judge's parse retries (a reply cut
short at the token limit included), its trace files and the judge block
key, and corrects the judge-cost wording for a retry that raised after a
billed reply. judge_cost's docstring says the same. Two judge-block key
order pins in the score and end-to-end tests gain the new key.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The alias test used `Queue`, which the JDK name set already drops, so it
stayed green with the alias drop removed. It now aliases to a name no set
covers, checks the name is kept without the import, and checks that an
alias of another organization's import is kept.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Calls heuristics.toggle_pinned_off_findings(files) at both sites where
release_shape_findings(files) already runs: the main path, where it is
spliced in at the chunk/sweep boundary alongside release_shape (moving
sweep_start with it, so it stays a CHUNK-side finding to
apply_sweep_dedup), and the summary-only (no-chunk) path, for symmetry,
threaded through _summary_only_run's new toggle_findings parameter.

Unlike release_shape, the toggle finding sits on a real line (> 0), so
the long comment above the main-path splice site is updated: the
reworded tier of apply_sweep_dedup (on with dedup_similarity) now can
compare the toggle finding against a chunk worker's own restatement on
the same line, keeping only the higher-ranked one of the two, instead
of never comparing it the way a line-0 finding is exempted. The
orchestrator module docstring and quality.py's module docstring (the
seventeenth-check sentence, now also naming the toggle check as the
eighteenth) are extended to match.

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Drives every #21 acceptance bullet through cli._run_review (and cli.main
for the exit code) on a two-chunk fixture patch read by the local-diff
forge, with the real openai-compat client, reviewer and orchestrator
against a local MockOpenAIServer. The server answers each review unit
(two chunks and the sweep) from its own script of canned replies and
counts that unit's calls, so counts do not depend on worker order, and
an unscripted extra call is an HTTP 500 that shows up as a wrong count.

Covered with PRXREF_LLM_PARSE_RETRIES unset (default 1), 0 and 2:
malformed then valid (2 calls, findings kept, run record and JSON 1,
meta parse_retries and first_error, the discarded reply in
chunk0.attempt1.response.json, the same request resent); malformed
twice (2 calls, the last of two different errors wins, exit 0 on a
partial and a total failure, exit 1 only with PRXREF_FAIL_ON=error);
{} twice (the no-findings error); a list, {} and an object without
findings retried; the issue's own stray-quote reply; a length stop
(1 call, the budget message, 8 meta keys) against the same bytes with
an honest stop (retried); 0 (a {} reply is clean, garbage costs 1 call,
record null, no attempt file, 8 meta keys, the empty-reply retry kept);
2 (3 calls); the sweep; and tokens and reported cost summed over every
attempt.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…endencies (#20)

Wire the jvm_lang and jvm_deps leaf modules into chunk_context, so a chunk
that touches .java, .kt or .kts files gets the same definitions and
dependencies blocks the js and python chunks already get.

- _language maps .java to "java" and .kt/.kts to "kotlin" through
  jvm_lang.jvm_language; _definition_regexes and _keywords return the
  jvm_lang tables for those two languages. Every other language is
  unchanged.
- referenced_definitions starts a Java or Kotlin entry at the annotation
  lines directly above the definition (at most MAX_ANNOTATION_LINES, clamped
  so the definition line always fits in max_lines_per_entry), reports the
  first annotation's line number, and counts the annotation lines toward the
  line cap. js and python decorators are not included, and non-JVM output is
  byte-identical.
- dependency_versions hands a JVM file to jvm_deps.dependency_lines before
  the manifest lookup and merges its g:a@v lines into the same sorted,
  deduplicated result. build.gradle.kts and settings.gradle.kts are skipped
  with no read.
- repo_crosschunk._SAME_FILE_LANGUAGES gains java and kotlin, so with a
  reader diff_definitions leaves a chunk's own JVM files to chunk context
  instead of repeating their definitions as diff-file entries.

Tests: a new tests/test_chunk_context_jvm.py, the Java/Kotlin rows of
test_repo_context_crosschunk.py re-pinned, and the chunk-context Java pin in
test_repo_context_defs.py flipped.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Adds tests/fixtures/issue22: the issue's make-fixture.sh, the PR as
pr.patch (format-patch main..feature/progress-notes under a fixed
placeholder identity and date, byte-reproducible), the PR head tree as
repo/, and the issue's description as desc.md.

tests/test_issue_22_acceptance.py drives the production local path,
cli._run_review with --diff-file, --repo-dir and --description-file,
against the real openai-compat client and a local scripted server that
answers every prompt with no findings, and asserts on the --format json
payload and the recorded prompts:

- at repo, the progress.py chunk's reader block lists
  assistant/history.py:9, assistant/state_store.py:15 and
  tests/test_engine.py:17, in that order, and nothing else in the prompt
  names either source file; the run record carries the same three
  shared-state reader rows;
- at diff and off, no prompt carries the reader block;
- at every level there is exactly one finding, the pinned-off toggle at
  assistant/progress.py:34, undropped and with the deterministic suffix;
  deleting the conftest pin, or setting it to "true", removes it;
- repo_readers.MAX_READER_ENTRIES = 0 removes exactly the reader block
  and nothing else from the prompts, and no reader row from the record;
- repo/ matches both the generator's head tree and the patch's
  post-image, and the patch carries only dev@example.com.

repo/ holds the fixture project's own tests and code that the project's
lint rules reject, so tests/fixtures/issue22/conftest.py keeps pytest
out of it (collect_ignore) and tests/fixtures/issue22/ruff.toml keeps
ruff out of it, while inheriting the project's settings.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ference (#20)

Leaving every js, python, Java and Kotlin file of a chunk out of
diff_definitions whenever a reader is given avoided showing a file's own
definitions twice, but it also lost a definition one file of a chunk
references and another file of the same chunk defines outside its hunks:
chunk_context.referenced_definitions only looks inside the referencing file.

diff_definitions now searches such a file only for the wanted names that do
not occur as an identifier (chunk_context._IDENT_RE) on that file's own
added lines. Every name referenced_definitions can show for the file is one
of those identifiers, so nothing is shown twice; a hit keeps reason
diff-file. When no name is left the file is skipped without a read, so a
one-file chunk still makes no own-file read here.

- repo_crosschunk module and diff_definitions docstrings describe the rule
  (the module's "knows no Java" line is gone).
- New TestTheChunksOwnFiles pins, for java, kotlin, python and js: the
  cross-file entry, no entry and no read for a file's own reference, zero
  reads for a one-file chunk, and a name both files reference shown once.
- test_repo_context_unit.py read lists and one read-cap entry count follow
  the removed own-file reads of one-file chunks.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…reads (#20)

Since #20 the chunk context reads Java and Kotlin files and probes the
directories above them for pom.xml, build.gradle.kts and build.gradle at
every repository-context level, off included, so the #17 fixture's forge
reads gained 23 JVM paths. The prompts are unchanged and the 0.15.0 golden
stays byte for byte.

- The off comparison now checks the prompts against the golden as before
  and the reads with every JVM path left out (_is_jvm: a .java, .kt or .kts
  basename, or pom.xml, build.gradle, build.gradle.kts, libs.versions.toml).
- A new test checks the golden itself holds no JVM path, so the filter
  cannot hide one of its reads, and another pins the off runs' JVM reads
  literally and in order (21 manifest probes, then ConnectorService and
  TransportConfig) for both design points and all three off variants.
- REPO_READS and the floor test's read list gain the same 23 JVM reads;
  the repo-mode control also shows the filtered comparison fails there.
- The two read-cap starvation tests count the chunk context's one read of
  each of the chunk's 21 own Java files, which bypasses the repository
  reader's cap; the route-contract test's last-read check becomes "SPEC is
  the only non-JVM read". Every other assertion is unchanged.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add the 0.17.0 CHANGELOG section (date left as YYYY-MM-DD for the
release): Java and Kotlin chunk context and dependency versions, Kotlin
in repository context, PRXREF_LLM_PARSE_RETRIES with the run record,
trace attempt files and the judge retry, the shared-state reader block,
and the pinned-off toggle check; the default-behaviour change for
replies without a findings list; the same-chunk definition fix; and the
known limitations, including the deferred #22 part 1.

docs/llm.md gains a Parse Retries section, retry costing under Which
calls count, JVM definitions and dependencies, the as-built own-file and
cross-file rule, and the reader block. docs/quality.md documents the
toggle check, docs/prompt-templates.md says nothing checks the reply
format's findings key, and the README repository-context table, the
config.py PRXREF_REPO_CONTEXT row and the orchestrate_review docstring
name the reader block and its first-attempt-only rule.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…0.16.0

Drives the real orchestrate_review over a new fixture monorepo: a Maven
multi-module build (parent properties plus a BOM import, a versionless
module dependency, a malformed module pom), a Gradle build with a
version catalog reached through version.ref and a garbage subproject
build.gradle, and a Python control, seven files, one per chunk.

Asserts the literal dependency lines (property-resolved, BOM-managed,
catalog), the skipped java.* and own-group imports, the annotated
definitions entries from their first annotation line, the Kotlin fun
and class, no keyword-named definition, the broken builds (no line, a
completed run in the record, the trace and the log, and a control that
reaches the enclosing build without them), one rendering per own-file
definition at diff and repo with a live cross-chunk control, and the
40-entry cap.

The regression golden was captured from the released prxref 0.16.0
installed from PyPI, never from the tree: the Python control and the
sweep prompts are byte-identical, and every JVM chunk prompt differs
only by the added dependency and definition blocks.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Bump pyproject.toml, prxref.__version__ and the uv.lock prxref stanza to
0.17.0. Date the [0.17.0] CHANGELOG section 2026-09-25, point [Unreleased]
at v0.17.0...HEAD and add the [0.17.0] link.

Correct the PRXREF_REPO_CONTEXT row of docs/env-vars.md: off adds no
repository context, but chunk context's Java and Kotlin definitions and
dependency versions do not depend on it, so off is no longer byte-identical
to 0.15.0; and diff finds Kotlin type declarations as well as Java ones.

Rewrite HANDOFF.md for 0.17.0: the module map for Java and Kotlin chunk
context (#20), the parse retry (#21) and the shared-state readers and
pinned-off toggle check (#22, parts 2 and 3), the lessons, the config key
counts (68 keys, 69 accepted names), the release shape, the live check,
and what is still open, including #22 part 1.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Severity consistency ran over every finding, the release-shape and
pinned-off toggle findings included, so a model finding in the same file
that shared a rare code token with one (or any finding with its normalized
title) could raise it. The live check on #22's fixture saw the toggle
finding, which has no model behind it, posted as an error in two of three
runs and as a warning in the third.

heuristics.is_deterministic marks a finding whose body ends with the
checks' "(deterministic check, no model)" suffix, and
apply_severity_consistency now leaves such a finding out the way it leaves
out a dropped one: it is never raised, never raises another finding, and
its text does not count toward a token's rarity. No other pass changes.

Tests: a new class in tests/test_quality.py covers the token rule, the
title rule, the finding as a would-be source, a transitive link and the
rarity count, with a control whose unmarked body is raised; a new
tests/test_deterministic_severity.py runs #22's fixture through the local
review path with a scripted model error on assistant/progress.py and
checks the toggle stays a warning, with a control that patches the helper
off and sees it raised to error.

Docs: the CHANGELOG Fixed entry; docs/quality.md's consistency row names
both grouping rules and the exclusion, and the deterministic-checks section
says which passes can still touch these findings (finding grouping can
still raise one); a HANDOFF lesson and the verified count. Also rewrites
the .env.example PRXREF_REPO_CONTEXT paragraph to match docs/env-vars.md,
and corrects the HANDOFF registration row: only REASONS' order ranks
repository-context entries.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The orchestrator comment and the heuristics module docstring still said a
deterministic finding flows through the quality passes exactly like a model
finding. Severity consistency now leaves them out, so both say so and name
heuristics.is_deterministic.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
@github-actions

Copy link
Copy Markdown

prxref automated review: Approved

PR: Release 0.17.0: JVM chunk context, parse retry, shared-state readers · files reviewed: 103

🟥 0 error · 🟧 3 warning · 🔍 0 spec · ⬜ 1 outofscope

  • 🟧 src/prxref/chunk_context.py:375 — Unreachable keywords branch after unconditional return
  • 🟧 src/prxref/eval_judge.py:296 — llm_calls undercounts when a retry's invoke raises
  • 🟧 src/prxref/jvm_deps.py:139 — Zero-score tie emits every candidate dependency line
  • ⬜ tests/test_parse_retries_config.py:5 — Stale docstring: claims the config key is unread, but this PR wires it

Reviewed by prxref · model=z-ai/glm-5.3-flash · 206245 tok · 130.9s

⚠️ Partial review: 7 of 9 chunks were reviewed; 2 failed. Findings may be incomplete.

  • chunk of 15 files (tests/test_jvm_maven.py, HANDOFF.md, CHANGELOG.md, +12 more): LLMError: all models failed: z-ai/glm-5.3-flash: timeout (Timeout)
  • chunk of 13 files (tests/test_repo_unit_readers.py, src/prxref/jvm_lang.py, src/prxref/heuristics.py, +10 more): LLMError: all models failed: z-ai/glm-5.3-flash: timeout (Timeout)

if language == "python":
return (_PY_DEF_RE, _PY_ASSIGN_RE)
return ()
return jvm_lang.definition_regexes(language)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 🟧 [WARNING] Unreachable keywords branch after unconditional return (src/prxref/chunk_context.py:375)

The hunk at @@ -355,1 +375,3 @@ replaces return () with return jvm_lang.definition_regexes(language) followed by if language in _JVM_LANGUAGES: return jvm_lang.keywords(language). The return is unconditional, so the keywords branch on lines 376-377 can never execute; if keywords were meant to be returned for JVM languages (the new docstring text distinguishes definition regexes from keywords), that behaviour is lost, and at minimum the dead code should be removed or the return guarded.


Reviewed by prxref · model=z-ai/glm-5.3-flash

Comment thread src/prxref/eval_judge.py
_write_trace_files(trace_dir, JUDGE_TRACE_LABEL, system, user, result.text, {**trace_meta, "error": reason})
return JudgeOutcome(**live, grades=None, error=reason)
_write_trace_files(trace_dir, JUDGE_TRACE_LABEL, system, user, result.text, trace_meta)
if retries:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 🟧 [WARNING] llm_calls undercounts when a retry's invoke raises (src/prxref/eval_judge.py:296)

In the retry loop, an exception from client.invoke returns via _live(base, unit, elapsed, retries, price_table), which sets llm_calls = 1 + retries. But the failing invoke is itself an additional request beyond the 1 + retries parsed-retry calls, so a case that retries once and then raises reports 2 llm_calls while making 3 requests. The docstring's own contract ("llm_calls is the number of judge requests this call made ... otherwise 1 plus parse_retries") is inconsistent with the actual request count in this path. [containment boundary not stated]


Reviewed by prxref · model=z-ai/glm-5.3-flash

Comment thread src/prxref/jvm_deps.py


def _matches(segments: Sequence[str], declared: Sequence[_Declared]) -> list[_Declared]:
candidates = [dep for dep in declared if _is_candidate(dep.group, segments)]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 🟧 [WARNING] Zero-score tie emits every candidate dependency line (src/prxref/jvm_deps.py:139)

In _matches (src/prxref/jvm_deps.py:131-137), when no candidate's artifactId token matches any import segment, all candidates tie at score 0 and every tied candidate's line is emitted. An import like com.google.common.collect with two unrelated declared deps sharing a 3-segment group prefix would render both dependency lines. This is the documented tie rule, but a zero-score tie carries no matching evidence and may attribute the wrong artifact versions to a file.


Reviewed by prxref · model=z-ai/glm-5.3-flash


Nothing reads this key yet -- it is wired to the reply-retry loop in a later
task. This file only pins the config surface: default, valid range, the exit-2
path shared with the other int keys, and that the docs surfaces agree with the

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 ⬜ [OUTOFSCOPE] Stale docstring: claims the config key is unread, but this PR wires it (tests/test_parse_retries_config.py:5)

The module docstring says 'Nothing reads this key yet -- it is wired to the reply-retry loop in a later task', yet the same PR's src/prxref/orchestrator.py changes read llm_parse_retries and thread parse_retries into every chunk worker and the sweep. The comment misleads future readers about the feature's state.


Reviewed by prxref · model=z-ai/glm-5.3-flash

@sblattj
sblattj merged commit 43ec560 into main Sep 25, 2026
4 checks passed
@sblattj
sblattj deleted the release/0.17.0 branch September 25, 2026 05:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Retry a worker call once when its reply doesn't parse Chunk context for JVM files: Java/Kotlin definitions and Maven/Gradle dependency versions

1 participant