Release 0.17.0: JVM chunk context, parse retry, shared-state readers - #24
Conversation
Adds the `llm_parse_retries` key (int, default 1, >= 0, no ceiling) across all four config surfaces: config._DEFAULTS/_INT_KEYS/_RANGES, the config.py module docstring, .env.example, and docs/env-vars.md (with its derived key/name/group counts). Nothing reads the key yet; a later task wires it into the reply-retry loop in reviewer.py. 🤖 Authored with Claude Code — Claude Sonnet 5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add heuristics.toggle_pinned_off_findings(files), a pure, deterministic check modelled on release_shape_findings. When a PR adds a toggle whose default is a literal true and, in the same PR, a suite-wide test-setup file (conftest.py, setupTests.*, jest.setup.*, vitest.setup.*) that sets a matching name to "false", "0" or "off", it returns one warning at the toggle's line with confidence 1.0 and the deterministic-check suffix. The green suite never runs the shipped default, which is why issue #22's fixture passed its own tests while breaking two unchanged readers. Both sides must be added lines, so a pre-existing toggle or pin never fires. A pin names a toggle when its lowercase name tokens (split on _ . -) end with the toggle's: ASSISTANT_PROGRESS_NOTES names progress_notes. The body quotes the toggle call so line alignment keeps the anchor on the toggle line, and it carries no hedge wording. The check is not wired into the review yet; tests cover the #22 fixture lines, every supported toggle and pin form, and the negatives (no pin, a true pin, a pin in a plain test, either side on a context line, unrelated names, a False default). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
review_chunk and review_systemic take a keyword-only parse_retries (library
default 0) and thread it into _invoke_and_parse, which now runs one shared
retry budget over the same request: an empty reply, an unparseable one, a
non-object, and, at parse_retries >= 1 only, an object without a findings
list are sent again while fewer than N retries have run. An empty reply
keeps its single retry at N=0, so a unit makes at most 1+max(N,1) calls.
Budget stops (finish_reason length/max_tokens) and raised calls are never
retried.
A reply still unusable when the budget is spent fails with its own error,
worded as before; at N >= 1 an object without findings fails with
"worker review JSON has no findings list". At N=0 the behaviour, log lines,
meta keys and trace files are unchanged, and {} stays a clean review.
Every call's usage folds through _fold_retry_usage. When a retry ran at
N >= 1, meta and meta.json gain parse_retries and first_error after the
eight base keys, each retry logs one "parse retry <k> of <N>" warning, and
each discarded reply is traced as <label>.attempt<K>.response.json beside
the four files of the used attempt. _write_trace_files keeps its
positional signature and gains a keyword-only attempts argument; a later
write of the same label removes attempt files it did not make.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ency context (#20) Add prxref/jvm_gradle.py, a stdlib leaf module (it imports nothing from prxref) that reduces a build.gradle or build.gradle.kts, plus its version catalog, to declared coordinates for the JVM dependency-version block. - String notation "g:a:v" and "g:a" in either quote style, anywhere outside a comment, so every configuration call and calls wrapped across lines count. Comments are stripped by a string-aware scan, so a commented-out declaration never becomes a pin and a glob such as '**/*.java' never opens one. A classifier or @ext is dropped; an interpolated version stays literal. - platform(), enforcedPlatform() and mavenBom are BOM owners, as a string or as a catalog accessor (platform(libs.spring.boot.bom)). - group = "g" / group "g" on its own line sets the own group; exclude group: and exclude(group = ...) never do. - gradle/libs.versions.toml (or libs.versions.toml beside the build file) via tomllib: [versions], and [libraries] in string, module and group/name form, with a string, rich, version.ref or { ref = } version. Every catalog library counts as declared. The catalog is probed only when the build file mentions libs. outside a comment: its own directory, then the root; the first non-empty read wins. - All reads go through an injected read(path); broken TOML, an unparseable file or a raising read contributes nothing and never raises. Map notation is out of scope. Tests: tests/test_jvm_gradle.py (39). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add src/prxref/jvm_maven.py, a stdlib-only leaf module that resolves one
pom.xml and its in-repository parents into groupId:artifactId coordinates
with a version, or with the owner that manages the version.
- Safety: text holding <!DOCTYPE or <!ENTITY (any case) or larger than
MAX_POM_BYTES (512 KiB of UTF-8, equal to chunk_context.MAX_FILE_BYTES)
is refused before xml.etree ever parses it. Namespaces are stripped. A
parse error, a non-<project> root or a failing read contributes nothing,
and resolve_pom never raises.
- Parents: <relativePath> (default ../pom.xml, a directory gets pom.xml
appended) resolved against the child's directory; an empty
<relativePath/>, a path outside the repository, an unreadable or
mismatched artifactId, MAX_POM_PARENTS = 5 or a revisit stops the walk.
- Properties: ${name} from <properties> merged child-over-parent plus
project.version, project.groupId and project.parent.version, in at most
MAX_PROPERTY_PASSES = 5 passes; unresolved placeholders stay literal and
a pass that would grow a value past MAX_INTERPOLATED_CHARS is abandoned.
- Dependencies and dependencyManagement are both read. A managed
type=pom scope=import entry is a BOM owner. A versionless dependency
takes the nearest managed version, else the first owner in Maven's own
precedence (external parent, then imported BOMs nearest-first); with
neither it renders no line. Lines are g:a@v or
g:a@(managed by bg:ba@bv).
All I/O goes through an injected read(path) -> str | None. Covered by
tests/test_jvm_maven.py (50 tests, counting readers for the walk and a
parse spy for the refusals).
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Brings the 0.16.0 trace-event docs fix (507a405) into the 0.17 line. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…rts (#20) Adds src/prxref/jvm_lang.py, a stdlib-only leaf module that imports nothing from prxref, so chunk_context can use it without the repo_context import cycle. It holds: - jvm_language(path): "java" for .java, "kotlin" for .kt and .kts. - Java definition regexes in the order a caller tries them: types (the repo_context type pattern, copied byte for byte), methods, fields and constants, then enum constants. A statement-keyword lookahead at the type position keeps return/new/throw/else/assert lines out of the method and field regexes. - Kotlin regexes for class (data, sealed, enum, value, annotation), interface, fun interface, object, typealias, fun (with extension receivers) and val/var, with modifier and annotation prefixes. - The Java keyword and JDK name sets (equal to repo_context's) and a Kotlin keyword set. - parse_import/parse_imports returning JvmImport(name, alias, static, wildcard) for any package, including static, wildcard and Kotlin `as`. - annotation_start: the lookback to at most 2 directly preceding annotation lines, where a definition entry then starts. tests/test_jvm_lang.py pins the equality with repo_context, the leaf-import property, every regex against the hard negatives, and Spring and Kotlin snippets end to end. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add prxref.repo_readers, a pure module that finds the unchanged code reading state a PR's added lines write to. shared_state_keys takes the keys from subscript stores (recv[k] = v) and write calls (recv.append/add/insert/save/put/push/store/write/extend/update/setdefault), resolving a bare alias such as data = run.root().data to its attribute. A key the added lines bind to a NEW value (a literal, a constructor or any call) is skipped as a local; a plain name or attribute chain on the right side (self.table = table) binds existing state and keeps the key, which the #22 fixture needs for both of its keys. reader_matches finds .key, ["key"] and key.<method>( reads that are not stores, and excerpts each from the enclosing definition down to the read in at most 8 lines, eliding the middle of a longer span. reader_candidates ranks same-language non-diff listing files by shared directory depth, and reader_entries walks them under MAX_READER_SCAN = 24 read attempts and MAX_READER_ENTRIES = 6 entries, both read at call time, returning reader / shared-state ContextEntry values. Wiring into repo_unit and registering the kind and reason in repo_context are separate changes. On the #22 fixture the progress.py chunk yields the keys table and data, and entries for history.py model_history (table.recent) and state_store.py load (f["data"]). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
judge_case gains a keyword-only parse_retries (N, library default 0). When parse_judge_response raises JudgeParseError, for any reason (empty, not JSON, not an object, no grades list, or a grades row that is not an object, has no human_id or has an unknown grade), the same request is sent again while fewer than N retries have run, so a case makes at most 1 + N calls. A raised call and a cache hit are never retried. When every attempt is rejected, the case fails with the last attempt's reason, worded as before. Every attempt's usage is folded into the outcome's unit through reviewer._fold_retry_usage before costs.run_cost, so tokens, time and cost cover all of them, and a retry that raises keeps the billed earlier attempts. llm_calls is 1 plus the retries made, and JudgeOutcome gains parse_retries (last field, default 0). Only the graded reply is cached. With a trace on at N >= 1, each discarded reply lands in judge.attempt<K>.response.json, and once a retry ran the meta adds parse_retries and first_error. At N=0 nothing changes from 0.16.0. tests/test_judge_parse_retry.py covers each of the parser's seven rejection kinds, retries running out, N=0, raised calls, usage and cost, cache hits and the trace files. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add prxref.jvm_deps.dependency_lines(path, added, read), which turns the imports on a .java, .kt or .kts file's added lines into sorted, deduplicated g:a@v and g:a@(managed by ...) lines, the same shape chunk_context.dependency_versions returns for other languages. It imports only the stdlib and the jvm_lang, jvm_maven and jvm_gradle leaf modules, so chunk_context can import it without a cycle. Nothing calls it yet. - Imports rooted at java, jdk, sun or kotlin are skipped (javax is kept); with nothing left, or on a non-JVM path, read is never called. - The walk climbs from the file's directory to the repository root, trying pom.xml, build.gradle.kts, then build.gradle at each level. The first non-empty read wins and ends the walk even if it is unparseable, so a broken manifest contributes nothing. Every read goes through one per-call cache, so no path is read twice (a parent relativePath pointing back at a probed miss is not re-read), and a raising or non-text read is a miss. - Imports under the project's own group (pom groupId, falling back to the parent's, or Gradle group) are skipped by segment prefix. - Matching: a dependency is a candidate when its groupId has at least 2 segments and prefixes the import, or shares at least 3 leading segments with it. The artifactId tokens (split on -._, minus groupId segments) that equal an import segment score the candidates; the best score wins and ties list every tied candidate, so a databind import picks jackson-databind over jackson-core and jackson-annotations. - Maven renders through MavenDependency.line(), keeping jvm_maven's owner precedence. A versionless Gradle dependency is owned by the platform whose group shares the most leading segments with it, first declared on a tie, rendered through the same MavenDependency/MavenCoordinate code. - Never raises; any failure yields []. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
At PRXREF_REPO_CONTEXT=repo, build_unit_context gains a fourth source after the resolver: repo_readers.reader_entries over the chunk's files, with the same guarded reader (so excluded paths are never read and the readers get only the chunk reads the resolver left), every PR diff path as diff_paths (a diff file is never a candidate), and the exclude predicate. It runs only with a reader and a file listing, never at diff or off, and reads repo_readers.MAX_READER_ENTRIES and MAX_READER_SCAN at call time, so zeroing MAX_READER_ENTRIES switches readers off. The entries have kind "reader" and reason "shared-state", both appended last to KINDS and REASONS, so the budget cuts them first. They render in their own last prompt block, "### Code elsewhere that reads state this chunk writes": UnitContext gains a trailing reader_lines field, _admit routes reader entries there (the omitted marker follows the first entry left out, readers included), render_context_blocks takes reader_lines last, and the orchestrator's _context_blocks passes it. With no reader lines every prompt renders byte-identically to before. docs/env-vars.md's PRXREF_REPO_CONTEXT row names the readers. The two tests that pin REASONS and KINDS are re-pinned, and the admission-rank test now includes a shared-state entry. New tests cover the levels, the listing, exclusion, diff-file candidates, the shared chunk cap, the runtime switch, the budget, record rows, rendering and the #22 shape. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…cord its total (#21) orchestrate_review gains a keyword-only llm_parse_retries (library default 0, which behaves as 0.16.0) and hands it to every review_chunk call, the timeout retry's included, and to review_systemic as parse_retries=. The CLI passes cfg["llm_parse_retries"], so a CLI run retries once by default. _invoke_chunk, _run_worker and _run_sweep each rebuilt the unit result from named keys and dropped the reviewer's parse_retries and first_error. All three now carry both when the meta has them (_retry_meta) and neither otherwise; a raised call and a crashed worker carry neither. The run record gains parse_retries: None when N is not an int of 1 or more, otherwise the sum over every chunk and the sweep (0 when nothing retried, including exits before any review unit). It is stamped after the sweep and before the total-failure exit, and every exit carries it through _run_record. --format json emits it after repo_context, and the README key list documents it. The timeout retry still replaces the first run's counts, the known gap in docs/llm.md. Test doubles for review_chunk/review_systemic with explicit keyword lists accept parse_retries=0, and the pinned run-record and JSON key lists add the new key. New tests in tests/test_parse_retry_orchestrator.py. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…pes-only (#20) repo_context's Java type regex, keyword set and JDK names are now the same objects as prxref.jvm_lang's JAVA_TYPE_RE, JAVA_KEYWORDS and JDK_NAMES, so there is one implementation; repo_contracts keeps importing _JAVA_KEYWORDS unchanged. The java/javax-only Java import regex stays, because jvm_lang.parse_import does not reproduce it exactly (it also accepts a semicolon-less import and rejects two imports on one line). Kotlin is answered here, before chunk_context is asked: - language_of maps .kt and .kts, in any case, to "kotlin"; - definition_regexes("kotlin") is (jvm_lang.KOTLIN_TYPE_RE,), types only, like Java: class of any flavour, interface, named object, typealias; - referenced_names for Kotlin uses the Kotlin keywords, drops JDK names and every segment (and the as-alias) of a kotlin.*, java.* or javax.* import on the added lines, and keeps only type-like names. kotlinx.* is kept. The module docstring now says chunk_context handles Java and Kotlin itself for the same-file context, and that this module keeps a types-only view for repository context with the shared facts in prxref.jvm_lang. A Kotlin file is therefore a same-language reader candidate for a Kotlin chunk. New tests in tests/test_repo_context_kotlin.py. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
eval run allowlists llm_parse_retries into run.json's config (seventeen keys, placed after llm_max_tokens with the other LLM settings). eval score passes the loaded PRXREF_LLM_PARSE_RETRIES to judge_case as parse_retries, so the judge re-sends a rejected reply as the review does, and score.json's judge block totals the retries as parse_retries, beside llm_calls. score.md's judge cost line names the retries only when there were any, so a scoring run without one renders byte-identically. The review's own retries are not totalled in score.json: a record's parse_retries is null at PRXREF_LLM_PARSE_RETRIES=0 on a healthy review, which the metrics' total/missing shape would misreport as missing, and each case's record.json already carries its count. docs/evals.md documents the key, the judge's parse retries (a reply cut short at the token limit included), its trace files and the judge block key, and corrects the judge-cost wording for a retry that raised after a billed reply. judge_cost's docstring says the same. Two judge-block key order pins in the score and end-to-end tests gain the new key. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The alias test used `Queue`, which the JDK name set already drops, so it stayed green with the alias drop removed. It now aliases to a name no set covers, checks the name is kept without the import, and checks that an alias of another organization's import is kept. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Calls heuristics.toggle_pinned_off_findings(files) at both sites where release_shape_findings(files) already runs: the main path, where it is spliced in at the chunk/sweep boundary alongside release_shape (moving sweep_start with it, so it stays a CHUNK-side finding to apply_sweep_dedup), and the summary-only (no-chunk) path, for symmetry, threaded through _summary_only_run's new toggle_findings parameter. Unlike release_shape, the toggle finding sits on a real line (> 0), so the long comment above the main-path splice site is updated: the reworded tier of apply_sweep_dedup (on with dedup_similarity) now can compare the toggle finding against a chunk worker's own restatement on the same line, keeping only the higher-ranked one of the two, instead of never comparing it the way a line-0 finding is exempted. The orchestrator module docstring and quality.py's module docstring (the seventeenth-check sentence, now also naming the toggle check as the eighteenth) are extended to match. 🤖 Authored with Claude Code — Claude Sonnet 5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Drives every #21 acceptance bullet through cli._run_review (and cli.main for the exit code) on a two-chunk fixture patch read by the local-diff forge, with the real openai-compat client, reviewer and orchestrator against a local MockOpenAIServer. The server answers each review unit (two chunks and the sweep) from its own script of canned replies and counts that unit's calls, so counts do not depend on worker order, and an unscripted extra call is an HTTP 500 that shows up as a wrong count. Covered with PRXREF_LLM_PARSE_RETRIES unset (default 1), 0 and 2: malformed then valid (2 calls, findings kept, run record and JSON 1, meta parse_retries and first_error, the discarded reply in chunk0.attempt1.response.json, the same request resent); malformed twice (2 calls, the last of two different errors wins, exit 0 on a partial and a total failure, exit 1 only with PRXREF_FAIL_ON=error); {} twice (the no-findings error); a list, {} and an object without findings retried; the issue's own stray-quote reply; a length stop (1 call, the budget message, 8 meta keys) against the same bytes with an honest stop (retried); 0 (a {} reply is clean, garbage costs 1 call, record null, no attempt file, 8 meta keys, the empty-reply retry kept); 2 (3 calls); the sweep; and tokens and reported cost summed over every attempt. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…endencies (#20) Wire the jvm_lang and jvm_deps leaf modules into chunk_context, so a chunk that touches .java, .kt or .kts files gets the same definitions and dependencies blocks the js and python chunks already get. - _language maps .java to "java" and .kt/.kts to "kotlin" through jvm_lang.jvm_language; _definition_regexes and _keywords return the jvm_lang tables for those two languages. Every other language is unchanged. - referenced_definitions starts a Java or Kotlin entry at the annotation lines directly above the definition (at most MAX_ANNOTATION_LINES, clamped so the definition line always fits in max_lines_per_entry), reports the first annotation's line number, and counts the annotation lines toward the line cap. js and python decorators are not included, and non-JVM output is byte-identical. - dependency_versions hands a JVM file to jvm_deps.dependency_lines before the manifest lookup and merges its g:a@v lines into the same sorted, deduplicated result. build.gradle.kts and settings.gradle.kts are skipped with no read. - repo_crosschunk._SAME_FILE_LANGUAGES gains java and kotlin, so with a reader diff_definitions leaves a chunk's own JVM files to chunk context instead of repeating their definitions as diff-file entries. Tests: a new tests/test_chunk_context_jvm.py, the Java/Kotlin rows of test_repo_context_crosschunk.py re-pinned, and the chunk-context Java pin in test_repo_context_defs.py flipped. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Adds tests/fixtures/issue22: the issue's make-fixture.sh, the PR as pr.patch (format-patch main..feature/progress-notes under a fixed placeholder identity and date, byte-reproducible), the PR head tree as repo/, and the issue's description as desc.md. tests/test_issue_22_acceptance.py drives the production local path, cli._run_review with --diff-file, --repo-dir and --description-file, against the real openai-compat client and a local scripted server that answers every prompt with no findings, and asserts on the --format json payload and the recorded prompts: - at repo, the progress.py chunk's reader block lists assistant/history.py:9, assistant/state_store.py:15 and tests/test_engine.py:17, in that order, and nothing else in the prompt names either source file; the run record carries the same three shared-state reader rows; - at diff and off, no prompt carries the reader block; - at every level there is exactly one finding, the pinned-off toggle at assistant/progress.py:34, undropped and with the deterministic suffix; deleting the conftest pin, or setting it to "true", removes it; - repo_readers.MAX_READER_ENTRIES = 0 removes exactly the reader block and nothing else from the prompts, and no reader row from the record; - repo/ matches both the generator's head tree and the patch's post-image, and the patch carries only dev@example.com. repo/ holds the fixture project's own tests and code that the project's lint rules reject, so tests/fixtures/issue22/conftest.py keeps pytest out of it (collect_ignore) and tests/fixtures/issue22/ruff.toml keeps ruff out of it, while inheriting the project's settings. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ference (#20) Leaving every js, python, Java and Kotlin file of a chunk out of diff_definitions whenever a reader is given avoided showing a file's own definitions twice, but it also lost a definition one file of a chunk references and another file of the same chunk defines outside its hunks: chunk_context.referenced_definitions only looks inside the referencing file. diff_definitions now searches such a file only for the wanted names that do not occur as an identifier (chunk_context._IDENT_RE) on that file's own added lines. Every name referenced_definitions can show for the file is one of those identifiers, so nothing is shown twice; a hit keeps reason diff-file. When no name is left the file is skipped without a read, so a one-file chunk still makes no own-file read here. - repo_crosschunk module and diff_definitions docstrings describe the rule (the module's "knows no Java" line is gone). - New TestTheChunksOwnFiles pins, for java, kotlin, python and js: the cross-file entry, no entry and no read for a file's own reference, zero reads for a one-file chunk, and a name both files reference shown once. - test_repo_context_unit.py read lists and one read-cap entry count follow the removed own-file reads of one-file chunks. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…reads (#20) Since #20 the chunk context reads Java and Kotlin files and probes the directories above them for pom.xml, build.gradle.kts and build.gradle at every repository-context level, off included, so the #17 fixture's forge reads gained 23 JVM paths. The prompts are unchanged and the 0.15.0 golden stays byte for byte. - The off comparison now checks the prompts against the golden as before and the reads with every JVM path left out (_is_jvm: a .java, .kt or .kts basename, or pom.xml, build.gradle, build.gradle.kts, libs.versions.toml). - A new test checks the golden itself holds no JVM path, so the filter cannot hide one of its reads, and another pins the off runs' JVM reads literally and in order (21 manifest probes, then ConnectorService and TransportConfig) for both design points and all three off variants. - REPO_READS and the floor test's read list gain the same 23 JVM reads; the repo-mode control also shows the filtered comparison fails there. - The two read-cap starvation tests count the chunk context's one read of each of the chunk's 21 own Java files, which bypasses the repository reader's cap; the route-contract test's last-read check becomes "SPEC is the only non-JVM read". Every other assertion is unchanged. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add the 0.17.0 CHANGELOG section (date left as YYYY-MM-DD for the release): Java and Kotlin chunk context and dependency versions, Kotlin in repository context, PRXREF_LLM_PARSE_RETRIES with the run record, trace attempt files and the judge retry, the shared-state reader block, and the pinned-off toggle check; the default-behaviour change for replies without a findings list; the same-chunk definition fix; and the known limitations, including the deferred #22 part 1. docs/llm.md gains a Parse Retries section, retry costing under Which calls count, JVM definitions and dependencies, the as-built own-file and cross-file rule, and the reader block. docs/quality.md documents the toggle check, docs/prompt-templates.md says nothing checks the reply format's findings key, and the README repository-context table, the config.py PRXREF_REPO_CONTEXT row and the orchestrate_review docstring name the reader block and its first-attempt-only rule. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…0.16.0 Drives the real orchestrate_review over a new fixture monorepo: a Maven multi-module build (parent properties plus a BOM import, a versionless module dependency, a malformed module pom), a Gradle build with a version catalog reached through version.ref and a garbage subproject build.gradle, and a Python control, seven files, one per chunk. Asserts the literal dependency lines (property-resolved, BOM-managed, catalog), the skipped java.* and own-group imports, the annotated definitions entries from their first annotation line, the Kotlin fun and class, no keyword-named definition, the broken builds (no line, a completed run in the record, the trace and the log, and a control that reaches the enclosing build without them), one rendering per own-file definition at diff and repo with a live cross-chunk control, and the 40-entry cap. The regression golden was captured from the released prxref 0.16.0 installed from PyPI, never from the tree: the Python control and the sweep prompts are byte-identical, and every JVM chunk prompt differs only by the added dependency and definition blocks. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Bump pyproject.toml, prxref.__version__ and the uv.lock prxref stanza to 0.17.0. Date the [0.17.0] CHANGELOG section 2026-09-25, point [Unreleased] at v0.17.0...HEAD and add the [0.17.0] link. Correct the PRXREF_REPO_CONTEXT row of docs/env-vars.md: off adds no repository context, but chunk context's Java and Kotlin definitions and dependency versions do not depend on it, so off is no longer byte-identical to 0.15.0; and diff finds Kotlin type declarations as well as Java ones. Rewrite HANDOFF.md for 0.17.0: the module map for Java and Kotlin chunk context (#20), the parse retry (#21) and the shared-state readers and pinned-off toggle check (#22, parts 2 and 3), the lessons, the config key counts (68 keys, 69 accepted names), the release shape, the live check, and what is still open, including #22 part 1. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Severity consistency ran over every finding, the release-shape and pinned-off toggle findings included, so a model finding in the same file that shared a rare code token with one (or any finding with its normalized title) could raise it. The live check on #22's fixture saw the toggle finding, which has no model behind it, posted as an error in two of three runs and as a warning in the third. heuristics.is_deterministic marks a finding whose body ends with the checks' "(deterministic check, no model)" suffix, and apply_severity_consistency now leaves such a finding out the way it leaves out a dropped one: it is never raised, never raises another finding, and its text does not count toward a token's rarity. No other pass changes. Tests: a new class in tests/test_quality.py covers the token rule, the title rule, the finding as a would-be source, a transitive link and the rarity count, with a control whose unmarked body is raised; a new tests/test_deterministic_severity.py runs #22's fixture through the local review path with a scripted model error on assistant/progress.py and checks the toggle stays a warning, with a control that patches the helper off and sees it raised to error. Docs: the CHANGELOG Fixed entry; docs/quality.md's consistency row names both grouping rules and the exclusion, and the deterministic-checks section says which passes can still touch these findings (finding grouping can still raise one); a HANDOFF lesson and the verified count. Also rewrites the .env.example PRXREF_REPO_CONTEXT paragraph to match docs/env-vars.md, and corrects the HANDOFF registration row: only REASONS' order ranks repository-context entries. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The orchestrator comment and the heuristics module docstring still said a deterministic finding flows through the quality passes exactly like a model finding. Severity consistency now leaves them out, so both say so and name heuristics.is_deterministic. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
prxref automated review: ApprovedPR: Release 0.17.0: JVM chunk context, parse retry, shared-state readers · files reviewed: 103 🟥 0 error · 🟧 3 warning · 🔍 0 spec · ⬜ 1 outofscope
Reviewed by prxref · model=z-ai/glm-5.3-flash · 206245 tok · 130.9s
|
| if language == "python": | ||
| return (_PY_DEF_RE, _PY_ASSIGN_RE) | ||
| return () | ||
| return jvm_lang.definition_regexes(language) |
There was a problem hiding this comment.
🤖 🟧 [WARNING] Unreachable keywords branch after unconditional return (src/prxref/chunk_context.py:375)
The hunk at @@ -355,1 +375,3 @@ replaces return () with return jvm_lang.definition_regexes(language) followed by if language in _JVM_LANGUAGES: return jvm_lang.keywords(language). The return is unconditional, so the keywords branch on lines 376-377 can never execute; if keywords were meant to be returned for JVM languages (the new docstring text distinguishes definition regexes from keywords), that behaviour is lost, and at minimum the dead code should be removed or the return guarded.
Reviewed by prxref · model=z-ai/glm-5.3-flash
| _write_trace_files(trace_dir, JUDGE_TRACE_LABEL, system, user, result.text, {**trace_meta, "error": reason}) | ||
| return JudgeOutcome(**live, grades=None, error=reason) | ||
| _write_trace_files(trace_dir, JUDGE_TRACE_LABEL, system, user, result.text, trace_meta) | ||
| if retries: |
There was a problem hiding this comment.
🤖 🟧 [WARNING] llm_calls undercounts when a retry's invoke raises (src/prxref/eval_judge.py:296)
In the retry loop, an exception from client.invoke returns via _live(base, unit, elapsed, retries, price_table), which sets llm_calls = 1 + retries. But the failing invoke is itself an additional request beyond the 1 + retries parsed-retry calls, so a case that retries once and then raises reports 2 llm_calls while making 3 requests. The docstring's own contract ("llm_calls is the number of judge requests this call made ... otherwise 1 plus parse_retries") is inconsistent with the actual request count in this path. [containment boundary not stated]
Reviewed by prxref · model=z-ai/glm-5.3-flash
|
|
||
|
|
||
| def _matches(segments: Sequence[str], declared: Sequence[_Declared]) -> list[_Declared]: | ||
| candidates = [dep for dep in declared if _is_candidate(dep.group, segments)] |
There was a problem hiding this comment.
🤖 🟧 [WARNING] Zero-score tie emits every candidate dependency line (src/prxref/jvm_deps.py:139)
In _matches (src/prxref/jvm_deps.py:131-137), when no candidate's artifactId token matches any import segment, all candidates tie at score 0 and every tied candidate's line is emitted. An import like com.google.common.collect with two unrelated declared deps sharing a 3-segment group prefix would render both dependency lines. This is the documented tie rule, but a zero-score tie carries no matching evidence and may attribute the wrong artifact versions to a file.
Reviewed by prxref · model=z-ai/glm-5.3-flash
|
|
||
| Nothing reads this key yet -- it is wired to the reply-retry loop in a later | ||
| task. This file only pins the config surface: default, valid range, the exit-2 | ||
| path shared with the other int keys, and that the docs surfaces agree with the |
There was a problem hiding this comment.
🤖 ⬜ [OUTOFSCOPE] Stale docstring: claims the config key is unread, but this PR wires it (tests/test_parse_retries_config.py:5)
The module docstring says 'Nothing reads this key yet -- it is wired to the reply-retry loop in a later task', yet the same PR's src/prxref/orchestrator.py changes read llm_parse_retries and thread parse_retries into every chunk worker and the sweep. The comment misleads future readers about the feature's state.
Reviewed by prxref · model=z-ai/glm-5.3-flash
Release 0.17.0: Java and Kotlin chunk context (#20), a bounded retry for unusable model replies (#21), and code that reads the state a change writes (#22, parts 2 and 3).
The user-facing account is
CHANGELOG.md[0.17.0]; the maintainer's map isHANDOFF.md.What ships
PRXREF_REPO_CONTEXTlevel. Also adds Kotlin type declarations in repository context.PRXREF_LLM_PARSE_RETRIES(default1). An unusable chunk or sweep reply is sent again. The newparse_retrieskey appears in the run record, in--format jsonand in the trace. Set it to0to handle replies exactly as 0.16.0 did.repo, a new### Code elsewhere that reads state this chunk writesblock (thereaderkind, theshared-statereason).Behaviour change at the defaults
.java,.ktor.ktsfile adds new prompt blocks and forge reads, even atPRXREF_REPO_CONTEXT=off.offstill adds no repository context. On Repository context outside the diff: cross-file definitions, contract excerpts, cross-chunk links #17's fixture, theoffprompts remain byte-identical to its 0.15.0 golden. Only its forge reads grew, all JVM paths, pinned.PRXREF_LLM_PARSE_RETRIES=0restores 0.16.0.Verification
uv run pytest -q: 8703 passed.uv run ruff check src tests: clean.domestic.flash), N=3 per arm, interleaved. One factor was varied: the reader cap, 0 vs the default.Issues
Closes #20. Closes #21.
#22 stays open: part 1 (a bounded follow-up call for more context) is deferred to a later release. It needs a decision on the project's single-shot rule.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459