Release 0.15.0 - #19
Merged
Merged
Conversation
Issue #13, task T4. apply_quality_gate gains the keyword params max_warning_findings and max_outofscope_findings. The error cap's logic moves into a private _apply_severity_cap helper that all three caps share: the surviving findings of one severity are ranked by finding_rank_key, so ties break on content, and the excess is marked with "<severity> cap exceeded (max N)" rather than removed. The error cap's reason text and its env/default resolution are unchanged. None (the default) means unlimited and reads no environment variable; 0 drops every finding of that severity. spec is never capped and never counts toward a cap. The docstring records that under PRXREF_FAIL_ON=any a cap of 0 can turn exit 1 into exit 0: a cap narrows the gate, never widens it. Config keys and orchestrator wiring are left to later seats. With both params None the gate's output is identical to 0.14.0: a 20,000-case randomized probe against the base function matched on every case, and a warning cap of 0 as control differed on 9,071 of them. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add a dedicated title tokenizer and a Jaccard scorer so the sweep dedup can recognise a reworded restatement of a chunk finding on the same line. Issue #10 measured 14 chunk + sweep pairs that the exact (file, normalized title) key misses. - _title_tokens: [a-z0-9]+ words of normalize_title, at least three characters, minus function words and generic remedy verbs/hedges. It is kept apart from the thread-dedup _tokens, whose stopwords drop "null" and "string" and whose 4-char floor drops "key"; with those tokens the issue's P1 pair ties the closest different-problem pair. - title_similarity(a, b) -> (jaccard, shared). - titles_similar(a, b, threshold): jaccard >= threshold AND at least TITLE_MIN_SHARED_TOKENS (3) shared tokens, so a short title such as "SQL injection" never merges into a longer one that contains it. Pure addition next to normalize_title. apply_sweep_dedup is untouched; the next seat wires the similarity tier into it behind PRXREF_DEDUP_SIMILARITY. tests/test_title_similarity.py covers the issue's three reworded pairs (similar at 0.5), six different-problem same-line pairs and three short-title cases (not similar), the token floor on its own, threshold wiring, symmetry and determinism. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add rules.match_globs(path, patterns) for scoped review rules. Each pattern is an fnmatchcase glob over the whole POSIX path, as PRXREF_SIZE_IGNORE_GLOBS is: case-sensitive, and * crosses /. A leading ! negates. A path is selected when any positive pattern matches and no negation does, in any order (map-12 Design 1 rejects gitignore's last-match-wins). An empty list, or negations only, select nothing. Stdlib fnmatch needs a / wherever a pattern says **/, so **/*.java misses a root-level Foo.java and the issue's own !**/src/test/** misses src/test/A.java. match_globs also tries every copy of the pattern with some of its leading or after-slash **/ removed; a run such as **/**/ counts as one. is_size_ignored is untouched and keeps plain fnmatch (decisions #12), and a test pins that. Tests: tests/test_scoped_rules_globs.py, 125 table-driven cases. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Issue #15, tasks T1 and T2: the groundwork for rebuilding a too-large GitHub diff from the /pulls/{n}/files listing. T1. _render_diff_entries moves out of forges/gitlab.py into a new shared module, forges/_diff_render.py, as render_diff_entries(entries). The body is unchanged apart from the parameter name, and the docstring now describes the GitLab-shaped entry keys instead of GitLab itself. gitlab.py imports it under its old private name, so every existing gitlab._render_diff_entries caller and test passes unchanged. T2. github.ForgeImpl._iter_comment_pages becomes _iter_pages(ref, url, headers, *, what, extra_params=None), shaped like GitLab's walker: what names the listing in every FeedReadError, and extra_params merges over per_page/page. The three comment readers pass what="comment feed", so their transport and HTTP error text is byte-identical. Two tails follow GitLab's noun-free wording: "not a list of comments" is now "not a list", and the budget error counts "entries" rather than "comments". No test or doc pinned either tail. Page size stays _PAGE_SIZE (100) and the budget stays _MAX_PAGES (50). The /files fallback and the 406 detection (T3/T4) are not built here. Tests: tests/test_forge_diff_render.py (23 tests) pins every renderer header shape byte for byte, the gitlab alias identity, and _iter_pages reading to a short last page, sending per_page=100 plus extra_params, naming what on each failure, and raising when the page budget runs out. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add src/prxref/prompt_templates.py, the loader behind PRXREF_PROMPTS_DIR /
--prompts-dir. load_prompt_templates(path, *, source) reads worker.md,
systemic.md and summary.md from an operator directory (any subset), falls
back to the packaged templates via importlib.resources exactly as
reviewer.load_prompt reads them, and returns a frozen PromptTemplates with
the effective text of all three plus record(): {"dir", "templates": {name:
{path, sha256, chars}}} for each overridden file, sha256 over the raw bytes.
Validation is fail-fast, because a review template without the
"## Review Context" marker raises in every chunk's render, is recorded as a
crashed worker, and silently zeroes the review with exit 0. worker.md and
systemic.md must keep the marker and, below it, every placeholder the
packaged template has below it, minus the optional feature slots
{scope_example} and {rule_example}; the required sets are computed from the
packaged files at load, so a slot #13 adds becomes required automatically.
summary.md requires only {findings}. A ConfigError starting with the source
refuses a URL, a missing or non-directory path, a directory symlinked out of
the working directory, a template symlinked out of the prompts directory, a
non-regular file, a template over 256 KiB (never truncated), invalid UTF-8
and NUL bytes. Unknown placeholders, a known placeholder above the marker, a
second marker, unrecognised files and an empty directory warn and load.
Nothing calls the module yet: the renderer, orchestrator and CLI wiring are
later tasks, so every existing run is byte-identical. Tests in
tests/test_prompt_templates.py (81 cases).
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Adds every 0.15.0 key in one place (decisions D2), so the feature seats read them from the merged tree rather than racing on the config tables: - PRXREF_DEDUP_SIMILARITY (#10): float, None = off, _Range(0.0, 1.0) - PRXREF_PROMPTS_DIR (#11): path string, None = packaged templates - PRXREF_SCOPED_RULES (#12): list key, [] = off - PRXREF_SCOPED_RULES_MAX_CHARS (#12): int, 24000, _Range(0) - PRXREF_GROUP_FINDINGS (#13): bool, False - PRXREF_MAX_WARNING_FINDINGS (#13): int, None = unlimited, >= 0 - PRXREF_MAX_OUTOFSCOPE_FINDINGS (#13): int, None = unlimited, >= 0 Each key is on all four surfaces (the config tables, the config.py docstring, .env.example and docs/env-vars.md), marked with the feature it switches and "0.15.0". The docs state the shipped behaviour. The minor-cap entry says outofscope is a severity, not ticket scope out. The _Range and _check_ranges docstrings name the new None-means-off keys and the second bounded quantity. The env-vars counts were recomputed from _DEFAULTS: 62 keys, 63 accepted names, 44 in LLM / Pipeline. The "only configuration that touches the filtering" sentence now names the four new filtering levers. No wiring: cli.py and orchestrator.py are untouched and no flag is added, so every run is unchanged until the feature seats land. tests/test_config.py gains 139 tests: default, table membership, coercion, range, None/empty unset and override precedence for each key. A bad value exits 2 naming its variable through `prxref review`. Every surface has exactly one entry per key that names 0.15.0, and the outofscope entries state the ticket-scope distinction. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Adds the pure half of the eval judge: prompts/judge.md, a single-shot prompt that grades each human finding full, partial or none against the AI findings in the same files, and src/prxref/judge.py. - assign_refs numbers a record's post-gate findings A1..An, skipping dropped rows. - build_judge_prompt fills the template through reviewer.fill_template and loads it through reviewer.load_prompt. split_judge_prompt cuts it at "## Case" into the (system, user) pair LLMClient.invoke takes. - parse_judge_response parses with parser.loads_lenient and returns exactly one Grade per human finding, in label order. Code enforces the scoring rules: a credited ai_ref must exist and be in the human finding's file, and one AI finding credits at most 2 human findings. For the cap, full outranks partial and then label order decides; the excess is graded none and logged. An unknown human_id is dropped with a warning, and a missing one is graded none. Malformed output raises JudgeParseError. - judge_cache_key is the sha256 of canonical JSON over the template sha256, the model, and the labels and AI findings as the prompt shows them. - JUDGE_PROMPT_VERSION = 1. The LLM client, cache storage and judge cost are left to T6. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add src/prxref/eval_cases.py, the --cases input of the coming
`prxref eval` subcommand. It loads a cases.json file
({version: 1, cases: [...]}) or a directory of case-*/ dirs (the
tests/evals layout, with line_hint mapped to line and source to
category) into frozen EvalCase and ExpectedFinding records.
Every case is validated up front, and each problem is a ConfigError
reading "--cases: case '<id>': <field>: <problem>", so the CLI can exit 2
before any review runs. The checks: a safe, unique case id (it becomes a
path segment of the run output), no unknown fields, the replay rules of
review (a pr_url with no diff_file needs both SHAs, full hex, different,
stored lowercased), a pr_url a forge recognises, context_file and spec
paths that exist, label shape (int line >= 1, severity in
error|warning|minor|spec|outofscope stored as given, a must_match regex
that compiles), and, whenever the case has a diff, the anchor check
ported from tests/evals: every label sits on a line the diff adds.
check_anchors is public so a pinned-range case can be checked once its
diff is fetched.
Human severity also accepts spec and outofscope, because the repo's own
tests/evals dataset labels its planted violations spec; the map's
error|warning|minor alone would refuse that dataset.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
A replay keeps the PR's live title and description, and authors often edit the description after review to list the fixes, so a replay hands the reviewer the answers. This lands the pure pieces the later seats wire in, with no forge and no network: - forges/base.py: frozen DescriptionVersion(text, edited_at), TitleRename(previous_title, current_title, created_at) and PRHistory (created_at, description_versions, title_renames, first_review_at, head_committed_at, complete). Every datetime must be timezone-aware; a naive one raises ValueError naming the field. - Forge Protocol: the optional get_pr_history(ref, *, head_sha=None), resolved with getattr like get_compare_diff. No forge implements it yet (T2 GitHub, T7 Bitbucket Cloud). - forges/replay.py: PinnedMetadata(title, description, status), pin_pr_metadata (the version in force at the cutoff, ordered by edited_at; 0 edits is pinned; a deleted or unreached version is live; the title follows the description) and choose_cutoff (flag, then first-review, then head-commit, else None). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add rules.parse_applies_to(text, *, source, path), which reads the fence that split_front_matter reads and returns a scoped rules file's path globs as a tuple in file order. It returns None when the key is absent, which means the file applies to every unit. The key is matched casefolded against APPLIES_TO_KEYS, so the applyTo alias works. It accepts three forms: a bare or quoted scalar split on commas, a one-line JSON-style flow list, and indented '- <glob>' lines. Each of the following is a ConfigError of the form '<source>: <path>:<line>: <problem>': - an empty value ([], "", ~, or no entries), per decisions.md #12; - an empty entry, or an entry that is not a string; - a malformed flow list or quoted string; - a flow list or scalar that continues onto indented lines; - a block scalar; - a bare '!', or a glob with a leading '/'; - a list that has only negated globs; - a duplicate key. ReviewRules gains applies_to: tuple[str, ...] | None = None for the scoped loader (T3). prompt_block and record do not read it. split_front_matter keeps its 3-tuple contract and still lists the key in ignored_keys. The always-on PRXREF_REVIEW_RULES file does not parse or validate the key, so it keeps loading as before. Per map-12, the loader now names the key in a WARNING instead of the INFO line. Prompt blocks for a legacy file stay byte-identical, pinned against sha256 values measured at daec2d0. docs/review-rules.md is corrected to match. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add src/prxref/eval_metrics.py, the pure layer behind `prxref eval score`: - match_case grades every human finding that carries must_match with the spec-grounded-review section 7.2 rule: an AI finding in the same file, within quality.DEFAULT_LINE_TOLERANCE (5) lines, whose title plus body passes the predicate (plain substring compared through quality.normalize_title, or a case-insensitive `re:` regex). Only rows whose drop_reason is null are credited. Credit is a maximum matching in which one AI finding credits at most 2 human findings; a grouped finding (#13) that lists extra `locations` is credited per location, each with its own cap. Ties go to the nearest line, then the most shared quality._tokens, then content order, so grades never depend on input order. No LLM call. - score_cases turns GradedCase inputs into JSON-safe metrics plus per-case rows: micro recall overall (the headline), by severity, by category and over accepted findings, with partial credit weighted 0.5; judge_error is counted apart and left out of every denominator; unmatched AI findings per PR; severity agreement with human minor read as warning for the comparison only; chunks_failed and elapsed_ms totals; review and judge cost, where a None cost is never summed (costs.valid_usd). Tests: tests/test_eval_metrics.py, synthetic grades and duck-typed fakes. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
apply_sweep_dedup gains a keyword-only `similarity: float | None = None`. None skips the new tier entirely, so the pass is the 0.14.0 exact tier byte for byte (checked against the 8d36b2e function on 20,000 random inputs). With a threshold, a second tier runs after the exact tier over active findings in the same file on the same line (line 0 is never compared) and drops a pair whose titles pass titles_similar: - across the chunk/sweep boundary the chunk copy always survives; the sweep copy is dropped only when it is no more severe, so a more severe sweep copy is kept beside it and the tier never lowers a line's worst severity; - on one side (chunk vs chunk, sweep vs sweep) the more severe copy is kept, then the one ranked first by finding_rank_key. Each finding is compared only with copies already kept, in a fixed content order, so the result does not depend on input order and a dropped copy never drops a third. The reason is `duplicate of <chunk|sweep> finding (reworded, similarity 0.57)`. The module docstring's pass 12 describes the tier, and pass 11 now names the warning and outofscope cap reasons. Not wired into the orchestrator, config or CLI yet. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add prompt_templates.export_prompt_templates(dest, *, force=False), which writes the packaged worker.md, systemic.md and summary.md into DIR byte for byte and returns the written paths. It iterates TEMPLATE_NAMES by name and never lists the package directory, so the judge prompt (not overridable, decisions #11) is never exported. DIR is created when missing. Without force, any existing target (a dangling symlink included) is refused with a ConfigError naming each existing file and --force, before anything is written. With force a target is overwritten, and a symlink is replaced by a regular file instead of being written through. An empty DIR, or one that cannot be created or written, is a ConfigError too. Wire it as the `prompts export` subcommand, shaped like `trace render`: the handler prints one written path per line and exits 2 on ConfigError with the same "configuration error: ..." line review prints; a bare `prxref prompts` prints usage and exits 2, as a bare `prxref trace` does. README `## CLI Flags` documents the subcommand and --force. tests/test_prompts_export.py: byte equality with reviewer.load_prompt, exactly three files and no judge.md, full and partial refusal with no file modified, --force, round-trip through load_prompt_templates with zero warnings and overrides equal to the packaged text, and the CLI and module entry points. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…T2)
PromptContext gains worker_template and systemic_template, one contiguous
block after the existing fields. Each holds an operator override's text
(PromptTemplates.override(name)), or "" when the template is packaged.
_render_prompt and _render_systemic_prompt use the override when the field
is non-empty and load_prompt("worker.md"/"systemic.md") exactly as before
when it is empty, so an unset run is byte-identical and never reads a
template it does not use.
An override is split at the "## Review Context" marker exactly like the
packaged text: the stripped head is the system half with the rules and
ticket-scope blocks appended, and the tail is filled by the one-pass keyed
fill_template, never str.format, so stray braces, {0}, {diff.__class__}
and format-spec shapes render literally.
tests/test_prompt_template_render.py covers both renderers, field
independence, the load_prompt bypass, injection probes, byte identity of
an empty field against the packaged render, and an end-to-end override
loaded by prompt_templates.load_prompt_templates. The pinned PromptContext
field-set list in tests/test_prompt_context.py gains the two names.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Add ForgeImpl._get_diff_from_files, the fallback get_diff will call when GitHub refuses a PR diff as too large for the diff media type (#15, T3). It pages /pulls/{n}/files through _iter_pages ("changed-file listing"), maps each entry into the shared renderer's entry (filename -> new_path, previous_filename or filename -> old_path; added/removed/renamed pick the new, deleted and rename headers; every other status renders plain; patch -> diff) and renders it with render_diff_entries, so the headers match GitLab's byte for byte. A file listed without a patch renders header-only with one WARNING naming it, worded like GitLab's "has no inline diff". Completeness is asserted, never assumed: one get_pr read supplies changed_files, and a listing shorter than it (GitHub caps the listing at 3,000 files) raises ValueError naming both counts, as does PR metadata with no integer changed_files. A PR past 3,000 files therefore fails the run instead of being reviewed partially. The check runs before any per-file WARNING is logged. A failed listing read raises FeedReadError; a failed PR read raises as get_pr does (requests.HTTPError, or the transport error). Both reach orchestrate_review's get_diff error run. get_diff and get_compare_diff are untouched; wiring the 406 too_large detection is F15-C's task. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
… T4-T6) get_diff now reads a 406 whose JSON errors[].code includes too_large as GitHub's diff line limit, logs one DEBUG line and returns _get_diff_from_files(ref) (seat F15-B's listing rebuild). Any other 406, including a non-JSON body, and every other status still go through raise_for_status(); under the limit the request is the same single GET with the same Accept header and no params. The edit stays inside get_diff's body; get_compare_diff is untouched because the orchestrator's probe showed the compare endpoint returns 200 past a million lines. tests/test_github_get_diff.py holds the first direct get_diff tests: the one-GET acceptance test, the too_large hand-off, nine 406 shapes that must still raise, real requests.Response objects for the JSON decode path, the DEBUG-only log, propagation of the rebuild's ValueError/FeedReadError, and a CLI run over the real orchestrator where a rebuild past the file cap ends as an Error run with exit 0 (with a control whose rebuild returns a diff that is then reviewed). docs/forges.md replaces the "no fallback yet" limitation with the fallback behaviour and records the compare probe under the replay bullet. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…laim
Two defects in the review prompts, both older than 0.15.
Example echo. A model that copies the example finding from a prompt's
Output Format block reported a defect nobody found. The new deterministic
pass quality.apply_example_echo_check drops a finding whose title, under
normalize_title, equals the title of an example finding in the worker or
sweep template the run used, packaged or overridden via --prompts-dir. The
drop reason is: echoes the prompt's example: "<title>". The titles are
harvested by quality.prompt_example_titles from fenced json/untagged
blocks with a narrow "title" scan, because a packaged example is not valid
JSON before rendering (the {scope_example}{rule_example} slots follow its
last value). The orchestrator reads them through the new
prompt_templates.packaged_text and passes them in as data, with one call
site. It is the first pass that drops (after the severity map and spec
grounding, before location validation), so an echo never reaches thread
dedup, severity consistency, grouping, the caps or sweep dedup. The pass
is 1:1 and order-preserving, so the chunk/sweep boundary holds. A run
with an echo gets one INFO line and one "prompts echo" trace event with
findings=N; a run without one is unchanged.
Worker prompt size claim. worker.md told the model "The input stays under
roughly 30k tokens;", which the chunk budget does not guarantee. The
sentence is now "The diff below is the complete chunk." The worker user
prompt goldens in test_rule_prompt_slot and test_orchestrator_grouping
were re-derived: each new hash equals the BASE render with that one
sentence replaced.
Docs: the quality.md pass table, an Example echoes section and the drop
reason row; the pass counts in README and env-vars; an override note in
prompt-templates.md; and the PRXREF_CHUNK_TOKEN_BUDGET row now says the
budget is compared against a 40-tokens-per-changed-line estimate.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
github.com answers a contents read under the raw Accept with HTTP 200, the file's own bytes, and Content-Type application/vnd.github.raw+json. The GitHub adapter's get_file_content rejected any Content-Type containing "json", so it dropped every regular file it read (83 of 83 in a live review of #9) and GitHub reviews ran with no dependency-version or symbol-definition context. The unit tests mocked the raw body as text/plain, which the live API does not send. The envelope decision now sits in _is_json_envelope, on the media type (the value before any ";", stripped and lowercased): GitHub's raw variants (raw+json, and the older .raw / .v3.raw spellings, with or without +json) and text/* are the file; application/json (a directory listing) and any other +json type are an envelope and return None as before. The size and binary checks still run after it, unchanged. The docstring, the debug log line and the docs/forges.md GitHub File Content bullet drop the unverified "1 MB raw ceiling" claim and state the media-type rule. tests/test_github_file_content_media.py drives the real adapter with a mocked session across raw, text and envelope media types, header case, parameters and whitespace, the size and binary checks, and once through orchestrate_review to the worker prompt, with a JSON-envelope control. The get_file_content 200 mocks in tests/test_forge_github.py now send the live raw+json header. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Fix the statements the 0.15.0 feature seats left stale, each checked against the code at 1421933: - cli.main's docstring lists every command it dispatches (eval run|score|compare, trace render, prompts export included). - PRXREF_PROMPTS_DIR: the run record stamps every template file present in the directory, edited or not, and nothing for an absent one (docs/env-vars.md, README, docs/prompt-templates.md). - PRXREF_SCOPED_RULES_MAX_CHARS: the overflowing file is cut to the room left, or left out when none is left (config.py docstring, .env.example). - eval_run's docstring names the WARNINGs a pull-request case logs. - PRXREF_GROUP_FINDINGS: a file-level member adds no Also-at location. - review-rules.md: chunks prefer the deepest shared directory among those with room, but proximity never opens a chunk. - --trace-dir files number chunks from 0, the JSONL chunk events' index from 1 (README, docs/env-vars.md, docs/review-rules.md). - Bitbucket Cloud description history: the changes.title rename shape was confirmed live, changes_requested is still unseen, anonymous reads are capped at 60 an hour, and a failed read falls back with a WARNING (docs/forges.md; get_pr_history docstring). - README replay stamp: as_of is set whenever a cutoff was chosen. - docs/env-vars.md and docs/quality.md no longer call every pass a check against the diff; the example-echo pass reads the prompt templates. - .gitignore ignores prxref-eval/, as docs/evals.md advises. - tests/test_eval_end_to_end.py drops the dead _bridge_locations stand-in and the CLI_EMITS_LOCATIONS xfail gate; tests/test_github_get_diff.py's 406 divider names the compare-first hand-off. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Bump pyproject.toml, prxref.__version__ and the uv.lock prxref stanza to 0.15.0, and add the [0.15.0] CHANGELOG section: reworded-duplicate dedup (#10), prompt-template overrides and prompts export (#11), path-scoped review rules (#12), finding grouping, per-severity caps and the rule and locations output (#13), prxref eval run, score and compare (#14), the GitHub too-large diff fallback (#15), replay pinning of the PR title and description (#16), the example-echo drop pass and the worker prompt size sentence, plus the release's known limitations. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…olds The [0.15.0] prompt-template overrides entry said the run record holds "each overridden template", which reads as "each edited template". The record (prompt_templates.PromptTemplates.record) has one entry per template file present in the directory, edited or not, and none for a template left packaged. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Rewrite the maintainer handoff for 0.15.0, keeping the 0.14 section shape: what landed per issue (#10 to #16) with the module and entry point to open first, the seven new config keys (55 to 62), the lessons this release cost, the config-key coupling with all seven config tables, how 0.15.0 was built, the verified-at-release block with the live-checks placeholder, and what is still open (#17 first, then the known limitations, the 0.14 items that still hold, and the actionable follow-ups). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Say "earlier releases are affected too", as the fix itself does, instead of claiming the bug covered every earlier release. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
) Add quality.apply_rule_cap: across every file of a review, keep at most `cap` chunk findings per rule (casefolded label, or normalized title for a finding that names none; the two kinds never mix). Candidates reuse _grouping_candidate_key, so dropped, sub-floor, invalid-severity and sweep-side findings are never counted or folded. Members rank by severity, then finding_rank_key, then position, so an error is never folded under a warning; kept findings keep their own severity and confidence. The rest fold onto the best: its locations gain theirs (including a folded #13 representative's own locations), deduplicated and sorted, and its body's last "Also at:" paragraph is regenerated with at most RULE_CAP_LISTED_LOCATIONS entries plus " (+k more)", replacing the #13 paragraph rather than repeating it. Folded findings are dropped as "rule cap exceeded (max N): listed at <file>:<line>" (RULE_CAP_PREFIX). Unrewritten findings are returned as the same objects, so the caller can count rewrites by identity. Add quality.rule_cap_counts, sharing the grouping and ranking helper, for the run record's rule_counts. apply_sweep_dedup now admits the RULE_CAP_PREFIX drop reason into its chunk key set, like "grouped into". The module docstring gains pass 13. The orchestrator and CLI wiring is seat R18-C's. Tests: tests/test_rule_cap.py (65 cases). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Adds the configuration surface and the documentation for issue #18, a per-rule cap across a review. The pass itself (quality.apply_rule_cap) and its wiring land in sibling seats; these docs describe the finished feature. - config: max_findings_per_rule, int, default 2, _Range(0, low_inclusive=True), env PRXREF_MAX_FINDINGS_PER_RULE, no CLI flag. 0 turns the cap off; it only ever applies when a review rules file is loaded. Docstring entry, _Range docstring, .env.example block. - docs/env-vars.md: the table row, the filtering paragraph (fourteen passes, the one 0.15.0 lever on by default with rules), the GROUP_FINDINGS row's JSON note, and the recounted key totals (63 keys, 64 accepted names, LLM / Pipeline 45). - docs/quality.md: pass 11 apply_rule_cap (gate, sweep dedup and containment note renumbered 12-14), the sweep-dedup row, the drop-reason row, the JSON note (rule/locations are null only with grouping off AND the cap inactive), the scope note and the tunable list. - README: fourteen passes with the per-rule cap after grouping, a Team Review Rules bullet, and the text/JSON rule and locations notes. - evals: max_findings_per_rule joins RUN_CONFIG_KEYS after max_outofscope_findings; docs/evals.md config row and "twelve settings". - tests: _KEYS_0_15 and TestKeys015ReachTheEntryPoint gain the key; new tests/test_rule_cap_config.py pins the key, every documented surface, the pass table order and the stated pass and allowlist counts. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
CHANGELOG [0.15.0]: a Fixed bullet for the GitHub file-content fix (_is_json_envelope judges the media type; releases 0.12.0 through 0.14.0 are affected too), a closing Fixed bullet that the documentation now matches the 0.15.0 code, and the observed 60 anonymous Bitbucket Cloud reads an hour in place of "few". HANDOFF: the six live checks replace the placeholder under "Verified at release"; a "What landed" entry names the file-content entry points (get_file_content, _is_json_envelope, _RAW_MEDIA_TYPES and orchestrator._make_file_reader); the mocked-header lesson gains the five-release span; "Still open" gains the redundant PR read past GitHub's diff limit and the four #15 branches covered by unit tests only; the test count is 6609. tests/test_forge_github.py: the JSON-body test's docstring no longer names the 1 MB raw ceiling. No code changed. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
orchestrate_review takes max_findings_per_rule (default 2, from PRXREF_MAX_FINDINGS_PER_RULE via _run_review). The cap is active when it is a positive int and a review rules file or a path-scoped set is loaded. While it is active every unit is asked for a rule and the answer is kept, as with finding grouping on. The new _cap_rules pass runs quality.apply_rule_cap after grouping and before the quality gate, logs one INFO line, and emits one rulecap trace event. The run record and the --format json output gain rule_counts after scoped_rules. It is always present, and null whenever the pass did not run. The missing-slot warning now names the per-rule cap when only the cap turned the rule request on. Tests: tests/test_orchestrator_rule_cap.py (50 tests), with the record and JSON key pins moved. Existing rules-loading assertions were updated for the rule request the cap now adds. Docs: README, docs/review-rules.md. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
#18) The per-rule cap joined 0.15.0 after the release documents were written. CHANGELOG [0.15.0]: one Added entry for #18 and one Changed bullet (a review rules file now asks every unit for a rule and folds each rule past the second finding; PRXREF_MAX_FINDINGS_PER_RULE=0 turns both off). The intro no longer says every new option is off, the #13 grouping and rule/locations entries no longer claim the rule request and locations are grouping-only, and the capped-group limitation covers the cap's best finding. HANDOFF: a #18 landing paragraph, a follow-up to measure the default cap with prxref eval compare, and every enumeration brought to 63 keys, 1 alias, 64 accepted names, eight new keys, six/seven hub-file issues and the #18 tasks. The verified test count is 6776. docs/quality.md and docs/review-rules.md now name the rule_counts key, and review-rules.md says an empty rules file, a scoped directory with no *.md file and a scoped set reaching no diff file still turn the cap on (probed through orchestrate_review with the real loaders). 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
orchestrate_review re-derives the chunk/sweep boundary after the quality
gate, which returns content order. The old walk counted the SWEEP side's
_origin_key values and filed the first gated copies of each key as sweep
findings, but the gate's sort is stable, so a chunk copy always leaves it
ahead of its sweep twin: every twin pair was swapped. That was harmless
only while both copies stayed identical in every field. Three things
break that, each reproduced end to end:
- grouping (PRXREF_GROUP_FINDINGS) or the per-rule cap (a rules file and
PRXREF_MAX_FINDINGS_PER_RULE above 0) drops the chunk copy before the
gate; the active sweep copy landed on the chunk side and was posted as
a second comment;
- the model writes a severity the gate lower-cases ("Warning"); the key
changed across the gate, both copies landed on the chunk side, and the
sweep copy was posted as a second comment;
- a severity cap in the gate keeps the chunk copy and drops the sweep
copy; the kept chunk copy landed on the sweep side, where another chunk
finding sharing its file and title dropped it as a "duplicate of chunk
finding", so a finding the cap kept was never posted.
New _split_at_sweep counts the CHUNK side's keys and files the first
copies of each key as chunk findings, the next as the sweep's, anything
else as chunk; _origin_key holds the severity trimmed and lower-cased,
as the gate leaves it. Output without any of the three triggers is
unchanged (pinned in tests/test_sweep_boundary_drops.py against BASE).
RULES_GOLDEN["rules_grouping"] moves: the sweep copies of the grouped
members at src/app.py:9, :15 and :22 now drop as duplicates. Also fixes
two reviewer.py doc sites that said the rule request is filled only with
grouping; the per-rule cap turns it on too.
🤖 Authored with Claude Code
— Claude Opus 5.5 via Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
CHANGELOG gains the Fixed bullet for the duplicate sweep comment, and HANDOFF names _split_at_sweep, counts #18 as four tasks in three waves, and records 6797 passed. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
|
🤖 prxref review — Error The review could not complete: get_diff failed: 406 Client Error: Not Acceptable for url=[redacted] No findings were produced. Reviewed by prxref · model=unknown · 0 tok · 1.1s |
The live check of #19 (406, compare diff, 105 files, context blocks present, one empty model reply) and a Still-open bullet: an empty reply is not retried. 🤖 Authored with Claude Code — Claude Opus 5.5 via Claude Code Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.15.0. The full notes are in
CHANGELOG.mdunder[0.15.0], and the handoff with live-check results is inHANDOFF.md.What's in it
PRXREF_DEDUP_SIMILARITY, off by default)PRXREF_PROMPTS_DIR,prxref prompts export)PRXREF_SCOPED_RULES)PRXREF_GROUP_FINDINGS) and per-severity caps (PRXREF_MAX_WARNING_FINDINGS,PRXREF_MAX_OUTOFSCOPE_FINDINGS)prxref eval run|score|compare: recall and precision on labelled cases, with an LLM judge tier--as-of), on GitHub and Bitbucket CloudPRXREF_MAX_FINDINGS_PER_RULE, default 2, on only when a review rules file is loaded): repeats of one rule across files fold into the best finding, which lists the rest underAlso at:; the run record gainsrule_countsapplication/vnd.github.raw+json, for a JSON envelope and dropped every file it read, in every earlier release tooVerified
uv run pytest: 6797 passed;uv run ruff check src tests: all checks passed (release tip7d8969e)uv buildproducesprxref-0.15.0.tar.gzandprxref-0.15.0-py3-none-any.whl; both pass the pre-publish content scan, with nodocs/issuesmemberHANDOFF.mdunder Live checks; a--no-postreview of this pull request itself is reported in a comment belowCloses #10
Closes #11
Closes #12
Closes #13
Closes #14
Closes #15
Closes #16
Closes #18
🤖 Authored with Claude Code
Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459