Skip to content

release: 0.14.0 — Azure DevOps, review rules, ticket scope, replay, CLI backends, cost, size advisory, spec grounding - #9

Merged
sblattj merged 108 commits into
mainfrom
release/0.14.0
Sep 24, 2026
Merged

sblattj merged 108 commits into
mainfrom
release/0.14.0

Conversation

@sblattj

@sblattj sblattj commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Release 0.14.0. The full notes are in CHANGELOG.md under [0.14.0], and the handoff with live-check results is in HANDOFF.md.

What's in it

Verified

  • uv run pytest -q: 4222 passed. uv run ruff check src tests: all checks passed. prxref --version prints 0.14.0.
  • Built sdist and wheel pass the release leak gate; the gate's control, the withdrawn 0.13.0 sdist, fails it as it should.
  • Live checks, read-only, on 2026-09-23: GitHub (fix CVE 2024 47081: manual url parsing leads to netloc credentials leak psf/requests#6963), a public Azure DevOps project read anonymously, and gitlab.com merge requests. Results for each issue are under ## Verified at release in HANDOFF.md.

Closes #1
Closes #2
Closes #3
Closes #4
Closes #5
Closes #6
Closes #7
Closes #8

🤖 Authored with Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459

Squash of the two rebased spec-branch patches (REB 0001 + 0002) onto the
0.13.0 release, so the foundation lands as one bisectable commit. Both
original commits were authored in session
68f0a704-ce12-4ecc-878a-f432b0aac171; their bodies follow.

feat: spec-grounded review (WIP, pre-rebase snapshot)

Adds a second review axis: the operator supplies scope/intent via repeatable
--spec / PRXREF_SPEC_SOURCES (web URLs, local files/dirs, Jira ticket URLs),
which a new src/prxref/specs.py fetches, prunes to a deterministic diff-relevant
constraint digest, and injects into the existing worker and systemic prompts.
Violations surface as a new 🔍 `spec` severity ranked below warning, non-blocking
on fetch failure, with a golden eval dataset under tests/evals/.

fix: post-rebase repairs for spec-grounded review on 0.12.0

Teach the release-shape-foldin test doubles the new spec_digest keyword, and
add `spec` to the severity vocabulary row in docs/quality.md, which landed on
main after the feature branch was cut.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…_review

The --spec flag and PRXREF_SPEC_SOURCES were loaded into the config but never
handed to the orchestrator, so every CLI and webhook review ran ungrounded.
The TestSpecFlag assertions now observe the orchestrator kwargs instead of
the llm factory's config, and a new test covers all six keys.

— Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The grounding note only reaches a posted summary, so a --no-post, dry-run,
or inline-only run said nothing when a spec source failed. Each failed
source now logs a redacted WARNING and the run logs one INFO line with the
fetched/constraint counts. Also pins where the digest lands under
--trace-dir: every unit's user.md, never its system.md.

— Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The spec severity must quote its violated constraint verbatim, and normative
text is conditional ("If a session already exists, the server MUST ..."), so
the if-still / unless-already rules dropped legitimate spec findings as
hedged. The Spec: "..." quote is now removed before the rules read the body;
a hedge in the model's own words still drops.

— Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
… kwargs

The spec branch passed spec_digest positionally after reader, the slot main
uses for all_files. The rebase already converted every call site to a
keyword; the `*` makes a future positional call a TypeError instead of a
silently swapped argument.

test_run_review_passes_only_real_orchestrate_kwargs captures the REAL
orchestrate_review signature before fake_runtime swaps the module, and
asserts every kwarg _run_review sends is a parameter. The fake accepts
**kwargs, so a misspelt kwarg otherwise passes the suite while the real call
raises TypeError, which `review` swallows to exit 0. Mutation-checked: an
injected bogus kwarg fails it.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The body of test_orchestrator.py's autouse _contract_stubs moves to
tests/conftest.py as the non-autouse fixture `contract_stubs`.
test_orchestrator.py keeps an autouse wrapper that requests it, so its 199
tests behave exactly as before (control: a wrapper that does not request it
turns 45 of them red). Other modules now opt in with
@pytest.mark.usefixtures("contract_stubs") instead of importing and
re-patching the stubs, or leave it off to run the real reviewer, which the
spec-grounding proof tests need.

The fixture imports the stubs from tests.test_orchestrator at call time, so
loading conftest never imports a test module. test_issue_06, test_issue_10
and test_issue_13 still define their own module-level `contract_stubs`
(chunk + sweep only); pytest's module-over-conftest override keeps them
unchanged.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
docs/issues/ holds local issue bundles written during triage; they are
private planning notes and must never be published. Three guards, each proven alone:

- .gitignore ignores docs/issues/.
- [tool.hatch.build.targets.sdist] excludes docs/issues. Observed with
  `uv build --sdist` and an untracked probe file: either guard alone keeps it
  out of the tarball; with neither, the probe ships.
- tests/test_repo_hygiene.py asserts `git ls-files docs/issues` is empty
  (control: a force-added file in a scratch repo turns it red), that the
  ignore rule matches, and that the sdist exclude is declared. The git checks
  skip without a git executable or outside a checkout of this project (an
  unpacked sdist), both observed. The test carries no deny list.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…BASE)

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
docs/issues/ is private and gitignored as of 0.14.0, so the module
docstrings and comments in the issue-03/05/12 regression tests that cited
docs/issues/... paths pointed at nothing a reader can open. Replace each
with a neutral "2026-09-04 inbox report, issue N (private, not tracked)"
citation that keeps the issue identity without the dead path. Test logic
and verbatim fixture strings are untouched.

Also rewrite the live-instance-verification followup doc's Origin
paragraph, which named a private `sharpen` skill and a session UUID, as a
neutral reference to "a retrospective of the 2026-08-30/31 live-instance
session" so the doc carries no private skill name or session identifier.

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
… URLs

SEC-2: _fetch_jira built basic auth whenever the email and token were set,
so any ticket-shaped URL (including plain http on a third-party host)
received the operator's Jira credentials. Auth is now sent only when
PRXREF_JIRA_BASE_URL, the email and the token are all set, which routes
every ticket to the host the operator typed. Credentials without a base
URL are withheld with a WARNING naming PRXREF_JIRA_BASE_URL; an http base
URL is honoured with a WARNING; neither logs a value. The anonymous hint
now also fires on 404 (Jira Cloud hides private issues behind 404), and
explains that credentials only go to the base URL when they are set
without one.

COR-2: _TICKET_PATTERNS take a 0-2 segment context path on the /browse/
and REST shapes (so https://host/jira/browse/KEY-1 reaches
https://host/jira/rest/api/2/issue/KEY-1 instead of being scraped as a
page), recognize the Cloud team-managed issue view, and read a Cloud
board's selectedIssue with parse_qs on /jira/ paths. Bitbucket Server
file URLs (four segments before /browse/) still do not match.

COR-4 (Type/Labels) and backlog 9b: the ticket text leaves out empty
Summary/Type/Labels lines instead of rendering them bare, a 200 whose body
is not JSON is a clean source error naming its content type, and a JSON
body without a fields object is one too.

Tests: the two tests/test_specs.py cases that asserted auth with an empty
base now set one. NEW tests/test_specs_jira.py covers every URL shape and
negative, where credentials go, the hints, the non-issue bodies, the
rendered text, and an entry-point test that drives cli._run_review against
a local recording server (no Authorization without a base URL; a control
with the base URL set sees Basic auth). The new file has 53 tests; 29 of
them fail against the pre-change specs.py, including the entry-point test.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Lands the LLM-layer foundation for 0.14.0 (seat F1-LLM). Every
signature and constant the #66/#67 feature seats build on is in
place; the feature bodies come later.

#61: litellm no longer needs PRXREF_LLM_BASE_URL. The base-URL check
now runs only for openai-compat/ferry/http. It still runs before the
models check, so a run with both values unset still names the
endpoint first. When the base URL is set on any other backend, it is
ignored with one INFO line and never forwarded. So a litellm
deployment that set a placeholder URL to get past the old check
keeps its routing after upgrading, and there is no api_base
pass-through (G4 L4). docs/llm.md gets the matching #61 sentences.

#66 (factory): three constants, OPENAI_COMPAT_BACKENDS, CLI_BACKENDS
and BACKENDS, plus DEFAULT_CLI_CONCURRENCY. An unknown
PRXREF_LLM_BACKEND is now a ConfigError that names the variable and
lists the six accepted values. The check is the factory's first step,
so it gives exit 2 instead of the old LLMError, a failed review that
exited 0. claude-cli and kiro-cli go to
llm_cli_backends.build_cli_client, which the factory imports lazily.
It passes the models, timeout, reasoning effort, PRXREF_LLM_CLI_PATH
and PRXREF_LLM_CLI_CONCURRENCY. The concurrency value is re-checked
here as an integer >= 1, defaulting to 2. An explicitly set
temperature or seed logs one WARNING saying it is not applied. The
new llm_cli_backends module is a fail-closed stub: both entry points
raise "PRXREF_LLM_BACKEND: <backend> is not wired in this build",
which exits 2.

#67 (fields): InvokeResult gains cost_usd (None = not reported,
never 0.0) and cost_source, after finish_reason.

The module and factory docstrings are written in their final form for
all three issues. The existing unknown-backend test now expects a
ConfigError, and test_litellm_selection unsets the base URL. The new
tests/test_llm_factory.py covers the vocabulary, the gated base URL,
the CLI wiring through a recorder, the stub contract (signatures,
fail-closed, stdlib-only) and exit 2 through cli.main. It asserts
only what still holds once the real CLI clients land.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…t, leaks no origin

The spec digest and the grounding note disagreed with each other and with
the prompt about what a grounded run is. Five fixes, one contract:

- specs.constraint_count (new, public) replaces orchestrator's private
  _SPEC_CONSTRAINT_RE/_spec_constraint_count. The regex is regrouped so the
  (MUST|SHOULD|MAY) label binds to spec lines only: [ticket:KEY] lines
  carry no label, so a Jira-only run counted 0 constraints (LIVE-5,
  INT-5, COR-4). Exact counts now: Jira-only 4, file-only 2, both 6.
  _spec_note and the INFO grounding log line call it; the log line's
  wording is untouched.
- build_spec_digest returns "" when sources were given but no unit was
  extracted from any of them (every source failed, or none held a
  constraint), so the prompt shows its no-specs text (LIVE-4). The test
  runs before the budget cut, so a too-small budget still yields the
  intro plus the truncation marker, and no sources still yields the
  intro.
- Failed sources get no "nothing diff-relevant kept" line, that line
  names the source by _origin_short, and _origin_short falls back to the
  bare hostname (never netloc) and strips a trailing slash from paths, so
  no URL query, userinfo, port, or full local path reaches the LLM (SEC-4).
- Ranking interleaves sections, so the digest re-emits a heading whenever
  the open section changes and writes "[spec:{short}] (heading) (no
  section)" before a headingless unit that follows a headed one; HTML
  <hN> elements become "#"*N heading lines, ignoring block tags nested
  inside the heading, dropping zero-width spaces and pilcrow permalinks,
  and falling back to plain text for a heading never closed (COR-3).
- _spec_note labels each failed source "source N (kind)", or "source N"
  when the kind is unknown, ahead of its redacted reason (SEC-5, note
  half). The "Spec fetch failed for N source(s):" prefix is unchanged.

tests/test_specs.py flips test_failed_source_explained (all-failed digest
is "") and test_empty_source_explanatory_line (short origin, no /docs/).
New tests/test_specs_digest.py covers the counts, the exclusions, the
empty digest, the no-leak digest and every real-reviewer prompt, the
section scoping and HTML headings, and the note labels.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…content

_real_response sets _content_consumed so a streamed read of the fixture
(iter_content, the fallback a later streaming reader takes when raw has
no read1) yields the same bytes that .json() parses. Without it the
non-JSON and no-fields cases would fail on a None raw stream once the
Jira fetch reads its body by streaming, which is a fixture failure and
not the clean source error those tests pin.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Leaf modules for 0.14.0 that the later seats build on (F1-LEAF):

- text_inputs: bounded, fingerprinted reads of operator-named text files.
  Streams in 64 KiB chunks, keeps at most max_chars, hashes the raw bytes
  (equals shasum -a 256), validates strict UTF-8 across the whole file,
  folds CRLF/CR to LF, refuses non-regular files before opening, and
  refuses a path under the cwd whose symlinks resolve outside it (SEC-3
  loader half). Plain OSError-family errors with separate strerror and
  filename so callers can report a path-free reason.
- costs: price-table parsing (inline JSON, a JSON file, or a mapping;
  strict validation with ConfigError messages), reported / estimated /
  unknown per-unit and per-run cost (unknown is None, never 0 or a partial
  sum), and the USD formatter and cost label.
- markers: the one read-only severity glyph table (outofscope is now the
  white square; the blue square is only the out-of-ticket prefix), the
  fallback marker, scope labels and marker_for.
- triage: SCOPE_IN / SCOPE_OUT / SCOPE_UNKNOWN, SCOPES and a strict
  normalize_scope with no synonym table.

Tests: tests/test_text_inputs.py, tests/test_costs.py and
tests/test_markers.py, including AST checks that each module stays a leaf.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…, normative tokens unscored

Spec constraint extraction matched one physical line at a time, so a
hard-wrapped MUST lost its tail, a paragraph holding several keywords
became one mislabelled unit, a `# comment` in a code fence was read as a
heading, and any line near a date became a MUST (COR-1 / LIVE-3, backlog 8).

_spec_units now groups lines into blocks first: a paragraph or list item
joins its wrapped and indented continuations; blank lines, headings,
setext underlines, fences, table rows and new list items end a block.
Table rows and fenced lines stay single-line units and are never headings.
Blocks over 400 chars or with two or more matching sentences are split by
an abbreviation-aware sentence splitter (e.g., i.e., initialisms, code
spans) into one capped unit per matching sentence with its own strength;
otherwise the block is kept whole. A prose `can`/`discouraged` or bare
pin keeps just its sentence. A unit ending its block with `:` carries the
following list items up to the cap. Units anchor L{n} and doc_idx on the
block's first original line.

_is_version_pin_line now uses _STANDALONE_PIN_RE after peeling list
markers and decoration, so only a pin on its own line becomes a MUST.

COR-5: _rank_units excludes the new _NORMATIVE_TOKENS from relevance, so
a `must` comment or `required=True` in the diff no longer reorders the
digest. quality._STOPWORDS is untouched.

New tests/test_specs_units.py covers the eval corpus (case-001/002/003),
block boundaries, sentence units, colon lead-ins, pins and COR-5.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…1-SPEC-UNITS)

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…F1-SPEC-DIGEST)

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…-SPEC-JIRA)

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…YGIENE)

🤖 Authored with Claude Code

— Claude Sonnet 5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Foundation stage C, seat F2-DATA (#63, #64, #67 plumbing).

reviewer.py: new public fill_template(template, values) fills every named
placeholder in ONE re.sub pass, so a PR description, ticket or rules text
that quotes {diff} (or any other placeholder) renders literally instead of
receiving that value; unknown braces stay literal. Both renderers use it.
New frozen PromptContext (rules_worker, rules_sweep, ticket_scope,
ticket_context, spec_digest; scope_active = bool(ticket_scope)) and
NO_PROMPT_CONTEXT replace the spec branch's spec_digest= keyword on every
hop. Assembly order per contract section 5.1: SYSTEM = template head, then
rules (worker or sweep framing), then ticket scope, joined by blank lines;
USER = Review Context, then {ticket_context} (always ending in one blank
line when set), then Spec constraints, then diff/digest. With nothing set
the system prompt is the template head and the user prompt is unchanged.
_finding_from/_invoke_and_parse gain accept_scope (scope read through
normalize_scope only when the prompt asked for it). The reviewer meta and
<unit>.meta.json reserve cost_usd=None and cost_source="" (8 base keys)
for W67A to fill.

triage.py: Finding.scope = "unknown", appended as the last field.

orchestrator.py: _run_workers, _run_worker (both _invoke_chunk calls,
timeout retry included), _invoke_chunk and _run_sweep thread
prompt_context; _coerce_finding(accept_scope) gates the dict-path scope;
orchestrate_review builds the one PromptContext from the spec digest.

prompts: worker.md and systemic.md gain the {ticket_context} slot between
Repo and Spec constraints.

tests: the 7 strict doubles in test_orchestrator.py take prompt_context;
test_reviewer.py's three spec_digest= call sites pass a PromptContext;
new tests/test_prompt_context.py (66 tests).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…reasons

SEC-3: a spec file or directory source under the working directory must
still resolve under it once symlinks are followed. _dispatch runs
text_inputs.confine_to_cwd before anything stats or reads the path and
uses only the resolved path afterwards, so a symlinked source, directory
root or parent component is refused (live or dangling alike), while an
absolute operator path outside the cwd is read as given. _fetch_dir walks
os.scandir and keeps only entry.is_file(follow_symlinks=False), skipping
every symlinked entry wherever it points. _read_capped now reads through
text_inputs.read_capped_file, so at most max_chars characters are held
(strict UTF-8, BOM dropped) and SOURCE_TRUNCATION_MARKER is appended
exactly when the file is longer than the cap.

Backlog (4): one unreadable or undecodable file is skipped on its own
instead of failing its directory; skipped names are logged at WARNING,
or become the source's error when nothing could be read.

SEC-5 (reasons): no source error carries a local path any more. A
missing path reads "not a URL or path", an empty directory "no
.md/.markdown/.txt/.adoc files in directory", and the fetch_specs fence
reports an OSError as its class plus strerror (other exceptions, and
OSErrors without strerror such as requests' ConnectionError, keep their
message).

New tests/test_specs_local.py: 29 cases, 20 of which fail against the
base specs.py, including an orchestrate_review end-to-end check that a
symlinked secret never reaches a prompt, with a regular-file control that
does.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…5/56)

Adds every remaining 0.14 config key to _DEFAULTS and its coercion/range
tables, the config.py module docstring, .env.example and docs/env-vars.md:

- #66 PRXREF_LLM_CLI_PATH, PRXREF_LLM_CLI_CONCURRENCY (int, > 0)
- #67 PRXREF_PRICE_TABLE (parsed at load by the new _check_price_table via
  costs.parse_price_table, so a malformed table is exit 2 naming its
  source), PRXREF_POST_COST (bool, literal "1")
- #68 PRXREF_SIZE_WARN_LINES / _FILES (int, None = off, >= 0) and
  PRXREF_SIZE_IGNORE_GLOBS (list)
- #63 PRXREF_REVIEW_RULES, PRXREF_REVIEW_RULES_MAX_CHARS (int, > 0)
- #64 PRXREF_TICKET_CONTEXT_FILE, PRXREF_TICKET_CONTEXT_MAX_CHARS (int, > 0)
- #62 PRXREF_AZURE_DEVOPS_TOKEN, PRXREF_AZURE_DEVOPS_WEBHOOK_SECRET

Rewords the changed keys: PRXREF_LLM_BACKEND gains claude-cli/kiro-cli,
PRXREF_LLM_BASE_URL is required only for openai-compat/ferry/http,
PRXREF_LLM_MODELS is comma- or whitespace-separated, and the Jira keys say
credentials are only ever sent to PRXREF_JIRA_BASE_URL (SEC-2). Documents
the list grammar once in the config docstring and the size thresholds as
the second "None means off" class in _check_ranges.

load_config now copies list defaults instead of sharing them: the shallow
dict(_DEFAULTS) let a caller that appended to its config's list mutate the
defaults every later load starts from (found by a new test; affected
llm_models and spec_sources before this change too).

CNV-1: docs/env-vars.md states 55 keys / 56 accepted names with bullets
37/10/3/5, and HANDOFF.md's count moves to 55/1/56.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…2-SPEC-LOCAL)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…keys

Parallel seat F2-DATA adds cost_usd (null when the backend reported no
cost) and cost_source ("" when cost_usd is null) to every
<unit>.meta.json written under PRXREF_TRACE_DIR, per reviewer
_write_trace_files on seat/F2-DATA. The PRXREF_TRACE_DIR row enumerates
the meta.json keys, so it names the two new ones.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…s right

SEC-6: a spec host that answers and then trickles its body a byte at a
time used to hold a review indefinitely, because iter_content blocks until
a whole 8 KiB chunk arrives and a trickle never trips the per-read timeout.
URL and Jira sources now read through one _read_stream that takes a single
socket read per call (raw.read1, content encoding undone; iter_content with
chunk_size=1 where read1 is absent) and checks a monotonic deadline of
SPEC_FETCH_BUDGET_S = 30 s, taken before session.get, between reads. The
body is capped at 4*max_chars+4 bytes. Jira now streams too: an over-cap
or non-JSON body is a clean source error and ticket text is cut at
max_chars with the truncation marker. The spec session retries once with
no backoff sleep and ignores Retry-After, so a 503 asking for 30 s is
skipped rather than slept on.

COR-6: a charset-less text/* page no longer decodes as requests'
ISO-8859-1 guess. The order is the Content-Type charset, then for HTML a
<meta> charset in the first 1024 bytes, then strict UTF-8 with a BOM
dropped, then cp1252 with replacement. A charset Python cannot decode
text with falls through instead of failing the source. _FakeResponse no
longer defaults to utf-8; it takes requests' own header-derived encoding.

Backlog 2: the truncation marker is appended after HTML stripping, so a
cut inside a <script> or an open tag no longer swallows it, and a cut
landing exactly on max_chars is still marked.

New tests/test_specs_reader.py drives a real 127.0.0.1 server:
close-delimited, chunked and Jira trickles, a Retry-After: 30 no-sleep
test with an honouring control, fast-body and gzip controls, charset
cases and marker placement. 29 of its 43 tests fail against the base code.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
_read_stream reads through resp.raw.read1, which bypasses requests'
exception mapping. A body shorter than its Content-Length, a bad gzip
body, or a host that sends headers and then stalls past the read timeout
raised urllib3's ProtocolError, DecodeError or ReadTimeoutError. None of
those is an OSError, so they reached the fetch_specs fence as a different
kind of failure than the same read through iter_content. The fence's SEC-5
reason rendering keys on OSError.

The reader now re-raises them as iter_content does: ChunkedEncodingError,
ContentDecodingError, ConnectionError for a read timeout, and SSLError.
Real-socket tests cover all three failures. The reader and fetch_specs
tests fail against the previous commit; an iter_content control raises
the same types on both.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
GitLab get_diff now logs one WARNING per too_large/collapsed entry, as
get_compare_diff already did, so a hunkless file in a live MR review is
visible in the log. A parallel call keeps get_compare_diff and the shared
renderer byte-identical; the rendered diff does not change.

Docs and docstrings brought in line with the code:
- forges.md: the Azure DevOps Diffs bullet names 410 alongside 404
  (_fetch_blob); the GitLab Diffs bullet describes the paged listing
  (_iter_pages, 100 per page, 50 pages), the FeedReadError on a short read,
  and the header-only warning.
- base.py: the FeedReadError docstring notes that GitLab get_diff raises it.
- llm.md: PRXREF_LLM_REASONING_EFFORT reaches openai-compat and claude-cli
  only; the seed is sent by openai-compat and litellm, never by the CLI
  backends (which warn when it is set); PRXREF_LLM_MAX_TOKENS is not applied
  by the CLI backends.
- .env.example and env-vars.md: an empty or unset PRXREF_LLM_SEED derives
  a per-process seed (_auto_run_seed); it is not omitted.
- test_forge_compare_contract.py: Azure DevOps fixture provenance bullet.
- test_evals.py and spec-grounded-review.md 7.1: test_eval_replay.py now
  runs each case through the pipeline; scoring is still manual.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
New tests/test_release_seams.py covers four seams that no single feature
seat could test, because each one needs code from more than one seat:

1. #67 config to CLI: test_malformed_price_table_exits_2_naming_the_variable
   (D67 §9.7). Four malformed values of PRXREF_PRICE_TABLE (invalid JSON, a
   negative price, an unknown field, a missing file) each make
   main(["review", ...]) exit 2. stderr names the variable, and nothing past
   config runs. A control proves a valid table reaches the orchestrator parsed.
2. #67 cost through the real reviewer (W67B follow-up). An LLM double
   reports a distinct cost for each call: three chunks plus the sweep. The
   run record's cost_usd equals their sum and cost_estimated is False. The
   null rule: one call with no figure, a chunk or the sweep, makes the run
   cost None, never a partial sum, and logs one INFO line naming the model.
   The same run is estimated when the parsed price table prices that model.
3. #68 x #67: with the size advisory and post_cost both on, the real
   summary.md carries each exactly once. The advisory comes first, before
   the heading. The cost label is the last field of the attribution (§5.3).
4. #63 x #64 x spec: --rules-file, --context-file and --spec go through
   main to the real worker and sweep prompts. The system prompt holds the
   rules, then the ticket scope. The user prompt holds the ticket context,
   then the spec constraints, then the diff or digest (§5.1).

Every seam went red under at least one mutation, and each mutated file was
restored byte-identical (checked with cmp).

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…, pass list

Release-wide sweep of stale sentences reported by wave-1/wave-2 seats,
each re-read against the code that the sentence describes.

- docs/deploy.md §5: the "deliberately no PRXREF_FAIL_ON" bullet was false.
  It now says the default `never` never gates and `error`/`any` is the
  explicit opt-in (cli._fail_on_exit). The exit-code table gains the `1`
  row that knob produces. The unrecognized-URL example used a Bitbucket
  Server URL that detect_forge now accepts, and quoted an outdated message.
  It now shows an issues link and the message cli._cmd_review prints.
- docs/deploy.md §2: the Bitbucket Cloud row listed the Server event keys.
  It now lists pullrequest:created / pullrequest:updated, and new Bitbucket
  Server / Data Center and Azure DevOps rows are derived from webhooks.py
  (_BITBUCKET_SERVER_EVENTS, _verify_azure_devops).
- docs/spec-grounded-review.md §3.2: constraint units are sentences matched
  inside blocks, not lines (specs._spec_units). Records the known
  limitation: unpunctuated keyword lines in one paragraph merge into one
  unit that carries the highest strength among them.
- README.md: the architecture diagram's forge box names all five adapters.
- quality.py module docstring: apply_severity_map is now listed as pass 1
  of thirteen, matching orchestrate_review's sequence.
- reviewer.review_systemic docstring now quotes prompts/systemic.md, and
  the duplicated OpenAICompatClient docstring sentence is removed.
- tests: TestSeverityMapStub is renamed to
  TestSeverityMapRewritesOnlyMappedWords, and the test_run_record docstring
  no longer calls the cost and size hooks inert.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
tests/test_release_seams_w2.py pins three seams that no single feature seat
could test, because each needs code from two merged seats.

- #65 replay x #62 Azure DevOps: cli._replay_forge accepts every built-in
  forge's real adapter (BUILTIN_FORGES of test_forge_compare_contract, so a
  new forge joins on its own) for a pinned-SHA replay and returns a
  ReplayForge, directly and through main(), on a session that fails on any
  request. A control shows a forge without get_compare_diff still exits 2.
- #65 replay x #67 cost: a --diff-file replay through main() and the real
  orchestrator, with an LLM double that reports its cost, emits the replay
  stamp and cost_usd / cost_estimated, the cost being the sum over every
  call. A normal run has no replay key in its run record, its JSON payload
  or its run-start trace event, as CONTRACT 3.3 says.
- #66 CLIs x #67 cost: the real ClaudeCLIClient and KiroCLIClient behind a
  full run. claude's total_cost_usd makes the run cost reported, winning
  over a price table. kiro leaves it None, never 0, even when the table
  prices its model, and one kiro unit makes a claude run unknown rather
  than a partial sum.

Each seam was shown to fail under a mutation of the product code, then
restored byte-identical.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Bump pyproject.toml, prxref.__version__ and the uv.lock prxref stanza to
0.14.0. Add the CHANGELOG [0.14.0] section: spec-grounded review and
issues #1-#8 on the new tracker (with a note that older entries cite the
previous tracker), the Changed/Fixed/Security entries including the
withdrawn releases and rewritten history, and a Known limitations
subsection. Rewrite HANDOFF.md for 0.14.0: what landed, what the build
taught, the foundation/wave/gate/REL release shape and the tag-driven
release workflow, the gate run at this commit, and what is still open.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ating exit codes, pass count)

Brings the last stale sentences in line with the code at release/0.14.0:

- spec-grounded-review.md §7: the As-built note and the §7.1 heading now say
  an offline replay run of every case shipped (tests/evals/test_eval_replay.py)
  while the designed harness and the scoring did not.
- tests/evals/test_evals.py: the spec-grounded pipeline is no longer "future".
- env-vars.md: PRXREF_LLM_REASONING_EFFORT and PRXREF_LLM_MAX_TOKENS rows say
  which backends apply them (OpenAICompatClient, ClaudeCLIClient._argv,
  build_cli_client, _CLIClient.invoke); the temperature and seed rows say the
  CLI backends send neither; the quality-pass pointer names the two relabel
  passes that run before the eleven.
- llm.md Determinism: the default temperature is sent by the two API
  backends only.
- README.md: the pass sentence names the team severity map and spec grounding
  ahead of the eleven passes, in orchestrate_review's order.
- README.md, deploy.md, env-vars.md: the PRXREF_FAIL_ON exit-code text says
  precisely which outcomes exit 1 (active findings that trip the policy, or
  an exception before a result returns); a verdict-Error review returns a
  result with no active finding and stays 0 (_fail_on_exit, _error_run).
- deploy.md §5 row 2 adds the allowed-vocabulary clause README already has.
- orchestrator._stamp_run_cost docstring: a unit that counted 0 input tokens
  is never estimated (costs.run_cost rule 3).

Docs and docstrings only; no behaviour change.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The CHANGELOG link list stopped at 0.12.1 and [Unreleased] compared from it.
HANDOFF listed two follow-ups REL-FIX has since landed, and described the
size-advisory/cost-label seam as never tested together while REL-T tests it
on the main summary post.

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The 0.4.0 promise was that under either gating value a review that fails
to complete also exits 1. Only a crash did: orchestrate_review turns a
forge that cannot be read, a diff that cannot be parsed or chunked, and a
total chunk failure into an Error run instead of raising, and
_fail_on_exit found no active finding in it and returned 0, so a gating
lane read those broken runs as green.

_fail_on_exit now returns 1 with a stderr note for a verdict-Error result
whenever the policy is not never. never still returns 0 first, and a
completed result without a findings list is still tolerated.

Tests: the test pinning the old behaviour is flipped (error and any), with
a never control and a completed-without-findings control, plus one
end-to-end case through the real orchestrator with a forge whose get_diff
raises (error -> 1, never -> 0). The README, deploy.md, env-vars.md,
llm.md, the config and cli docstrings are brought back to the promise,
and the replay notes say a blank diff file or empty pinned range is gated
too. CHANGELOG [0.14.0] Fixed gains one bullet.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…-FAILON)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The #64 scope ask lives in the SYSTEM prompt, but the ## Output Format JSON
example ends the USER prompt and its finding had no "scope" key. A live replay
of tests/evals/case-002 with its ticket on openai/gpt-4.1-mini returned 0 of 4
responses with a "scope" key (10 of 10 findings "unknown"), while
anthropic/claude-haiku-4.5 answered scope on 8 of 8 with identical prompts:
the model copies the example it read last.

Both templates gain one {scope_example} slot glued to the example finding's
last value. reviewer._render_prompt and _render_systemic_prompt fill it
through the existing fill_template call with ',\n      "scope": "in"' only
while PromptContext.scope_active, and with "" otherwise, so a no-ticket
prompt is byte-identical to BASE (worker and sweep, system and user) and the
system prompt keeps exactly one ## Ticket scope block.

Live after the fix, gpt-4.1-mini with the ticket, 2 runs: 4 of 4 responses
carry a decoded "scope" key and 12 of 12 findings are in/out (12 in, 0 out).
The no-ticket control's trace prompts cmp identical to BASE's; haiku still
scopes 8 of 8 (7 in, 1 out). Spend $0.0166.

Tests: tests/test_issue_64_rendering.py pins the no-ticket prompts free of
"scope" and of any leftover {placeholder}, the active-ticket example parsing
with "scope" == "in" as its last key for worker and sweep, the scope line
being the only user-prompt change, the unchanged system prompt, and the
same through the real orchestrator for every ticket state.
tests/test_prompt_context.py's 0.13 chained-replace oracle strips the new
slot as it already strips {ticket_context}.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…ribution

Issue #7 row of the 0.14 contract: claude-cli's total_cost_usd is the
call's price at API list rates, not what a subscription is invoiced, but
the -v line and the posted attribution printed it as a bare "$0.0202".

- costs.cost_label gains keyword-only api_equivalent: a reported figure
  renders "$0.0202 (API-equivalent)"; "cost unknown" and "~$x (est.)"
  are unchanged (an estimate keeps "(est.)").
- costs.api_equivalent_run(units): true when at least one received unit
  reported a cost and every such unit's cost_source is "claude-cli".
- orchestrator._stamp_run_cost derives it once into
  run_inputs["cost_api_equivalent"]; _cost_label uses it at every exit,
  the cost-accounting crash path resets it, and _run_record writes it
  into the result ONLY when True (like replay), so a run not priced by
  claude-cli keeps its exact record, and the JSON payload (a whitelist)
  gains no key.
- cli._fmt_cost reads the same record key for the -v line.
- docs/llm.md: where the label shows, with the exact string; the Kiro
  timeout sentence gains the 2026-09-23 live figures (24-27 s one-line,
  28-37 s review call; 120 s is the floor).
- tests/test_issue_67_cost.py: label forms, the flag rule, end to end
  through the real orchestrator (summary and error notice, mixed and
  estimated runs unlabelled, JSON keys and trace events unchanged) and
  through main() for -v and --format json.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
Three small fixes from the 0.14.0 live verification seats.

Total LLM failure counts: when every chunk worker failed but the systemic
sweep answered, orchestrate_review's total-failure exit reported
chunks_reviewed 0 and chunks_failed == chunk_count. _error_run gains a
keyword-only chunks_reviewed (default 0, so the get_pr, get_diff, parse and
build_chunks exits keep their counts), and the total-failure exit passes the
units that succeeded. The verdict (Error), the reason and the posted notice
are unchanged; PRXREF_FAIL_ON still gates on the verdict. Six assertions in
tests/test_orchestrator.py pinned the old count (the contract sweep stub
succeeds there) and now read the true one.

GitLab diffs: get_diff no longer sends access_raw_diffs. gitlab.com returned
byte-identical /diffs bodies with and without it on 4 of 4 MRs; only the
deprecated /changes endpoint reads it. The paging test now asserts it is not
sent, and docs/forges.md drops it from the Diffs bullet.

docs/forges.md: the GitLab Thread List bullet says thread dedup needs
PRXREF_GITLAB_TOKEN even on a public gitlab.com project, because gitlab.com
answers anonymous /notes and /discussions with 401.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…SCOPE)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…OST)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…LF1)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
The forge.get_diff trace span recorded len(raw), a count of characters,
under the key "bytes". A diff with non-ASCII text or a byte-order mark
read short: a live Azure DevOps PR showed 15647 against 15651 UTF-8
bytes. It now records len(raw.encode("utf-8")). New
tests/test_release_docs_w3.py drives orchestrate_review with a trace
file, as TestRunTrace does, and checks a non-ASCII diff, a BOM diff and
an ASCII control.

Docs:
- docs/forges.md: the Azure DevOps pinned-range method was run live on
  2026-09-23; replace the stale "not been run against a live server"
  sentence with what that run showed.
- docs/env-vars.md, .env.example: list the "(API-equivalent)" cost form
  beside "(est.)", with "(est.)" taking precedence.
- docs/deploy.md: the all-chunks-failed example now shows 1/3, because
  the sweep is a review unit and can answer over a dead worker pool.
- README.md, docs/deploy.md exit-code row 0: drop the redundant and
  over-broad "An empty diff is not an error at all" sentence; a replay's
  blank diff file or empty pinned range is an Error run.
- README.md: the scope example in the worker and sweep output format;
  link CONTRIBUTING.md from the Quickstart.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
CHANGELOG [0.14.0]: add the access_raw_diffs Changed bullet and the
total-failure sweep-count Fixed bullet, the forge.get_diff byte-count
Fixed bullet, the gitlab.com thread-token note in the documentation
corrections, and the claude-cli "(API-equivalent)" cost label in the
dollar-cost bullet. Known limitations gain spec crowd-out and the
unrecognized --pr-url exit 0 under gating, and the Azure DevOps bullet
now lists what was verified live.

HANDOFF: the gate count comes from the release tip, a public "Live
checks" list replaces the placeholder comment, and Still open drops the
done access_raw_diffs and CONTRIBUTING.md items and adds the GitLab
token note, spec crowd-out, the gating exit and the scope-fixture
follow-up.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…-DOCS-A)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
…-DOCS-B)

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
HANDOFF's gate block now shows the merged tip's 4222 passed. docs/env-vars.md
scopes its exit-0 sentence to the default PRXREF_FAIL_ON=never and says
"empty PR diff", matching the exit-code tables: a blank replay diff is an
Error run.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
@github-actions

Copy link
Copy Markdown

🤖 prxref review — Error

The review could not complete: get_diff failed: 406 Client Error: Not Acceptable for url=[redacted]

No findings were produced.

Reviewed by prxref · model=unknown · 0 tok · 0.8s

The 0.14.0 release PR's own CI review ended with verdict Error:
GET /pulls/{n} with the diff media type answered HTTP 406, "the diff
exceeded the maximum number of lines (20000)", code too_large. Every
release has this limit. docs/forges.md, CHANGELOG Known limitations and
HANDOFF's Still open list now say so, and HANDOFF names the fix: a
fallback that rebuilds the diff from the paged /pulls/{n}/files
listing.

🤖 Authored with Claude Code

— Claude Opus 5.5 via Claude Code

Claude-Session-Id: f915b0ff-b27c-46f3-bc7c-430dca61d459
@sblattj
sblattj merged commit daec2d0 into main Sep 24, 2026
4 checks passed
@sblattj
sblattj deleted the release/0.14.0 branch September 24, 2026 17:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment