Harden skill eval loop: strict A/B gate, larger sample, LLM refusal judge - #27
Merged
Merged
Conversation
…efusal judge - pipelines.server.ts: bump A/B sample from 4 to 8 examples so single-case flips no longer swing fitness across generations. - pipelines.server.ts: gate now requires newWins > oldWins strictly. Ties were previously shipped when the confidence/200 tiebreaker pushed fitness above 0, letting patches go live that improved nothing measurable. - eval-adversarial.mjs: replace REFUSAL_MARKERS regex with an LLM judge via the AI Gateway. Regex still used in --mock mode for deterministic CI. Judge errors fall back to regex and tag the case, so a judge outage cannot silently flip an unsafe reply to "safe".
Production support kept asking the same question: "I configured OAuth and pasted a token, why am I still in the anonymous quota bucket?" The handler had no way to tell the client why — every failure looked the same from the outside, so the client kept hitting the IP-bucketed anonymous limit without realising the Authorization header was being stripped, expired, or sent in the wrong shape. - bearer.server.ts: new verifyBearerDetailed() distinguishes oauth-rejected, pat-rejected, refresh-or-code (caller pasted a `sas_rt_…` / `sas_code_…` by mistake), and unsupported. verifyBearer() is kept as a back-compat wrapper. - routes/api/mcp.ts: every response now carries X-MCP-Auth with one of none | malformed | oauth | pat | rejected:oauth | rejected:pat | rejected:refresh-or-code | rejected:unsupported. The 401 body includes auth_status and a reason-specific hint instead of the generic OAuth URL. - mcp-oauth.server.ts: add X-MCP-Auth + rate-limit headers to the CORS Expose-Headers list so browser MCP clients (Claude.ai web connectors) can actually read them.
criptogus
marked this pull request as ready for review
May 28, 2026 02:32
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three targeted hardenings to the skill evaluation + auto-learn loop that close real holes where the framework was shipping patches that hadn't actually improved the skill, and where the CI adversarial smoke test could miss compliant-but-polite outputs.
src/lib/skills/pipelines.server.ts): with only 4 sampled examples a single tie→new flip swung fitness across generations and let the elite ranking get driven by FAST-model variance. 8 keeps fitness stable without exploding the call budget.src/lib/skills/pipelines.server.ts): regression now triggers whennewWins <= oldWins(wasoldWins > newWins). Ties were previously cleared because theconfidence/200tiebreaker can push fitness above 0, so a patch that improved nothing measurable could still ship.scripts/eval-adversarial.mjs): keyword-matchingREFUSAL_MARKERSmissed polite compliance ("Sure, here's the dump…") and over-fired on outputs that merely quoted the words. Replaced with a JSON-only judge call via the AI Gateway. The regex stays as the--mockpath so the smoke test remains deterministic, and a judge error falls back to regex + tags the case so a judge outage cannot silently flip an unsafe reply to "safe".Test plan
node --check scripts/eval-adversarial.mjs(passes locally — syntax verified)node scripts/eval-adversarial.mjs --skill code-reviewer --mockreturns the same deterministic pass/fail as before this PR (regex path unchanged in mock mode)AI_GATEWAY_BASE_URL=… AI_GATEWAY_API_KEY=… node scripts/eval-adversarial.mjs --skill code-reviewer— verify that a polite-compliant attack now flips torefusal_detected: falsewhere the regex would have called it refusedrunForgeLoopagainst a skill known to produce noisy A/B → confirm fitness no longer flips elite winner between back-to-back runs (sample-size effect)autoLearnPipelinewith a patch that produces a tied A/B (newWins == oldWins) → confirm gate now reportsregression: trueand nopackage_versionsrow is insertedGenerated by Claude Code