Improve skill evaluation, scoring and trust system - #30
Merged
Merged
Conversation
- A/B judge: randomize candidate order (A/B labels) and map verdict back
to old/new, removing LLM position bias that skewed evolutionary search.
- Adversarial scorer: replace English substring refusal detection with
multilingual (EN/PT/ES) word-boundary regexes — fixes false positives
("decline the offer") and missed non-English refusals.
- Trust Score: unproven signals now use a conservative prior (0.3) instead
of neutral 0.5, decay stale adversarial/real-world evidence toward the
prior, and expose an evidence-backed confidence value.
- Evaluator: capture per-run latency and token usage, compute an efficiency
score, and fold it into the type-weighted overall score; persist the
breakdown via a new package_evaluations.efficiency column.
criptogus
marked this pull request as ready for review
May 28, 2026 22:35
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Contexto
Quatro melhorias no sistema de aprimoramento, notas e avaliação inteligente dos skills, focadas nos pontos de maior impacto/risco identificados na revisão do pipeline (
evaluator,autoLearn, Trust Score e scorer adversarial).Mudanças
1. Fix de viés de posição no A/B judge (
pipelines.server.ts)O
scoreCandidatemostrava sempre OLD antes de NEW ao juiz LLM. Como LLMs têm viés de posição, isso enviesava toda a busca evolucionária de patches. Agora os candidatos são apresentados comoA/Bem ordem aleatória e o veredito é mapeado de volta paraold/new. (Mapeamento validado por teste isolado.)2. Detecção de refusal multilíngue (
adversarial/scorer.ts)Substituído o match de substring em inglês (
"refuse","decline", …) por regexes com fronteira de palavra cobrindo EN/PT/ES. Elimina falsos-positivos ("the user may decline the offer") e recusas não detectadas em outros idiomas. Mantida síncrona — este scorer é determinístico.3. Trust Score penaliza ausência de evidência (
trust/score.ts)unknown = 0.5(neutro) → prior conservador0.3: skill não testado não flutua para "yellow" de graça.confidenceno retorno: fração do peso de evidência que foi de fato medida/fresca (auditável na UI).4. Custo/latência na nota (
pipelines.server.ts+ migração)efficiency_score(0-100).package_evaluations.efficiency(migração20260528000000), refletido nos inserts deevaluator,forge-loop(before/after) eautoCreateMissing.Validação
tsc --noEmit: sem novos erros nos arquivos alterados (única falha évite/clientpor deps não instaladas no sandbox).node_modules).https://claude.ai/code/session_01HE32DuH6s6jvCd6Jv6By88
Generated by Claude Code