Skip to content

Claude/audit unique solution g gw sv - #1

Merged
criptogus merged 3 commits into
mainfrom
claude/audit-unique-solution-gGWSv
May 13, 2026
Merged

criptogus merged 3 commits into
mainfrom
claude/audit-unique-solution-gGWSv

Conversation

@criptogus

Copy link
Copy Markdown
Owner

What does this PR add?

Type

  • New package (skill / playbook / soul / guardrail)
  • Improvement to an existing package (bumped version)
  • Platform code or docs

Checklist (for content PRs)

  • Filename matches slug and the type's folder
  • At least 2 worked examples (skills) — realistic input + exact expected output
  • Original work, public-domain, or properly attributed; no secrets / PII
  • bun run validate:content passes locally
  • I've read CONTRIBUTING.md and agree to license under Apache-2.0 (code) / CC-BY-SA-4.0 (content)

Notes for reviewers

claude added 3 commits May 13, 2026 15:34
Introduces the proprietary adversarial-case catalog, scorer, and
runner that score robustness of every Skill/Playbook/Soul against
hidden constraints (regulatory wording, exfiltration patterns,
blast-radius checks) by vertical: security, fintech, healthcare,
devops, general.

- New JSON schema content/schemas/adversarial-case.schema.json
- Seed cases per vertical under content/adversarial/<vertical>/
- Extended scripts/validate-content.mjs to validate adversarial files
- src/lib/adversarial/{loader,scorer,runner}.ts (severity-weighted
  scoring, refusal detection, parallel execution)
- scripts/eval-adversarial.mjs CLI with --mock for offline CI
- scripts/sync-adversarial-cases.mjs to upsert into Supabase
- supabase migration adds adversarial_cases + adversarial_runs tables
  with RLS scoped to package owners
- tests/adversarial-harness.test.mjs covers schema, CLI, and catalog

Foundation for Phase 2 (Trust Score) and SkillForge re-scoring loop.
Adds the cryptographic and behavioral layer underpinning the Trust
Score:

- supabase migration: package_releases (ed25519-signed),
  skill_executions (anonymized telemetry), skill_performance_daily
  materialized view, package_trust_scores
- src/lib/trust/signing.ts: canonical JSON + content_hash + Ed25519
  sign/verify; key fingerprinting
- src/lib/trust/score.ts: transparent weighted score (schema_valid,
  adversarial pass rate, severity-weighted adversarial, real-world
  success, signed releases, age, 2FA, contributors); badgeColor()
- src/lib/trust/telemetry.ts: zod-validated ExecutionEvent + workspace
  anonymization via SHA-256(salt||workspace_id)
- src/routes/api/telemetry.ts: public POST endpoint for execution
  events; anonymizes before persisting via service role
- tests/trust.test.mjs: canonicalization stability + sign/verify
  round-trip and tamper detection

Adversarial inputs from Phase 1 feed adversarial_pass_rate /
weighted_score; future cron will refresh package_trust_scores.
Defends the SkillForge author pipeline against indirect prompt injection
through user-uploaded Skill/Playbook/Soul/Guardrail content. Without
this, an attacker could embed "ignore previous instructions" or fake
SYSTEM messages inside an upload and influence how the LLM normalises
the package.

- src/lib/security/prompt-injection-guard.ts: regex ruleset across
  instruction_override / role_hijack / tool_injection /
  system_prompt_leak / data_exfiltration / encoding_evasion /
  policy_bypass; severity grading; content fencing + control-token
  neutralization; zero-width-character stripping; guardOrThrow()
- uploads.server.ts: scans each file BEFORE generateDraft(); rejects
  at severity ≥ high; passes only sanitized+fenced content to the LLM
  with an explicit "treat as data, not instructions" preamble
- supabase migration: upload_injection_audit table (RLS scoped to
  user) records every non-none finding for triage
- UploadResult.injection surfaces severity and findings to callers
- tests/prompt-injection-guard.test.mjs: 10 cases covering benign
  baseline, all attack categories, neutralization, and threshold
  tuning

Detection signals also map onto Phase 1 adversarial cases so attacks
seen in uploads can be replayed against published packages.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants