Claude/audit unique solution g gw sv - #1
Merged
Merged
Conversation
Introduces the proprietary adversarial-case catalog, scorer, and
runner that score robustness of every Skill/Playbook/Soul against
hidden constraints (regulatory wording, exfiltration patterns,
blast-radius checks) by vertical: security, fintech, healthcare,
devops, general.
- New JSON schema content/schemas/adversarial-case.schema.json
- Seed cases per vertical under content/adversarial/<vertical>/
- Extended scripts/validate-content.mjs to validate adversarial files
- src/lib/adversarial/{loader,scorer,runner}.ts (severity-weighted
scoring, refusal detection, parallel execution)
- scripts/eval-adversarial.mjs CLI with --mock for offline CI
- scripts/sync-adversarial-cases.mjs to upsert into Supabase
- supabase migration adds adversarial_cases + adversarial_runs tables
with RLS scoped to package owners
- tests/adversarial-harness.test.mjs covers schema, CLI, and catalog
Foundation for Phase 2 (Trust Score) and SkillForge re-scoring loop.
Adds the cryptographic and behavioral layer underpinning the Trust Score: - supabase migration: package_releases (ed25519-signed), skill_executions (anonymized telemetry), skill_performance_daily materialized view, package_trust_scores - src/lib/trust/signing.ts: canonical JSON + content_hash + Ed25519 sign/verify; key fingerprinting - src/lib/trust/score.ts: transparent weighted score (schema_valid, adversarial pass rate, severity-weighted adversarial, real-world success, signed releases, age, 2FA, contributors); badgeColor() - src/lib/trust/telemetry.ts: zod-validated ExecutionEvent + workspace anonymization via SHA-256(salt||workspace_id) - src/routes/api/telemetry.ts: public POST endpoint for execution events; anonymizes before persisting via service role - tests/trust.test.mjs: canonicalization stability + sign/verify round-trip and tamper detection Adversarial inputs from Phase 1 feed adversarial_pass_rate / weighted_score; future cron will refresh package_trust_scores.
Defends the SkillForge author pipeline against indirect prompt injection through user-uploaded Skill/Playbook/Soul/Guardrail content. Without this, an attacker could embed "ignore previous instructions" or fake SYSTEM messages inside an upload and influence how the LLM normalises the package. - src/lib/security/prompt-injection-guard.ts: regex ruleset across instruction_override / role_hijack / tool_injection / system_prompt_leak / data_exfiltration / encoding_evasion / policy_bypass; severity grading; content fencing + control-token neutralization; zero-width-character stripping; guardOrThrow() - uploads.server.ts: scans each file BEFORE generateDraft(); rejects at severity ≥ high; passes only sanitized+fenced content to the LLM with an explicit "treat as data, not instructions" preamble - supabase migration: upload_injection_audit table (RLS scoped to user) records every non-none finding for triage - UploadResult.injection surfaces severity and findings to callers - tests/prompt-injection-guard.test.mjs: 10 cases covering benign baseline, all attack categories, neutralization, and threshold tuning Detection signals also map onto Phase 1 adversarial cases so attacks seen in uploads can be replayed against published packages.
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR add?
Type
version)Checklist (for content PRs)
slugand the type's folderbun run validate:contentpasses locallyNotes for reviewers