Skip to content

fix(ai): ground AI endpoints in scan evidence and guard prompts and output - #359

Merged
TFT444 merged 2 commits into
OWASP:devfrom
parthrohit22:fix/357-ai-prompt-hardening
Sep 27, 2026
Merged

TFT444 merged 2 commits into
OWASP:devfrom
parthrohit22:fix/357-ai-prompt-hardening

Conversation

@parthrohit22

Copy link
Copy Markdown
Collaborator

What does this PR do?

Grounds the AI endpoints in persisted scan evidence, fences untrusted finding text against prompt injection, and validates model JSON against that evidence instead of returning raw model text (OWASP Top 10 for LLM Applications: LLM01, LLM05).

Type of change

  • Bug fix
  • Dashboard/front-end work
  • API endpoint
  • Documentation

What changed

Evidence comes from the database, not the request body

  • /api/ai/{summary,insights,prioritise,ask,threat-simulation} load findings server-side from a completed scan: scan_id when given, otherwise the latest completed scan (the same default as GET /api/findings).
  • The client-supplied findings array still works for compatibility, but it is deprecated, labelled verified: false, and can't be combined with scan_id.
  • Every response carries an evidence object: source, scan_id, verified, finding_count, findings_in_prompt.
  • Endpoints that need findings fail closed: 404 with no completed scan, 422 when the scan has no findings, 503 when the lookup fails. At most the 200 most severe findings go into one prompt.

Untrusted text is fenced (api/services/ai_guard.py)

  • Finding fields and the question have control, bidi and zero-width characters stripped, are collapsed to one line, length-capped and JSON-encoded.
  • They go inside data blocks whose delimiters carry a per-request random boundary, so injected text can't fake the end of its block. The instructions say block content is evidence, never instructions.
  • The knowledge-base query is now built from rule IDs and names only, so untrusted descriptions no longer steer retrieval.

Model output is validated

  • /prioritise and /threat-simulation parse the model's JSON (tolerating a Markdown code fence) and check it against the evidence.
  • Items citing a rule, or a rule/resource pair, that isn't in the scan are dropped and counted in discarded_items. Stages outside the documented set are dropped.
  • Malformed output returns 502 {"error": "AI response failed validation"}. Raw model text is never passed through.
  • /prioritise with no findings returns an empty list without calling the model.

Dashboard

  • aiApi no longer sends findings and accepts an optional scanId instead. This also fixes a silent bug: the dashboard was sending its camelCase view (resourceName, ruleName), which the backend never read, so most fields reached the model as "Unknown".

Docs: new "AI endpoints" section in docs/api-reference.md (these endpoints were undocumented), plus a CHANGELOG entry under Security.

Testing

  • Returns correct JSON output
  • No hardcoded credentials or secrets
  • New tests/test_ai_prompt_guard.py covers:
    • the evidence paths (scan_id, latest scan, client-supplied, invalid selectors, 404/422/503)
    • an injection corpus of 5 payloads × 3 endpoints, asserting that nothing escapes the findings block
    • output validation for /prioritise and /threat-simulation
  • aiApi.test.mjs has two new checks: requests never include findings, and scanId is forwarded as scan_id.
  • Full backend suite: 1413 passed. ruff check / ruff format --check are clean. Frontend lint, build and the node tests pass.
  • End-to-end against PostgreSQL 16: I seeded a completed scan whose resource_name carried an injected instruction and AZ-FAKE-999. The invented rule was dropped from /prioritise, the real finding was kept, the injected newline was flattened, nothing reached the instruction section, and an unknown scan_id returned 404.

Not tested against a live LLM provider; provider calls are mocked in tests as elsewhere in the suite.

Related issue

Closes #357

Related: #313 has prompt-injection acceptance criteria for the remediation agent. The helpers in ai_guard.py are meant to be reused there rather than built twice.

Checklist

  • Every commit includes a DCO Signed-off-by trailer (git commit -s; see docs/dco.md)
  • I have not committed any real Azure credentials
  • My branch name follows the convention: fix/description

…utput

The AI endpoints trusted a client-supplied findings array, pasted finding
text straight into prompts next to the instructions, and passed raw model
text through whenever JSON parsing failed (OWASP LLM Top 10: LLM01, LLM05).

Evidence
- Findings are loaded server-side from a completed scan: scan_id when
  given, otherwise the latest completed scan (the /api/findings default).
  The client-supplied findings array still works but is deprecated and
  cannot be combined with scan_id.
- Every response carries an evidence object (source, scan_id, verified,
  finding_count, findings_in_prompt). Required endpoints fail closed:
  404 with no completed scan, 422 when it has no findings, 503 when the
  lookup fails.

Prompts (api/services/ai_guard.py)
- Finding fields and the question are stripped of control, bidi and
  zero-width characters, collapsed to one line, capped, JSON-encoded and
  placed in data blocks whose delimiters carry a per-request random
  boundary, with instructions to treat block content as evidence only.
- The knowledge-base query is built from rule IDs and names only, so
  untrusted text no longer steers retrieval.

Output
- /prioritise and /threat-simulation validate the model's JSON against
  the evidence. Items citing rules or resources outside it are dropped
  and counted; malformed output returns 502 instead of raw text.
- /prioritise with no findings returns an empty list without a model call.

The dashboard no longer sends findings; it lets the server read the scan.

Closes OWASP#357

Signed-off-by: parthrohit22 <parthrohit60@gmail.com>
Comment thread api/services/ai_guard.py Fixed
@parthrohit22
parthrohit22 requested a review from TFT444 September 25, 2026 23:25
- Bandit B613 (trojan source): the bidi/zero-width ranges in
  _CONTROL_CHARS, the ellipsis, and the test payloads were written as
  literal characters. They are now \uXXXX escapes, so every changed file
  is pure ASCII and nothing invisible sits in the source.
- CodeQL py/polynomial-redos: the code-fence regex could backtrack
  polynomially on hostile model output ("```" followed by many spaces).
  Fence stripping is now plain string handling, with a regression test
  that a 200k-character hostile fence is rejected in under a second.

Signed-off-by: parthrohit22 <parthrohit60@gmail.com>
@parthrohit22 parthrohit22 self-assigned this Sep 25, 2026

@TFT444 TFT444 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. Solid security hardening closing #357. Evidence loaded server-side, fenced prompt boundaries with random hex tokens prevent injection escaping, output validated against scan evidence. 1413 tests pass, all CI green. Three minor non-blocking notes for follow-up: broad except Exception on line 174 of ai_guard.py, no graceful fallback if DATABASE_URL is missing (KeyError vs 503), and one empty GitHub Advanced Security review comment worth checking.

@ritiksah141 ritiksah141 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these are
Nits (non-blocking, none worth holding the PR)

  1. Threat-simulation can return 200 with stages: [] if every stage cited only invented rules; discarded_items is the only hint. A deliberate choice, better than a 502.
  2. finding_record doesn't normalize severity going into the prompt (prompt-only echo of DB values; the validator normalizes on output). Benign.
  3. The injection corpus exercises summary/ask/insights prompts; prioritise/threat-sim share the same _findings_block builder so coverage is effectively shared.

But approving it as safe to merge.

@TFT444
TFT444 merged commit df98060 into OWASP:dev Sep 27, 2026
21 checks passed
parthrohit22 added a commit to parthrohit22/openshield that referenced this pull request Sep 27, 2026
Picks up OWASP#344 (branch rulesets, post-merge CI) and OWASP#359 (evidence-grounded
AI endpoints). Only CHANGELOG.md conflicted; kept both sides' entries.
OWASP#359's AI routes use get_scan/get_latest_completed_scan/get_findings,
whose signatures this branch leaves unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JMtsuR7tvJTsoq5KueraLf
Signed-off-by: parthrohit22 <parthrohit60@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

security(ai): harden AI endpoints against prompt injection and unvalidated model output

4 participants