fix: bound PDF text extraction and recover parser panics in intake [sec-check] - #142
Conversation
…in intake ExtractPDF handed the whole document to Reader.GetPlainText, which materializes every page's decompressed text in one buffer — a crafted FlateDecode-heavy PDF within the 25MB upload cap could expand into multi-GB allocations (OOM DoS on the authenticated /api/intake route). The reader-level pdf entry points also panic on malformed input with no recover in the caller. Extract page by page instead, accumulating through writeLimitedRunes and stopping at extractRunesLimit (the caller keeps at most MaxReturnedRunes anyway), and convert parser panics into ordinary errors. Refs #141 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: sec-check <sec-check@hive.kubestellar.io>
|
Important Held for human review by the hive's ACMM level gate. This PR was opened by the "sec-check" agent while Hive policy required a human checkpoint for that agent. Non-outreach agents are held at ACMM L3–L5; the Hive will automatically remove the |
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
…ranches (#145) pkg/intake was 68.8% covered with the upload type-confusion gate (sniffMatchesDocument 50%, sniffMatchesAudio 44%) and the RTF hex-escape decoder (parseHexByte/hexVal 0%) untested — directly behind the /api/intake surface hardened by #142. Adds table tests for both sniff gates (every extension arm, accept+reject), StripRTF \'xx escapes, cleanText whitespace branches, MaxUploadBytes env override, Error.Error, and HandleUpload branches (success, unsupported type, oversize, and non-multipart). pkg/intake: 68.8% -> 88.3%. Signed-off-by: hive-quality <sec-check@hive.kubestellar.io> Co-authored-by: hive-quality <sec-check@hive.kubestellar.io> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Security Fix
Files/functions claimed:
pkg/intake/intake.go(ExtractPDF),pkg/intake/intake_test.go(new malformed/multi-page PDF tests).ExtractPDFhanded the whole upload topdf.Reader.GetPlainText(), which materializes every page's decompressed text into a single buffer before returning. PDF content streams are FlateDecode-compressed, so a crafted PDF within the 25 MB upload cap can expand ~1000:1 into multi-GB allocations — an OOM DoS reachable by any signed-in user viaPOST /api/intake. The reader-level parser entry points (NewReader/NumPage/Page) also panic on malformed input, andExtractPDFhad no recover.This change:
writeLimitedRuneshelper and stopping atextractRunesLimit(20 001 runes) — the caller only ever returnsMaxReturnedRunesanyway, so whole-document materialization was pure downside;Refs #141 (partial fix — a bomb concentrated in a single page still materializes that page's text inside the library before the cap applies; a complete fix needs a streaming-capped extractor or subprocess isolation, tracked in the issue)
Filed by sec-check agent (ACMM L4/L5 — hold-gated mode). Hold-gated: human review required.
— hive: agent=sec-check backend=copilot model=claude-fable-5 copilot=1.0.88