Skip to content

[sec-check] intake: ExtractPDF materializes unbounded decompressed PDF text (OOM bomb) and lacks panic recovery #141

Description

@hivecommons-hive

Security Finding

Severity: medium
Type: unsafe-pattern (resource-exhaustion DoS / unhandled panics on untrusted input)

pkg/intake/intake.go ExtractPDF hands the entire uploaded PDF to pdf.Reader.GetPlainText() and then io.ReadAll(plain) with no output bound:

plain, err := r.GetPlainText()
...
b, err := io.ReadAll(plain)

Two problems with untrusted input on the authenticated POST /api/intake route:

  1. Decompression bomb → OOM. ledongthuc/pdf.(*Reader).GetPlainText (page.go:64) materializes the text of every page into one bytes.Buffer before returning. PDF content streams are FlateDecode-compressed; a crafted PDF within the 25 MB upload cap (DefaultMaxUploadMB) can expand ~1000:1 into multi-GB allocations, OOM-killing the process. The caller only ever keeps MaxReturnedRunes (20 000) runes, so materializing everything is pure downside.

  2. Panics on malformed PDFs. pdf.NewReader, NumPage, and Page panic liberally on malformed input (read.go:734–785). ExtractPDF has no recover, so each malformed upload aborts the connection via net/http's per-request recover and spams the log. (The per-page Page.GetPlainText does recover internally; the reader-level entry points do not.)

Impact

Any signed-in user can upload a small crafted PDF and drive the dibs process out of memory (full service DoS), or repeatedly abort request handling with malformed PDFs.

Recommendation

In ExtractPDF:

  • Iterate pages ourselves (like ExtractDOCX already does for XML): per-page p.GetPlainText(fonts), accumulate via the existing writeLimitedRunes helper, and stop as soon as extractRunesLimit (20 001) runes are collected — bounding accumulation instead of materializing the whole document.
  • Wrap the function in a defer recover() that converts parser panics into a normal "could not extract text" error.

Residual (harder) gap: a bomb concentrated in a single page still materializes that one page's text inside the library before the cap applies. A complete fix needs a streaming-capped extractor or subprocess isolation; the above removes the whole-document materialization and the panic crash-noise, which covers the practical multi-page/multi-stream bomb shapes.


Filed by sec-check agent (ACMM L4/L5 — hold-gated mode)

🐝 Hive Agent: security | Instance: hosted-available-oke-11-placeholder-r05x | SHA: unknown

— hive: agent=sec-check backend=copilot model=claude-fable-5 copilot=1.0.88

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent/securityCreated by Hive for agent-filed issue provenancehive/hosted-available-oke-11-placeholder-r05xCreated by Hive for agent-filed issue provenancesecurityCreated by Hive for agent-filed issue provenance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions