Security Finding
Severity: medium
Type: unsafe-pattern (resource-exhaustion DoS / unhandled panics on untrusted input)
pkg/intake/intake.go ExtractPDF hands the entire uploaded PDF to pdf.Reader.GetPlainText() and then io.ReadAll(plain) with no output bound:
plain, err := r.GetPlainText()
...
b, err := io.ReadAll(plain)
Two problems with untrusted input on the authenticated POST /api/intake route:
-
Decompression bomb → OOM. ledongthuc/pdf.(*Reader).GetPlainText (page.go:64) materializes the text of every page into one bytes.Buffer before returning. PDF content streams are FlateDecode-compressed; a crafted PDF within the 25 MB upload cap (DefaultMaxUploadMB) can expand ~1000:1 into multi-GB allocations, OOM-killing the process. The caller only ever keeps MaxReturnedRunes (20 000) runes, so materializing everything is pure downside.
-
Panics on malformed PDFs. pdf.NewReader, NumPage, and Page panic liberally on malformed input (read.go:734–785). ExtractPDF has no recover, so each malformed upload aborts the connection via net/http's per-request recover and spams the log. (The per-page Page.GetPlainText does recover internally; the reader-level entry points do not.)
Impact
Any signed-in user can upload a small crafted PDF and drive the dibs process out of memory (full service DoS), or repeatedly abort request handling with malformed PDFs.
Recommendation
In ExtractPDF:
- Iterate pages ourselves (like
ExtractDOCX already does for XML): per-page p.GetPlainText(fonts), accumulate via the existing writeLimitedRunes helper, and stop as soon as extractRunesLimit (20 001) runes are collected — bounding accumulation instead of materializing the whole document.
- Wrap the function in a
defer recover() that converts parser panics into a normal "could not extract text" error.
Residual (harder) gap: a bomb concentrated in a single page still materializes that one page's text inside the library before the cap applies. A complete fix needs a streaming-capped extractor or subprocess isolation; the above removes the whole-document materialization and the panic crash-noise, which covers the practical multi-page/multi-stream bomb shapes.
Filed by sec-check agent (ACMM L4/L5 — hold-gated mode)
🐝 Hive Agent: security | Instance: hosted-available-oke-11-placeholder-r05x | SHA: unknown
— hive: agent=sec-check backend=copilot model=claude-fable-5 copilot=1.0.88
Security Finding
Severity: medium
Type: unsafe-pattern (resource-exhaustion DoS / unhandled panics on untrusted input)
pkg/intake/intake.goExtractPDFhands the entire uploaded PDF topdf.Reader.GetPlainText()and thenio.ReadAll(plain)with no output bound:Two problems with untrusted input on the authenticated
POST /api/intakeroute:Decompression bomb → OOM.
ledongthuc/pdf.(*Reader).GetPlainText(page.go:64) materializes the text of every page into onebytes.Bufferbefore returning. PDF content streams are FlateDecode-compressed; a crafted PDF within the 25 MB upload cap (DefaultMaxUploadMB) can expand ~1000:1 into multi-GB allocations, OOM-killing the process. The caller only ever keepsMaxReturnedRunes(20 000) runes, so materializing everything is pure downside.Panics on malformed PDFs.
pdf.NewReader,NumPage, andPagepanic liberally on malformed input (read.go:734–785).ExtractPDFhas norecover, so each malformed upload aborts the connection via net/http's per-request recover and spams the log. (The per-pagePage.GetPlainTextdoes recover internally; the reader-level entry points do not.)Impact
Any signed-in user can upload a small crafted PDF and drive the dibs process out of memory (full service DoS), or repeatedly abort request handling with malformed PDFs.
Recommendation
In
ExtractPDF:ExtractDOCXalready does for XML): per-pagep.GetPlainText(fonts), accumulate via the existingwriteLimitedRuneshelper, and stop as soon asextractRunesLimit(20 001) runes are collected — bounding accumulation instead of materializing the whole document.defer recover()that converts parser panics into a normal "could not extract text" error.Residual (harder) gap: a bomb concentrated in a single page still materializes that one page's text inside the library before the cap applies. A complete fix needs a streaming-capped extractor or subprocess isolation; the above removes the whole-document materialization and the panic crash-noise, which covers the practical multi-page/multi-stream bomb shapes.
Filed by sec-check agent (ACMM L4/L5 — hold-gated mode)
🐝 Hive Agent:
security| Instance:hosted-available-oke-11-placeholder-r05x| SHA:unknown— hive: agent=sec-check backend=copilot model=claude-fable-5 copilot=1.0.88