Skip to content

Users/pgo/performance issues - #6

Merged
sahebansari merged 11 commits into
sahebansari:masterfrom
pgourlain:users/pgo/performance-issues
Sep 27, 2026
Merged

sahebansari merged 11 commits into
sahebansari:masterfrom
pgourlain:users/pgo/performance-issues

Conversation

@pgourlain

Copy link
Copy Markdown
Contributor

Performance work across the library

This PR makes TerraPDF 2–30× faster in realistic workloads and cuts allocations by 60–95%. Output is byte-identical except for three intended changes, listed at the end. It adds a benchmark suite, a throughput harness and verification tools, so the gains can be measured and checked.

Throughput in a 2-CPU / 1 GB Docker container (benchmarks/TerraPDF.Throughput):

  • Invoice (1 page): 356 → 10,200 pages/s
  • Annual report (19 pages): 3,100 → 25,200 pages/s
  • With a different logo in every document (image cache cold): about 670 and 8,800 pages/s.
Area What changed Gain (BenchmarkDotNet, baseline → now)
Text layout Word width, font and colour computed once; wrapped lines memoised per publish; no second layout pass without page numbers; one Tj per run of words 500 pages: 422 → 59 ms, 1.19 GB → 89 MB
Tables Row heights memoised; page slices only visit their own rows; ~70% fewer allocations per cell 10,000 rows: 283 → 82 ms, 645 → 79 MB
Images PNG decoded at save time, once per distinct image; RGB/palette PNGs embedded without decoding; faster decoder; converted images cached across documents 40× same PNG: 94 → 0.13 ms; DecodePng 8.1 → 2.4 MB
Custom fonts Fast path when no Devanagari shaping is needed; subsetting without table copies Lato document: 15.1 → 6.9 ms; subset allocations −68%
QR codes Flat matrix, masks applied in place, penalties scored bit-parallel Generation 2–6× faster; 100-QR document 25.0 → 10.7 ms
Vector canvas Fixed-point number formatting; parallel page compression for large documents Dense canvas: 26.2 → 5.0 ms
Encryption / output Content streams compressed straight from the buffer; less data to compress and encrypt 20 pages AES-256: 15.5 → 2.7 ms

Intended output changes

  • Text runs: consecutive words in the same style share one Tj. PDFs are up to 14% smaller; glyphs sit within 0.01 pt of before.
  • PNG passthrough: RGB and palette PNGs are embedded with /Predictor 15 instead of being decoded and recompressed.
  • Font widths (fix): built-in font widths now match the Adobe AFM metrics; 25 entries were wrong. Characters that can't be encoded are measured as the ? drawn in their place. Text containing these characters may wrap differently.

Verification

  • Tests: 602 tests on net8.0, net9.0 and net10.0, including new tests for AFM widths, number formatting, the image cache and parallel compression.
  • Output checks: sample PDFs compared byte-for-byte, visually and by glyph position. PNG round-trips are pixel-exact. The 458-symbol QR reference is unchanged.
  • Tooling: everything needed to re-run these checks is in tools/pdf-compare.

Full analysis and numbers: benchmarks/benchmarks-analysis-2026-09-26.md. How to run: docs/benchmarks.md.

@sahebansari sahebansari left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One of the test cases is failing, which is currently blocking this PR from being merged. Could you please take a look and fix it?

@pgourlain

Copy link
Copy Markdown
Contributor Author

Why ShowIfFalseHeadingIsNotInTableOfContents was changed

What failed

On every target framework (net8.0, net9.0, net10.0), CI failed with:

Assert.Equal() Failure: Values differ
Expected: 2
Actual: 0
at BehaviourTests.ShowIfFalseHeadingIsNotInTableOfContents() BehaviourTests.cs:158
It also fails locally on current master, so it isn't a CI environment problem.

Cause: the test was out of date, the library is fine

The test assumed text was drawn one word at a time, so it looked for (Public) Tj in the content stream. The performance commits (7bba9a4, 12314b9, 85e0bd5) changed text rendering to draw one Tj per line. For this document the output is now:

(1 Public chapter) Tj ← table of contents entry
(Public chapter) Tj ← heading in the body
(Public) Tj never appears, so the count was 0. What the test is really checking still works: the hidden "Secret chapter" heading is drawn nowhere, and the visible heading is drawn twice.

There was a second problem. Assert.DoesNotContain("(Secret) Tj", content) could no longer fail, because the renderer never writes that string any more, not even when the heading is drawn. It would have missed a real regression.

The change

  • // Text is emitted word by word; the TOC page and the body each draw the visible heading once.
  • Assert.DoesNotContain("(Secret) Tj", content);
  • Assert.Equal(2, content.Split("(Public) Tj").Length - 1);
  • // Each line is emitted as one run; the TOC page ("1 Public chapter") and the body each draw
  • // the visible heading once.
  • Assert.DoesNotContain("Secret", content);
  • Assert.Equal(2, content.Split("Public chapter) Tj").Length - 1);
    "Secret" must appear nowhere. This check fails if the hidden heading shows up anywhere, in the table of contents or the body, however the text is split into runs.
    "Public chapter) Tj" must appear exactly twice. It matches both the numbered table-of-contents entry and the body heading.

@sahebansari sahebansari left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for quick fix.

@sahebansari
sahebansari merged commit 094db98 into sahebansari:master Sep 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants