Skip to content

Memory Benchmark Matrix

luojiyin edited this page Sep 26, 2026 · 5 revisions

Memory Benchmark Matrix

This page records peak RSS across file sizes and worker counts.

Tracking issue: #183

Related adaptive concurrency work: #78

Test conditions

  • Platform: Linux x64
  • Node.js: v24.15.0
  • Host memory: 29 GiB RAM and 8 GiB swap
  • Tool: GNU /usr/bin/time -v
  • Benchmark: scripts/benchmark-memory.mjs
  • Fixture shape: one title and one long paragraph
  • Modes: lint and --fix
  • Worker counts: 1, 2, and 4
  • File sizes: 1, 3, 5, and 10 MiB
  • Adaptive mode: disabled
  • Size filtering: disabled

Piscina uses worker threads. GNU time therefore measures the total Node process RSS.

Test process

The first plan used eight files and five runs per cell.

The 1 MiB and one-worker cell produced these results:

  • Median peak RSS: 852.7 MiB
  • Maximum peak RSS: 887.2 MiB
  • Median wall time: 32.44 seconds

The initial result made the full plan unsafe. The test stopped before larger cells ran.

The exploratory matrix used one file per worker. This isolated the active concurrent working set.

Each exploratory cell used a fresh CLI process and one run.

Command shape:

npm run benchmark:memory -- \
  --files <worker-count> \
  --bytes-per-file <bytes> \
  --threads <worker-count> \
  --runs 1

The fix matrix added --fix.

Lint peak RSS

File size 1 worker 2 workers 4 workers
1 MiB 745 MiB 1,506 MiB 2,674 MiB
3 MiB 1,603 MiB 2,951 MiB 5,118 MiB
5 MiB 1,833 MiB 3,875 MiB 6,327 MiB
10 MiB 1,967 MiB 3,487 MiB 7,104 MiB

Fix peak RSS

File size 1 worker 2 workers 4 workers
1 MiB 696 MiB 1,644 MiB 2,741 MiB
3 MiB 1,773 MiB 2,616 MiB 5,351 MiB
5 MiB 1,813 MiB 3,969 MiB 5,925 MiB
10 MiB 1,843 MiB 3,677 MiB 7,249 MiB

Findings

  • Peak RSS grows approximately with worker count.
  • The worker relationship is noisy but not severely superlinear.
  • File-size growth is not linear.
  • Single-worker RSS starts to level near 3 to 5 MiB.
  • Fix mode does not consistently exceed lint mode.
  • Four 10 MiB files need approximately 7.1 GiB peak RSS.

The current adaptive policy prevents the largest measured four-worker peaks.

max < 1 MiB       -> availableParallelism()
1 MiB <= max < 5  -> at most 2 workers
max >= 5 MiB      -> 1 worker

Measured examples:

  • 3 MiB with two workers used 2.6 to 3.0 GiB.
  • 5 MiB with one worker used approximately 1.8 GiB.
  • 10 MiB with one worker used approximately 1.9 GiB.

Important boundary

The policy has a sharp boundary below 1 MiB.

A 0.99 MiB file can still use every available CPU.

The next matrix should test 256 KiB, 512 KiB, and 900 KiB.

It should test 1, 2, 4, 8, and 16 workers.

Limitations

  • Each exploratory cell has one run.
  • The data shows trends, not precise thresholds.
  • The fixture is a synthetic long paragraph.
  • Other Markdown shapes can produce different memory use.
  • The test covers one Linux host.
  • The first five-run cell used eight files.
  • Its result is not directly comparable with the exploratory cells.

Run important boundary cells five times before changing policy thresholds.

Core 2.5.0 comparison

This comparison measures the CLI before and after the core 2.5.0 update.

Test conditions

  • Before: CLI commit f263ac8, @lint-md/core@2.3.1
  • After: CLI master commit 2ad0bbe, @lint-md/core@2.5.0
  • Node.js: v24.15.0
  • Platform: Linux x64
  • Input: 8 generated Markdown files, 64 KiB per file
  • Workers: 2
  • Runs: 5 per version
  • Command: npm run benchmark:memory -- --files 8 --bytes-per-file 65536 --threads 2 --runs 5

Results

Metric Core 2.3.1 Core 2.5.0 Change
Average wall time 1.684 s 0.610 s -63.8%
Average peak RSS 822,962 KiB 360,906 KiB -56.1%

Core 2.3.1 wall times were 1.69, 1.65, 1.69, 1.73, and 1.66 seconds.

Core 2.5.0 wall times were 0.59, 0.60, 0.61, 0.62, and 0.63 seconds.

Core 2.3.1 peak RSS values were 806,180, 809,060, 818,748, 859,624, and 821,196 KiB.

Core 2.5.0 peak RSS values were 327,772, 344,796, 350,140, 392,448, and 389,376 KiB.

The comparison shows lower wall time and lower peak RSS after the update. The result uses one host and one synthetic input shape. Run more input sizes and worker counts before setting general performance limits.

Core 2.5.0 to current master comparison

This comparison measures the core package at tag v2.5.0 and after the following performance changes:

  • #286 shared text classification.
  • #288 selective ellipsis mapping.

Test conditions

  • Before: core tag v2.5.0, commit 72984e7.
  • After: core master, commit 05bd8d0.
  • Node.js: v24.15.0.
  • Platform: Linux x64.
  • Input: 1 MiB per case.
  • Runs: 7 measured runs per case.
  • Warmup: 2 runs per case.
  • Benchmark: scripts/benchmark-memory.mjs.
  • Shapes: low-match-density and high-match-density.
  • Cases: parser-only, parse-traverse, text-scanner-rules, all-rules, and fix-mode.

The four benchmark commands ran in separate processes at the same time. The values show trends. They do not provide noise-free microbenchmarks.

Average wall time: low density

Layer v2.5.0 Current master Change
parser-only 35.056 ms 37.797 ms +7.8%
parse-traverse 38.416 ms 38.108 ms -0.8%
core infrastructure 3.360 ms 0.311 ms -90.7%
text rules 76.917 ms 39.210 ms -49.0%
remaining rules 74.214 ms 41.065 ms -44.7%
all-rules 112.630 ms 79.173 ms -29.7%
fix amplification -8.983 ms 0.147 ms noisy

Average wall time: high density

Layer v2.5.0 Current master Change
parser-only 173.130 ms 176.338 ms +1.9%
parse-traverse 172.901 ms 175.428 ms +1.5%
core infrastructure -0.228 ms -0.911 ms noise
text rules 319.700 ms 219.402 ms -31.4%
remaining rules 252.173 ms 194.099 ms -23.0%
all-rules 425.074 ms 369.527 ms -13.1%
fix amplification 712.607 ms 655.668 ms -8.0%

Median wall time: high density

Layer v2.5.0 Current master Change
parser-only 174.984 ms 180.522 ms +3.2%
parse-traverse 171.924 ms 175.579 ms +2.1%
text rules 302.476 ms 212.373 ms -29.8%
remaining rules 252.520 ms 196.179 ms -22.3%
all-rules 424.444 ms 371.757 ms -12.4%
fix amplification 733.728 ms 644.808 ms -12.1%

Findings

  • Current master lowers high-density all-rules time by approximately 13.1%.
  • Current master lowers high-density text-rule time by approximately 31.4%.
  • Parser time changes by less than 2% on average in high-density input.
  • Parse-traverse adds little time beyond parser-only in both versions.
  • Core infrastructure does not appear to be a stable hotspot.
  • Fix amplification remains large in high-density input.
  • Earlier fix profiling found that the second runLint dominates fix amplification.
  • Earlier profiling found that applyFix uses approximately 3.2% of total fix time.

The results support stopping small core rule optimizations. Future parser work needs a new parser profiler.

Core 2.5.2 to 2.5.3 comparison

This comparison measures the current CLI lock against core 2.5.3.

Test conditions

  • Test date: 2026-09-26.
  • Before: core v2.5.2, commit 829dd8f.
  • Before parser: @lint-md/parser@0.2.1 from the CLI lock file.
  • After: core v2.5.3, commit eecb933.
  • After parser: @lint-md/parser@0.3.1.
  • Node.js: v24.15.0.
  • Platform: Linux x64 under WSL2.
  • Logical CPUs: 16.
  • Harness: the core 2.5.3 benchmark scripts for both versions.
  • Main cases: 7 measured runs with one warmup per child process.
  • Long-code case: 3 measured runs without selector warmup.
  • The test ran each version in sequence.
  • The scenario order changed between version pairs.

End-to-end core results

The table shows average wall time.

Shape and case Input Core 2.5.2 Core 2.5.3 Change
mixed-markdown, parser-only 256 KiB 1,267.2 ms 481.9 ms -62.0%
mixed-markdown, all-rules 256 KiB 1,333.5 ms 494.3 ms -62.9%
high-match-density, all-rules 64 KiB 46.6 ms 40.6 ms -12.9%
overlapping-fixes, fix-mode 64 KiB 37.8 ms 32.2 ms -14.7%
large-code-block, no-long-code selector 10 MiB 6.16 ms 3.05 ms -50.4%

The report counts matched for both versions. The fix case produced 5,042 reports in two rounds. Both versions reached stable convergence.

Peak RSS

The table shows the average process peak RSS.

Shape and case Core 2.5.2 Core 2.5.3 Change
mixed-markdown, parser-only 300.9 MiB 337.3 MiB +12.1%
mixed-markdown, all-rules 304.1 MiB 336.5 MiB +10.7%
high-match-density, all-rules 105.4 MiB 104.0 MiB -1.3%
overlapping-fixes, fix-mode 81.0 MiB 80.7 MiB -0.4%

The mixed-markdown result shows a speed and memory tradeoff. The other measured shapes keep peak RSS effectively unchanged.

SourceCode index benchmark

Core 2.5.3 builds the line-start index on its first use.

Input Core 2.5.2 create Core 2.5.3 create First position change
64 KiB 0.292066 ms 0.000115 ms -1.0%
256 KiB 1.902948 ms 0.000075 ms +2.6%
1 MiB 7.069443 ms 0.000050 ms -1.0%

The change removes index work when a run does not request a position. The first position lookup still pays the index build cost.

Space-around-link benchmark

This benchmark uses many sibling links in one paragraph.

Links Siblings Core 2.5.2 Core 2.5.3 Change
1,000 2,001 2.243 ms 2.372 ms +5.8%
5,000 10,001 17.616 ms 8.308 ms -52.8%
10,000 20,001 60.122 ms 18.879 ms -68.6%

The sibling index improves scaling for large sibling lists. The 1,000-link result remains within microbenchmark noise.

Findings

  • Parser 0.3.1 causes most of the mixed-markdown speed gain.
  • High diagnostic density improves by approximately 13%.
  • Fix mode improves by approximately 15% in the measured shape.
  • The no-long-code selector improves by approximately 50%.
  • Large sibling lists receive the largest rule-specific gain.
  • Mixed Markdown peak RSS increases by approximately 11% to 12%.
  • The performance change depends on the Markdown shape.

Core 2.5.3 is faster in every end-to-end case in this comparison. Monitor memory when processing large, complex Markdown documents.