Skip to content

Compute the clutter filter in cache-resident blocks - #70

Merged
Purple10101 merged 4 commits into
mainfrom
20260915-wienerhopf-blocks
Sep 20, 2026
Merged

Purple10101 merged 4 commits into
mainfrom
20260915-wienerhopf-blocks

Conversation

@Purple10101

Copy link
Copy Markdown

Summary

The Wiener-Hopf clutter filter ran seven transforms the length of the whole CPI: four of 1,000,000 points for the correlations and three of 1,016,064 for the convolution. Those arrays are 16 MB against the Pi 5's 2 MB L3, so every pass went to DRAM. blah2 is memory-bandwidth-bound, measured at roughly 75% of the bus, so those transforms were paying in the one currency the machine is short of.

Both halves computed the right answer far more expensively than they needed to. The four correlation transforms existed to produce 410 autocorrelation lags and 410 cross-correlation lags, and the inverses discarded 99.96% of what they computed. The convolution applied a 410-tap filter using a transform the length of the CPI, so the tap array held 410 non-zero values in 1,016,064 points.

Overlap-save gives both, exactly, with every transform inside L2.

Measured

In situ on owl-ded9, reading the stage time blah2 reports for itself, against the code being replaced:

clutter_filter
before 491.7 ms
after 172.6 ms
2.85x

Live across the fleet on this branch, at ±300 Hz:

node clutter_filter duty cycle CPI rate
jonathan-node-1 176.9 ms 97.9% 2.0000 CPI/s
fairforest-b 168.2 ms 96.8% 2.0000 CPI/s
owl-ded9 173.1 ms 100.0% 2.0007 CPI/s

All three now hold real time, including the 2 GB boards, which previously could not.

Soak

Soaked on three nodes from 10:34 UTC 2026-09-19, sampled every 5 minutes.

node window restarts RSS slope RSS now CPIs chol failures Hermitian warnings
jonathan-node-1 23.3 h 0 -0.000 MB/h 146.0 MB 168,173 0 0
fairforest-b 23.3 h 0 -0.005 MB/h 142.1 MB 168,160 0 0
owl-ded9 12.3 h 0* +0.000 MB/h 147.2 MB 90,227 0 0

All three held exactly 2.0000 CPI/s for the whole window, which is real time at a
0.5 s CPI, so every node stayed data-limited rather than compute-limited.

Hourly RSS means are flat, not merely flat by linear fit. jonathan-node-1 ranged
143.9-146.3 MB across 24 hourly buckets and fairforest-b 142.6-143.7 MB.
owl-ded9 settled over its first two hours and then sat pinned at 147.2 MB for
10.3 h with a slope of exactly +0.0000 MB/h.

For scale, the tracker leak retired in #69 ran at 1.5-2.0 MB/h, which would have
added 35-47 MB over this window.

* owl-ded9's window is shorter and its container restarted once because I
deliberately used it for an unrelated doppler-span experiment partway through.
That excursion is excluded from its figures above; the restart is mine, not a
fault of this branch.

This is not bit-identical, and that was quantified rather than assumed

The same terms are summed in a different order, and that perturbation reaches the Cholesky solve, which amplifies it by cond(A).

cond(A) is 2.5e3 to 1.5e4, measured by rebuilding the reference from the live nodes' own reported spectra. That estimate is exact rather than approximate, because the autocorrelation is the inverse transform of the periodogram and phase never enters A.

Against the shipped filter, through the real ambiguity processor:

  • cancellation identical to 1.8e-15 dB
  • the delay-Doppler map differs by 2.5e-10 absolute, 198 dB below the rms map cell a CFAR threshold sits on
  • an injected target lands in the same delay and Doppler bin

A synthetic sweep puts the onset of anything that could matter at cond(A) near 1e13, nine orders above what the fleet sees.

Constants

Both were swept on a live node with the real binary under the real pipeline, because benchmarks running alongside blah2 contend for the memory bus differently from blah2 contending with its own second stage, and the optimum had already moved once between idle and contended.

block length clutter_filter, 1 thread
1024 195.2 ms
2048 172.6 ms
4096 184.6 ms
8192 186.1 ms
16384 209.5 ms

2048 is a clean minimum rather than a plateau, and one thread beats two there by 31% (172.6 against 226.8).

kClutterFftLength is removed with this change: padding a CPI-length transform is not something this filter does any more. kFrontStageThreads is removed too, because it stopped governing anything once the filter began planning its own transforms, and a constant that names a stage and controls nothing is worse than no constant.

Tests

TestClutterFft gained multi-block geometries. Every case it had was short enough to run as a single block, so partial-sum accumulation was never exercised, which is exactly where this formulation goes wrong.

It also now captures cerr around process() and requires silence. That caught a real regression: the block form multiplies two different sequences, so nothing forces the zero-lag autocorrelation to be exactly real, and the roundoff residue produced 173 armadillo Hermitian warnings in five minutes on jonathan-node-1 against zero on main. The residue only trips armadillo's tolerance past roughly 300 blocks, so a shipped-geometry case at 1,000,000 samples and 611 blocks was added. That is the first test in this file to run at the size the radar actually uses.

Risk

The filter is the largest stage in the pipeline and this rewrites it. The mitigations are the direct-convolution oracle in TestClutterFft, the quantified output comparison above, and the soak table.

🤖 Generated with Claude Code

Purple10101 and others added 4 commits September 19, 2026 11:21
The filter ran seven transforms the length of the whole CPI: four of
1,000,000 points for the correlations and three of 1,016,064 for the
convolution. Those arrays are 16 MB and the Pi 5's L3 is 2 MB, so every
pass went to DRAM. blah2 is memory-bandwidth-bound, measured at roughly
75% of the bus: an idle board gives an external probe 11.1 GB/s against
2.8 with blah2 running. So the transforms were paying in the one currency
the machine is short of.

Both halves computed the same quantities far more expensively than they
needed to. The four correlation transforms existed to produce nBins
autocorrelation lags and nBins cross-correlation lags, 820 numbers, and
the inverses discarded 99.96% of what they computed. The convolution
applied a filter of only nBins taps using a transform the length of the
CPI, so filtW held 410 non-zero values in 1,016,064 points.

Overlap-save gives both, exactly, with every transform inside L2. At the
shipped geometry the filter goes from 337 ms to 171 ms, saving ~167 ms
per CPI at 1.96x, measured against this code on a Pi 5 with the image's
NEON FFTW.

Not bit-identical: the same terms are summed in a different order, and
that perturbation reaches the Cholesky solve, which amplifies by cond(A).
Quantified rather than assumed. cond(A) is 2.5e3 to 1.5e4, measured by
rebuilding the reference from the live nodes' own spectra, an estimate
that is exact because the autocorrelation is the inverse transform of the
periodogram and phase never enters A. Against the shipped filter, through
the real ambiguity processor: cancellation identical to 1.8e-15 dB, the
delay-Doppler map differs by 2.5e-10 absolute, 198 dB below the rms map
cell a CFAR threshold sits on, and an injected target lands in the same
delay and Doppler bin. A synthetic sweep puts the onset of anything that
could matter at cond(A) near 1e13, nine orders above what the fleet sees.

The blocks are 2048 points, too small to pay for thread sync, so the
filter now plans at one thread while the code it replaces wanted two.
That also stops the planner thread count changing its output at all,
which TestStageOrder now reports as exactly zero rather than 2.4e-16.

kClutterFftLength goes with it. Padding a CPI-length transform is not a
thing this filter does any more.

TestClutterFft grew multi-block geometries. Every case it had was short
enough to run as a single block, so partial-sum accumulation was never
exercised, and that is exactly where this formulation goes wrong: a
correlation numerator taking the full window on both sides, or a
normalisation by the CPI length rather than the block length, both
survive a single-block test and are both badly wrong. Mutation tested.
The one mutation the numerical test cannot catch is removing the modulo
on the reference window, because that guards memory rather than
arithmetic; ASAN catches it as a heap-buffer-overflow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Zero lag is sum |x[n]|^2, real by definition, and it lands on the
diagonal of A, which has to be real for the matrix to be Hermitian and
for arma::chol to accept it without complaint.

The single-transform code got that for free. It formed X * conj(X),
whose imaginary part is exactly zero in IEEE arithmetic, so the array
was exactly real and so was its transform at index 0. The block form
multiplies two different sequences, one windowed and one padded to the
hop, so nothing forces the cancellation and a roundoff residue survives.

On jonathan-node-1 that was 173 armadillo warnings in five minutes
against zero on main. It was cosmetic, chol still succeeded and the
filter still ran, but it is a regression against main and it buries
anything real in the log.

Discarding the residue restores an exact property rather than
approximating one.

The test could not see it. The oracle solves with arma::solve, which
does not care about Hermitian-ness, and chol only warns rather than
failing, so the return value cannot see it either. It now captures cerr
around process() and requires silence. That alone was still not enough:
the residue only trips armadillo's tolerance past roughly 300 blocks and
every geometry in the suite was small enough to stay under it, so a
shipped-geometry case was added at 1,000,000 samples and 611 blocks.
That case is the first test in this file to run at the size the radar
actually uses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both constants were picked from benchmarks running alongside blah2, which
contends for the memory bus differently from blah2 contending with its own
second stage. Since the bus is what this workload is short of, and the
optimum had already moved once between idle and contended (8192 to 2048),
neither number was safe to assume.

Swept on a live node with the real binary under the real pipeline, reading
the stage time blah2 reports for itself. Both hold:

  block   clutter_filter (ms, 1 thread)      block   1 thr   2 thr
   1024       195.2                           2048   172.6   226.8
   2048       172.6  <- minimum               4096   184.6   195.1
   4096       184.6
   8192       186.1
  16384       209.5

2048 is a clean minimum rather than a plateau, and one thread beats two by
31% there. Against 491.7 ms for the code being replaced, that is 2.85x.

Comments only. No behaviour change: this records evidence for numbers that
were already in the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
kFrontStageThreads no longer governed anything. The front stage's only FFTW
consumer is the clutter filter, and since it started planning its own
transforms at kClutterBlockThreads the value set before its constructor was
overwritten before it ever reached a plan.

Removed rather than left unused, because a constant that names a stage and
controls nothing is worse than no constant: the next person to tune threads
would reasonably change it and see no effect.

What is left now says what it does. kBackStageThreads governs the ambiguity
processor and the spectrum analyser, which are the only plan sites outside the
filter and are both genuinely in the back stage. The comment claiming the
spectrum analyser runs in the front stage was wrong; it runs in the back, as
the loop at the bottom of blah2.cpp has always shown.

WienerHopf now restores the planner to kBackStageThreads rather than to the
count it was handed, so a plan site added later without an explicit count
cannot silently inherit the single thread that suits 2048-point blocks and
nothing else.

No behaviour change: every plan is still built with the count it was built
with before. testClutterFft and testStageOrder pass on aarch64, and the latter
still reports the planner thread count moving the output by exactly zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Purple10101
Purple10101 merged commit 8d2ac7d into main Sep 20, 2026
4 checks passed
@Purple10101
Purple10101 deleted the 20260915-wienerhopf-blocks branch September 20, 2026 10:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant