Repository navigation
[Renderer] Add virtual shadow maps with cached pages and scalable many-light shadows #132
Description
Activity
Directional VSM milestone implemented and qualified behind the explicit
BLOOM_VSM=1gate. The issue remains open for dynamic-page rendering, receiver-driven demand, local lights, and tier rollout.Implemented:
- Fixed 3-level virtual address space: 32x32 pages per level, 128x128 interior texels, 2-texel gutters, 132x132 physical pages.
- Deterministic fixed-budget LRU residency with current-frame eviction protection, missing-page CSM fallback, signature invalidation, eight-frame CSM-to-VSM residency fade, and a hard device/config capacity bound.
- Center-first, interleaved level scheduling so an invalidating near level cannot starve mid/far pages.
- Opt-in GPU resources only: R32Uint page table, Depth32Float physical array, bounded dedicated per-page uniforms, and a zero-work/default allocation path.
- Physical page rendering reuses the existing opaque and alpha-cutout shadow pipelines, performs per-page caster culling, preserves gutters, and uses a dedicated uniform buffer so deferred queue writes cannot corrupt CSM matrices.
- Built-in scene and material-ABI receiver sampling. Missing, dirty, over-budget, and dynamically unsafe pages resolve through CSM, never unshadowed.
- Lazy fallback sampling: resident steady-state pixels pay the 4-tap VSM kernel, not both 16-tap CSM and VSM. This fixed the first A/B performance regression before landing.
- Stable-cache fast path eliminates repeated 224-page cache walks and page-table/parameter uploads after residency/fade stabilizes.
- Telemetry now reports active/fallback state, capacity, physical/overhead/total bytes, hit/miss/eviction/denial/invalidation/render counts, and per-level residency.
- Quality captures now emit
virtual-shadow-pages.pngandvirtual-shadow-physical.pngoccupancy views.
Safety contract:
- Default mode remains the canonical CSM shader/layout and allocates no VSM GPU resource or per-pixel branch.
- Continuous skinned/dynamic casters conservatively set
dynamic_fallback: true; the exact live CSM path remains authoritative until dynamic VSM overlays are implemented. - Alpha-tested static casters render through the cutout page pipeline.
Evidence on Apple M1 Max / Metal:
- Sponza static run:
active: true, 224/224 resident, 224 hits, 0 dirty, 0 denied, 0 evictions; level residency 144/64/16. - Sponza VSM-vs-CSM final: SSIM 0.990153, luminance RMSE 0.011009, OKLab mean delta 0.002991, edge delta 0.002022. The visible change is the intended higher-resolution/sharper shadow boundary; the diff gate passes.
- Latest same-machine Sponza main pass: 1.498 ms CSM reference vs 1.302 ms VSM, while steady shadow-pass CPU returned to about 0.040 ms after the cache fast path. Whole-frame GPU values remain scheduler/noise dominated and are not claimed as a hard gate yet.
- Skinned/alpha motion:
dynamic_fallback: true; VSM-gated vs CSM final is effectively identical: SSIM 1.00000, luminance RMSE 0.00003, zero pixels above tolerance. - Quality runner failure in both cases is only the repository-wide missing approved portable baseline, not capture, validation, or image-threshold failure.
cargo test --manifest-path native/shared/Cargo.toml --lib --no-fail-fast: 181 passed, 0 failed, 1 ignored.- Default
golden_lit_primitives_3d: passed. - Focused VSM suite: 15 tests, covering coordinate bounds, LRU and eviction protection, memory bounds, page-table age/fallback, crop/gutter math, fair scheduling, fast-path accounting, debug output, and WGSL parsing for scene/material variants.
Still required before closing #132:
- Receiver/depth-driven page marking and GPU compaction/submission.
- Page-granular dynamic/skinned overlays instead of whole-feature CSM fallback.
- True snapped directional clipmaps independent of the CSM oracle.
- Spot/point-light virtual projections and shared pool arbitration for 100+ submitted shadow lights.
- Bistro/moving-light/small-pool stress qualification, seam/ring/pop gates, quality-tier integration, and high-tier default enablement only after those gates pass.
Receiver-driven demand and dynamic/skinned safety milestone completed and qualified. #132 remains open for true clipmaps, GPU compaction/submission, local-light projections, and tier rollout.
Implemented in this milestone:
- Camera-visible receiver AABBs are projected into every directional level and expanded by a one-page filter/jitter guard.
- Requests are ranked deterministically by receiver coverage and center distance, capped at 144/64/16 pages, and interleaved across levels. Missing or omitted pages continue to sample CSM.
- Receiver-to-page results are cached by receiver-bounds signature and level matrices. The first implementation rebuilt the coverage hash every frame and regressed Sponza shadow CPU from about 0.04 ms to 0.333 ms; that version was rejected. The cached implementation is back to 0.048-0.057 ms in steady static captures.
- Moving/skinned receivers are excluded from the static-demand signature so animation does not rebuild static page ranking each pose.
- Dynamic casters now have conservative light-space page masks with a two-page guard. Masked pages use live CSM and never sample stale static VSM depth.
- Small receiver footprints automatically keep the proven whole-frame CSM fallback because page-mask construction and VSM indirection cannot repay their cost. Larger footprints can retain VSM on unrelated static pages. Unbounded casters preserve whole-frame fallback.
- Dynamic fallback telemetry now reports
whole-frame-csm,full-demand-csm, orper-page-csm, plus page count. - Disabled VSM telemetry now reports actual zero physical capacity, bytes, overhead, and render budget rather than the internal one-slot CPU sentinel.
Evidence on Apple M1 Max / Metal:
- Final static Sponza:
active:true,demand_source:receiver-bounds, 224 resident/requested/hits, 0 misses, 0 denied, 0 evictions, 0 dirty; 144/64/16 pages by level. - Static Sponza main pass: 1.268 ms in the latest run. Shadow CPU 0.0566 ms. Output versus the preceding receiver-cached capture: SSIM 1.00000, luminance RMSE 0.00010, 0.00% above tolerance.
- Skinned/alpha motion: 67 receiver-demand pages, adaptive
whole-frame-csm, shadow CPU 0.0327 ms versus 0.0312 ms CSM reference. Output versus CSM: SSIM 1.00000, luminance RMSE 0.00003, 0.00% above tolerance. - Default
golden_lit_primitives_3d: pass. - Full library suite: 188 passed, 0 failed, 1 ignored.
- Focused VSM suite: 22 passed, including bounded/unique/fair receiver demand, deterministic signatures, guard coverage, offscreen rejection, per-page dynamic masking, and unbounded-caster fallback.
- Quality runner status is fail only because the repository does not contain the approved portable baseline; capture and explicit image comparisons pass.
Local-light audit / required nonbreaking contract:
Bloom currently exposes only unshadowed point lights. There is no spot-light type and no public or native shadow-request flag. Existing content, including the 192-point-light stress case, depends on that behavior. Automatically treating current point lights as shadowed would be a visual compatibility break and an unbounded performance regression, so this milestone does not do that implicitly.
The safe local-light implementation should therefore:
- Add an explicit, default-false shadow request and priority to point lights, plus a first-class spot-light API. Preserve existing
addPointLightbyte-for-byte and route new options through additive FFI entry points. - Extend the virtual address key with projection kind and face. Use 2D spot projection and a documented six-face cube or octahedral point representation; do not overload directional clip levels.
- Keep one physical page allocator for directional, spot, and point owners. Arbitration must be deterministic and visibility/priority aware, with per-light minimum/coarse fallback and a hard global byte/page cap.
- Add per-light page-table metadata outside the existing lighting UBO so unshadowed lights do not grow or slow the canonical loop.
- Qualify two negative controls: the current 192 unshadowed lights remain image/performance equivalent, and 100 explicitly shadow-requesting lights stay bounded with reported admitted/deferred/page cost.
- Only then connect point/spot sampling in the clustered and material paths; missing/deferred pages must resolve to conventional coarse shadow or an explicitly documented no-shadow policy chosen by the caller, never silently change existing lights.
This preserves the no-regression constraint while leaving an independently implementable local-light contract.
Dynamic/skinned directional page-overlay milestone completed and qualified in
6b4c256with evidence/docs in75cdfe0. #132 remains open for true snapped clipmaps, GPU compaction/submission, local-light projections/arbitration, and rollout qualification.Implemented:
- Bounded dynamic/skinned caster AABBs are projected with a two-page PCF/jitter guard; pages nearest each caster core are selected first, including separated multi-caster footprints.
- Up to 4 affected pages and 64 total page draws are rebuilt per frame with both static and current-frame dynamic geometry. Opaque, MASK/cutout, foliage-motion, and skinned pipelines reuse the live CSM geometry, cutout bindings, joint palette, and wind parameters.
- Prior/current affected pages are invalidated before page-table upload. Only successfully rendered pages become sampleable; deferred, over-budget, missing, and dirty pages use current CSM. There is no stale animated VSM depth or unshadowed fallback.
- Small receiver sets (<128 pages) and unbounded casters retain whole-frame CSM.
BLOOM_VSMunset remains the zero-allocation, zero-pass canonical CSM path. - Hard telemetry now exposes overlay footprint, rendered pages, draws, deferred pages, and page/draw budgets.
- Added opt-in
quality-motion --vsm-dynamicoracle and operational documentation indocs/virtual-shadow-maps.md.
Final Apple M1 Max / Metal evidence:
- 224 demanded/resident pages; 151 guarded dynamic pages; 4 overlay pages rendered with 8 draws; 75 demanded guard pages left dirty for CSM; 0 denied and 0 evicted. Total VSM memory remains fixed at 19,951,632 bytes.
- Repeated overlay capture: SSIM 1.000000000, luma RMSE 0.000005030, 0% above 0.02 tolerance.
- Prior per-page-CSM control to overlay: SSIM 0.992135584, RMSE 0.008894581; reviewed differences are the intended shadow/foliage/contact-edge detail, without page rectangles or missing shadow regions.
- VSM-disabled old/new control: SSIM 1.000000000, RMSE 0.000012179, 0% above tolerance, and telemetry remains zero VSM bytes/work.
- Interleaved timing medians: wall 13.873404 ms control vs 13.760610 ms overlay; shadow GPU 4.008879 vs 4.100516 ms (0.091637 ms absolute); bounded shadow CPU 0.054876 vs 0.180561 ms. Four passes/64 draws are hard ceilings and excess work uses already-rendered CSM.
- 331 shared unit tests + headless device test + 59 GPU goldens + 4 render-target tests passed; 24 focused VSM tests passed. Contracts, strict Clippy, file-size ratchet, native release, and wasm Web-feature checks pass.
Evidence:
docs/evidence/issue-132-dynamic-vsm-overlay-v1.mdand.json.Acceptance checklist: only the already-proven fixed memory-budget item is newly checked. Moving-light behavior, independent clipmaps, 100 local lights, VSM-specific automated foliage/skinned corpus coverage, constrained-adapter qualification, and full camera-path seam/pop gates remain deliberately open.
Independent page-snapped directional clipmap milestone completed in
3861584, with exact-revision qualification/docs in9851407.Implemented:
- Replaced VSM reuse of fitted CSM matrices with three independent camera-centered orthographic clip levels. Light-space X/Y origins snap to one virtual-page footprint; sub-page camera motion keeps every matrix and cached address byte-stable.
- Added split-derived coverage guard plus quantized scene-depth pancaking. Receiver demand, physical-page rendering, scene shading, and material shading now use the same clipmap matrices.
- Preserved the independently fitted CSM matrices as live fallback data. Missing, dirty, deferred, or out-of-volume VSM samples use the original CSM projection; no stale or accidentally unshadowed sample is exposed.
- Expanded the sampling ABI from 16 to 208 bytes for three matrices plus policy words, with exact Rust/WGSL size/alignment tests and all scene/material bind layouts updated.
- Added telemetry projection identity:
camera-centered-page-snapped-clipmap. - Kept the entire path opt-in. With
BLOOM_VSMunset, clipmaps are not computed, VSM resources/passes remain zero, and canonical shaders remain uninjected.
Exact Apple M1 Max / Metal evidence:
- Dynamic fixture: 224 demanded/resident pages, 153 guarded dynamic pages, 4 overlay pages / 8 draws, 119 deliberately dirty demand pages using CSM, 0 denials, 0 evictions, and fixed 19,951,824-byte VSM memory.
- Static Sponza: 224 resident pages, zero dirty pages, denials, evictions, or steady-state page renders.
- Three repeat captures: worst SSIM
0.999999940, luma RMSE0.000021062, and 0% above the 0.02 tolerance. - Fitted-control to clipmap: dynamic SSIM
0.994245768/ RMSE0.006658406; static Sponza SSIM0.996063292/ RMSE0.005528215. Full-image and heatmap review showed localized shadow-edge/foliage-detail changes with no page rectangles, missing shadows, or new unshadowed geometry. - VSM-disabled control: SSIM
1.000000000, RMSE0.000013567, 0% above tolerance, and zero VSM bytes/work. - Median versus the preceding fitted-overlay qualification: wall
13.760610 -> 13.025648 ms, shadow CPU0.180561 -> 0.172756 ms, shadow GPU4.100516 -> 4.029285 ms. No measured steady-state regression. - Quick CI lane passed: all-platform FFI/schema parity, strict Clippy, file-size ratchet, WASM Web-feature check, 336 shared tests + headless device + 59 GPU goldens + 4 render-target tests. The focused VSM suite is 29/29.
Evidence:
docs/evidence/issue-132-directional-clipmap-v1.mdand.json.No additional acceptance checkbox is being marked yet. A snapped-origin crossing currently invalidates the affected level and safely falls back to CSM while it rebuilds. Rolling/toroidal page-table preservation plus moving-camera/light transition qualification remain the next directional milestone; GPU request compaction/submission, local-light projections/arbitration, constrained-adapter qualification, and quality-tier rollout also remain open.
Directional clipmap camera-rebase milestone complete
Implemented and qualified in draft PR #147 through
5abdf70.What landed:
- camera-centered clip levels now carry integer planar page origins plus exact stable projection keys;
- an origin-only rebase remaps virtual owners while preserving physical layers, depth, age, and content signatures;
- pages leaving the 32x32 address space are the only static pages freed;
- prior dynamic-overlay pages are invalidated before remap;
- light-basis, scale, depth, content, or unexpected matrix changes still invalidate the affected level and resolve through CSM;
- new telemetry reports level rebases and preserved/dropped page counts;
- stable frames bypass all transition classification (
f4da021), retaining the original fast path; quality-motion --vsm-dynamic --vsm-scrollis an opt-in exact light-plane boundary oracle.
Exact Apple M1 Max / Metal transition result, repeated three times:
- 1 level rebase, 152 pages preserved, 0 occupied pages dropped;
- 224 requests / 224 hits / 0 misses;
- 0 denials and 0 evictions;
- 13 safe invalidations, 8 page renders, 8 pending pages;
- fixed allocation remains 19,951,824 bytes;
- repeat RMSE <= 0.000009965, SSIM 1.0, 0% above tolerance;
- visual review: no page rectangles, missing shadow bands, newly unshadowed regions, or foliage discontinuities.
Stable-camera parity against the previous clipmap checkpoint was RMSE 0.000009865 / SSIM 1.0. The VSM-disabled control was RMSE 0.000005068 / SSIM 1.0 and reported zero allocation/work.
Five-run stationary medians versus the previous checkpoint:
- wall 13.025648 -> 12.766289 ms;
- shadow CPU 0.172756 -> 0.171354 ms;
- shadow GPU 4.029285 -> 3.899865 ms;
- page CPU 0.031248 -> 0.032132 ms (0.000884 ms noise inside lower total shadow CPU);
- page GPU 3.696149 -> 3.605031 ms.
Final quick lane passed FFI parity including Linux/Web, strict Clippy, 340 shared tests + 1 ignored, headless device construction, 59 GPU goldens + 2 policy ignores, 4 render-target tests, WASM, quality/cooker tools, and 20 examples. Focused VSM suite: 33/33.
Evidence:
issue-132-clipmap-scroll-v1.mdand machine-readable JSON.No acceptance checkbox is being claimed yet: fixed Sponza/Bistro motion and moving-light qualification remain. Next directional step is the conservative moving-light transition path, followed by GPU request/caster work.
Moving-light transition milestone complete
Qualified in PR #147 at
2f73135; evidence committed atdf9dbdb.The new opt-in
quality-motion --vsm-dynamic --vsm-light-motionoracle changes the primary directional basis every 30 frames and captures frame 240 as it returns to the ordinary light. A basis change cannot satisfy the clipmap scroll key, so all directional levels take the conservative invalidation path before page-table upload.All three Apple M1 Max / Metal repeats reported identical safety counters:
- 108 clean pages invalidated in the transition frame;
- 236 resident owners, 228 dirty/unsampleable pages after the bounded 8 renders;
- 224 requests / 224 ownership hits / 0 misses;
- 0 rebases, 0 preserved pages, 0 denials, 0 evictions;
- fixed 19,951,824-byte allocation.
Repeat envelope: RMSE <= 0.000023677, SSIM >= 0.999999821, 0% above tolerance. Transition versus settled same-final-light: RMSE 0.006145390 / SSIM 0.999100208. Same-history CSM control: RMSE 0.008730293 / SSIM 0.992735803. Full-image and heatmap review found no page boundaries, missing bands, newly unshadowed regions, or stale old-direction shadows.
Evidence:
issue-132-moving-light-v1.mdand JSON.The moving-caster/light acceptance checkbox is now complete: bounded dynamic overlays cover the moving-caster half, and this basis-change oracle covers moving light. Next implementation milestone is GPU-driven receiver marking, request compaction, caster culling, and page submission.
Fixed-address receiver request compaction complete
Implemented at
b4e96f0; evidence committed at4392dd0in PR #147.Directional receiver ranking now uses one fixed 1,024-entry R32-compatible coverage domain per level plus a sparse first-touched address list. This removes per-level hash tables while avoiding a dense scan for sparse scenes. The caps and complete request order remain 144/64/16 and 224 total.
Safety/equivalence:
- the previous hash implementation remains as a test oracle;
- complete ordered output matches across overlapping, outside, invalid, duplicated, and tied coverage;
- no GPU pass, readback, persistent allocation, residency change, or fallback change was introduced;
- stationary, moving-light, and disabled images all stayed at SSIM 1.0 with 0% above tolerance;
- VSM counters matched exactly and default-off remained zero bytes/work.
The moving-light workload recomputes coverage every 30 frames. Three-run medians all improved:
- wall 14.385408 -> 12.873031 ms;
- shadow CPU 0.218197 -> 0.181323 ms;
- page CPU 0.044228 -> 0.034917 ms;
- shadow GPU 2.969117 -> 2.864068 ms;
- page GPU 2.541226 -> 2.505326 ms.
Full quick lane passed: Linux/Web/all-platform FFI parity, strict Clippy, 342 shared tests + 1 ignored, headless device, 59 GPU goldens + 2 policy ignores, 4 render-target tests, WASM, quality/cooker tooling, and 20 examples. Focused VSM: 35/35.
Evidence:
issue-132-request-compaction-v1.mdand JSON.This is the bounded CPU oracle and storage ABI, not a claim that GPU marking is complete. Next: capability/workload-gated compute marking and compaction without same-frame blocking readback, then GPU page-caster culling/submission.
Async GPU receiver-marking checkpoint is implemented and pushed in PR #147 (
ad6e7c6, evidence/docsdb61815).Delivered boundary:
- capable native adapters use compute marking only for 1,024–4,096 camera-visible receiver bounds;
- smaller scenes retain the fixed CPU oracle with zero marker resources, dispatches, copies, or readbacks;
- resources and pipeline are lazy;
- two 12,288-byte readbacks are mapped asynchronously; production uses
Poll, never a same-frameWait; - first GPU output must exactly equal the complete ordered CPU demand or the backend disables itself and retains CPU demand;
- projection changes stay exact synchronous CPU transitions; continuous same-projection motion consumes the newest completed result, with missing addresses safely using CSM.
Qualification:
- real Metal dispatch test: complete 1,024-receiver GPU result exactly equals the CPU oracle;
- 1,140 moving receivers: 36 dispatches, 35 completions, validated=true, 0 validation failures, 102,896 lazy bytes;
- GPU vs CPU image: RMSE 0.000018280, SSIM 1.0, 0% above tolerance; cache/fallback counters match;
- paired large-workload CPU frame 18.615338→18.306413 ms, GPU frame 40.886617→37.220795 ms, shadow CPU 0.637508→0.386642 ms;
- the experimental 256 threshold was rejected after GPU timestamp inflation at ~285 receivers, so that workload remains CPU-only;
- default-off and ordinary two-receiver fixtures remain effectively identical with zero marker work;
- complete quick lane passed: all-platform FFI parity, strict lint, 345 shared unit tests + headless device, 59 GPU goldens, 4 render-target tests, WASM, quality/cooker, and 20 examples.
Detailed evidence: https://github.com/Bloom-Engine/engine/blob/codex/rendering-wip-integrated-20260727/docs/evidence/issue-132-async-gpu-receiver-v1.md
Architecture status (kept deliberately precise):
- bounded, capability/workload-gated GPU coverage marking without blocking readback
- GPU-resident request compaction and page-residency scheduling (next; current dense result is compacted by the exact CPU oracle after async readback)
- GPU page-caster culling and submission
Correction: GPU receiver backend retained explicit opt-in
Follow-up direct pass instrumentation invalidated the automatic performance qualification in the preceding comment. The correction is pushed in PR #147 at
9fe0371(default behavior) andb05c8b8(evidence/docs).What changed:
- production now defaults to the fixed CPU receiver oracle;
- GPU marking requires explicit
BLOOM_VSM_GPU_RECEIVER=1in addition to the existing capability and 1,024–4,096-bound gates; - the failed GPU rank/compact prototype was removed before commit;
- ordinary/default runs allocate 0 marker bytes and issue 0 marker dispatches;
- the exact GPU-vs-CPU validation and safe async fallback remain available as an experimental architecture oracle.
Why the earlier total-frame conclusion was rejected:
- the earlier paired run showed a
0.250866 msshadow-CPU reduction, but its window-server-throttled total GPU/wall values could not attribute marker cost; - direct timestamps measured the unchanged dense GPU mark at
0.485908 ms; - experimental rank and compact dispatches added
0.436717 msand0.504675 ms; combining the three in one compute pass still measured1.510075 ms; - that is not an across-the-board net win, so there is no qualified automatic workload crossover.
Verification after the correction:
- focused VSM suite: 38/38, including exact real-Metal GPU demand equality;
- production example with the 1,140-bound stress fixture and no GPU env opt-in reports
receiver_marking_backend: fixed-cpu,gpu_receiver_enabled: false, 0 dispatches, and 0 bytes.
Corrected evidence: https://github.com/Bloom-Engine/engine/blob/b05c8b8/docs/evidence/issue-132-async-gpu-receiver-v1.md
Issue status remains precise: fixed CPU request compaction is complete; automatic GPU request compaction/residency and GPU page-caster submission remain open and must pass direct pass-level timing before activation.
Implemented and qualified the next milestone in draft PR #147 at code revision
7962bee(evidence/docs at87b9d2e).Completed slice: bounded rigid-opaque page indirect submission
- Directional VSM pages with at least 48 visible eligible draws now submit compact
multi_draw_indexed_indirectranges. - Classification remains the exact CPU page-frustum oracle; the implementation performs a single partition pass and precomputes one clip matrix per caster.
- Only static rigid opaque geometry in the shared geometry arena is eligible. Cutout, skinned, foliage-motion, dynamic-overlay, dedicated-buffer, overflow, small, unsupported, Web, and GPU-driven-disabled paths remain on the compatibility renderer.
- Lazy allocation: 51,200 bytes in the 512-caster fixture; small VSM, VSM-off, GPU-driven-off, and
BLOOM_VSM_GPU_CASTERS=0controls allocate zero bytes. - Telemetry is exposed at
renderer_paths.vsm_gpu_casters.
Qualification
Five interleaved 60-warmup / 180-measured-frame runs on Apple M1 Max / Metal:
- wall-frame mean median: 11.893637 -> 11.457747 ms (-3.66%)
- CPU-frame mean median: 9.824165 -> 8.851101 ms (-9.90%)
- GPU-frame mean median: 20.292364 -> 19.262754 ms (-5.07%)
- VSM-page CPU: 0.106694 -> 0.094448 ms (-11.48%)
- VSM-page GPU: 1.092336 -> 0.582765 ms (-46.65%)
- render-total CPU: 4.921546 -> 4.435739 ms (-9.87%)
CPU-control and indirect captures are byte-identical across 3,686,400 pixels: RMSE 0, SSIM 1, max error 0, OKLab delta 0, edge delta 0. The full
scripts/ci-check.sh --quicklane passes, including FFI parity, Linux/Web builds/checks, 347 shared tests, 59 GPU goldens, WASM check, file-size ratchet, and quality/tooling gates.Explicit remaining boundary
This does not claim full GPU caster culling. A Metal fixed-count Cartesian prototype and a compute safety-cull prototype were rejected before commit. Remaining milestone-3 work is true GPU classification + indirect-count compaction on qualified backends, plus separate cutout/skinned/foliage/dynamic/instanced/dedicated-buffer qualification. Spot/point projections, shared arbitration, broader scenes, and tier rollout also remain open.
- Directional VSM pages with at least 48 visible eligible draws now submit compact
Completed and checked the geometric-contact-detail acceptance criterion in draft PR #147.
Code/tool checkpoint:
c1131d0; evidence/docs:3050fd7.Automated VSM-vs-CSM oracle
- New opt-in
quality-motion --vsm-contact-detailfixture: 266 colored 0.055 m posts over a neutral receive-only ground, fixed camera/light/exposure, 120 warm-up frames so all 224 demanded VSM pages are resident and clean. - New dependency-free
tools/quality/shadow_detail.pygate evaluates the same neutral-ground pixel intersection in both images and excludes colored geometry/sky/TAA chroma edges. - Negative controls prove sharp-vs-blurred passes, the reversed candidate fails, and colored geometry is excluded.
Five deterministic captures per mode:
- strong shadow-edge pixels: CSM 161, VSM 8,110 (50.37x)
- edge p99: 0.071697 -> 0.116095 (1.619x)
- 95th-to-5th percentile shadow contrast: 0.284005 -> 0.291565 (+0.007560)
- all VSM captures shared one SHA-256; all CSM captures shared one SHA-256
Measured incremental allocation is explicit: 19,951,824 VSM pool/metadata bytes + 51,200 caster-indirect bytes = 20,003,024 bytes. Five-run medians show GPU-frame mean effectively flat (+0.014120 ms / +0.03%), GPU p50 -2.90%, wall -2.71%, and CPU mean -5.82%. The VSM sampling cost itself is visible (
main_hdr_pass+0.231049 ms); lower whole-frame totals are not claimed as an optimization.This checkpoint changes no production renderer behavior: it adds only an opt-in example branch, offline semantic gate, tests, and evidence. The full quick lane passes.
Remaining #132 acceptance work is still Bistro/seam/ring qualification, alpha/skinned automation, constrained fallback qualification, debug-view completeness, and local-light VSM/arbitration.
- New opt-in
Forced lower-tier VSM fallback complete
Implemented in c34743e; evidence/docs are pushed at d306164 in draft PR #147.
What changed:
- Renderer startup records the accepted capability tier before VSM shader/layout/resource selection.
- BLOOM_VSM=1 on forced baseline or modern now compiles canonical CSM and reports lower-tier-csm-fallback.
- Telemetry separates requested, capability_eligible, enabled, active, and selection_reason.
- The public capability report exposes the same state under runtime_support.virtual_shadows.
- The capability table now states CSM fallback for baseline/modern and VSM-with-CSM-fallback for high-end.
Qualification:
- Before fix: forced baseline still activated 256 VSM pages and allocated 19,951,824 bytes.
- After fix: 120/120-frame forced-baseline request is byte-identical to the CSM control at 2560x1440; both SHA-256 596a8b611496b56b9f2e5747f95b706effa2b85e136e84e481afb1d0bc987d6f.
- Lower-tier VSM pool, page work, receiver work, and caster-indirect work are all zero.
- Forced modern also exactly matches its CSM control.
- High-tier VSM retains the previously qualified SHA-256 81c6d840a275ee4fe3ff9fbeb87a5a1e21b23bb1c087a4e24c024c2f778b0bde with all 224 pages resident and clean.
- Five counterbalanced timing pairs show no shadow-path cost: shadow CPU 0.087858 -> 0.088146 ms (+0.000288 ms, noise); no timing improvement is claimed.
- Full quick lane passes: 348 shared tests + 1 ignored, headless device, 59 GPU goldens + 2 policy ignores, 4 render-target tests, strict Clippy, FFI parity, file ratchet, WASM, quality tooling, and 20 examples.
- Supported Web crate, iOS aarch64, and Android aarch64 checks pass.
Evidence: https://github.com/Bloom-Engine/engine/blob/d306164/docs/evidence/issue-132-capability-fallback-v1.md
JSON: https://github.com/Bloom-Engine/engine/blob/d306164/docs/evidence/issue-132-capability-fallback-v1.jsonThe forced fallback acceptance box is now checked. Remaining acceptance work is fixed Sponza/Bistro motion quality, alpha/skinned automation, complete debug views, and 100 local shadow-requesting lights/shared arbitration.
Completed the alpha-tested foliage + skinned caster acceptance checkpoint in
862981f; evidence is recorded in3015dcb.What this adds:
- VSM telemetry now reports alpha-tested and skinned page draws independently, so an unloaded/missing caster cannot silently pass.
quality-motion --vsm-alpha-no-castis a negative control that keeps Sponza's MASK curtain visible while removing only its shadow submission.tools/quality/vsm_caster_coverage.pycompares that control and an 8-frame Fox pose offset against fixed ground ROIs. It fails closed on missing/malformed telemetry, unchanged shadows, weak coverage, or opaque replacement coverage.
Real Apple M1 Max / Metal result at 1600x900, 120 warm-up + 120 measured frames:
- Alpha-tested shadow: 19,931 / 71,280 ROI pixels changed (27.962%), 4.646 segments/occupied row, p95 luminance delta 0.178709.
- Skinned shadow: 4,465 / 51,840 ground-only ROI pixels changed (8.613%), p95 luminance delta 0.201384.
- Full telemetry: 4 cutout + 4 skinned page draws, 4 rendered pages / 4-page budget, 8 draws / 64-draw budget, 119 deferred pages covered by live CSM fallback, 19,951,824 VSM bytes.
- Repeat variance was bounded to a few 1-code-value pixels: max 1/255, 0% above 0.02, SSIM 1.0.
No render pass, shader, resource, binding, page, or draw was added. The only production-path work is two integer classifications inside the already opt-in VSM page draw loop; default CSM behavior is unchanged.
Regression result: complete quick lane passed (348 shared + 1 ignored, 59 GPU goldens + 2 policy ignores, 4 render-target tests, 25 quality tests, strict Clippy, FFI, ratchet, Web/WASM, and 20 examples).
Evidence: https://github.com/Bloom-Engine/engine/blob/3015dcb/docs/evidence/issue-132-caster-coverage-v1.md
I marked only this acceptance checkbox complete. The fixed Sponza/Bistro motion corpus, 100 shadow-requesting local lights, and complete debug views remain open.
Completed the fixed Sponza/Bistro camera-path acceptance checkpoint.
Implementation:
2ee8903adds opt-in 30-frame light-plane camera paths to both established examples plustools/quality/vsm_motion_corpus.py.aa36327makes all fixture bookkeeping/path math conditional on the opt-in flag.91e7843records the exact captures and machine-readable evidence.
The oracle uses four matched captures per scene (settled VSM/CSM + motion VSM/CSM) and evaluates
(motion VSM - motion CSM) - (settled VSM - settled CSM). This removes shared auto-exposure/TAA/SSGI/SSR history and isolates only VSM transition behavior.Apple M1 Max / Metal results at 1600x900:
- Sponza: residual RMSE 0.002120, p99 0.008394, 0.0052% pixels above 0.03, largest connected artifact 20 pixels, no seam/ring-like component.
- Bistro: residual RMSE 0.004746, p99 0.018491, 0.4667% pixels above 0.03, largest connected artifact 225 pixels (0.0156% of frame), no seam/ring-like component.
- Missing-shadow/stale-shadow ratios stayed below 0.1% in both scenes.
- Sponza transition: 2 rebases, 228 pages preserved, zero dropped/dirty/rendered/denied/evicted pages.
- Bistro transition: 3 rebases, 202 preserved + 54 safely dropped, 8 bounded renders, 26 dirty pages excluded from sampling, zero denials/evictions.
- Both retained the fixed 19,951,824-byte allocation.
- Independent repeat SSIM: 0.999997 Sponza, 0.999954 Bistro, with identical cache/work counters.
Six seeded failures prove page seams, clipmap rings, missing-page flashes, stale/doubled shadows, missing rebases, and malformed telemetry fail closed. A seventh test proves matched controls remove unrelated temporal/exposure history.
No renderer/pass/shader/resource/API behavior changed in this checkpoint. The paths and offline oracle are opt-in. The exact pushed state passed the full quick lane: FFI, strict Clippy, ratchet, Web/WASM, 348 shared tests, 59 GPU goldens, 4 render-target tests, 32 quality tests, and 20 examples.
Evidence: https://github.com/Bloom-Engine/engine/blob/91e7843/docs/evidence/issue-132-full-scene-motion-v1.md
Scope note: Sponza is the complete checked-in asset. Bistro is the governed 96 unique-mesh camera-visible subset from the pinned source, not a claim of full 2,909-instance streaming coverage.
I marked only the fixed full-scene motion checkbox complete. The 100+ shadow-requesting local-light and complete debug-view criteria remain open.
Completed the full debug-view acceptance criterion in
840eb49with evidence in20ac21e.What is now exposed by an opt-in intermediate capture:
virtual-shadow-pages.png: the three 32×32 virtual clip levels, stacked near-to-far;virtual-shadow-physical.png: physical pool occupancy in slot order;- distinct amber never-rendered misses and magenta previously-rendered invalidations;
- stable near/middle/far clip-level colors plus
virtual-shadow-legend.png; - a same-frame
virtual-shadow-report.jsoncontaining the debug contract and a directional per-light cost row (requests/hits/misses/invalidations/renders/residency/dirty/rebases/draws/owned and shared bytes/budget).
The fail-closed validator (
tools/quality/vsm_debug_views.py) requires exact palette and geometry, whole cells, matching virtual/physical page counts, report-consistent occupancy/dirty/clip levels, and reconciled per-light counters and bytes.Real Apple M1 Max / Metal captures passed:
- early fill: 212 visible never-rendered misses;
- light transition: 102 visible invalidations plus 126 never-rendered misses, with 236 resident / 228 dirty pages;
- settled static scene: 144/64/16 clean pages across levels 0/1/2 and zero dirty pages.
This is capture-only and changes no shader, pass, GPU resource, binding, draw, page policy, sampling, or fallback.
scripts/ci-check.sh --quickpasses (349 shared unit tests + 1 ignored, device negotiation, 59 GPU goldens + 2 policy ignores, 4 render-target tests, Web/WASM, 39 quality tests, FFI parity, and 20 examples).Evidence:
docs/evidence/issue-132-debug-views-v1.{md,json}.With this checkbox complete, #132 is 7/8; only the 100 shadow-requesting local-lights acceptance criterion remains.
Point-light portion of the final local-light acceptance criterion is now qualified on PR #147.
Implemented and pushed:
a01cf8a— bounded shadow-required point-light API, deterministic visibility/admission, six-face shared VSM pages, retained/immediate/froxel sampling, fail-closed suppression, telemetry, 128-light fixture, and validator.630cb23— uniform no-request early-out before dynamically indexed local metadata.93ebdc3/6a70ca7— test-only source split required by the file-size gate.e27793e— checked-in human- and machine-readable evidence.
Measured on Apple M1 Max / Metal:
- 128 submitted and visible; 5 admitted; 123 budget-suppressed.
- 30 local pages maximum (5 x 6), sharing the existing 256-page pool and existing 8-page/frame render budget.
- Cold: exactly 8 local pages rendered; pending lights remained zero-contribution; dirty directional pages used CSM fallback.
- Warm: all 30 local pages clean, all 5 admitted lights active, zero page renders.
- Controlled five-light shadowed/unshadowed A/B changed 0.858% of pixels, localized behind occluders (RMSE 0.01009, SSIM 0.99437); visual review found no cube seams, page blocks, missing bands, or frame-wide shift.
- Both >=100-light telemetry reports pass the fail-closed validator.
- Exact pushed head passed
scripts/ci-check.sh --quick: 355 shared tests, 59 GPU goldens, 4 render-target tests, Web/WASM and all-platform FFI parity, 39 quality tests, visual metrics, asset cooker, and 20 canonical examples.
Evidence:
docs/evidence/issue-132-local-lights-v1.mddocs/evidence/issue-132-local-lights-v1.json
I marked the >=100 local-light criterion complete. I am leaving #132 open because its architecture/milestone text still explicitly includes a spot-light virtual projection; this delivery qualifies bounded point-light cube projections, not spot lights.
Completed the final spot-light milestone and closing #132.
Pushed on PR #147:
0bcd2c7— publicaddShadowedSpotLight(...), one-page perspective projection, circular smooth cone sampling, deterministic shared-pool admission, fail-closed suppression, all-platform/Web ABI handling, telemetry, fixture, and validator.d91d9c6/69b7403— human/machine-readable evidence and follow-up-scope documentation.
Apple M1 Max / Metal qualification:
- 128 submitted and visible; 5 admitted; 123 suppressed before retained/immediate shading and froxel assignment.
- Exactly 1 page per admitted spot and 5 local pages total, versus 6/30 for point cubes; the existing 256-page pool and 8-page render budget are reused.
- Warm state: all 5 spots active, 5/5 pages clean, zero page renders, zero denials.
- Same-scene/same-cone caster-on/off oracle: RMSE 0.01644, SSIM 0.98078, 2.027% pixels above 0.02. Differences are localized to occluder silhouettes; visual review found no page rectangles, projection seams, missing bands, cone discontinuities, or frame-wide shift.
- No persistent GPU bytes or graph passes added over the point milestone. Directional-only requests retain the uniform early-out.
scripts/ci-check.sh --quickpasses: 356 shared tests + 1 ignore, headless device construction, 59 GPU goldens + 2 hardware-policy ignores, 4 render-target tests, FFI/Web/Wasm, strict Clippy/format/ratchet, quality/visual/cooker suites, and all examples. The ordinary immediate and clustered many-point-light goldens pass exactly. The native Web crate check also passes.
Evidence:
docs/evidence/issue-132-spot-lights-v1.mddocs/evidence/issue-132-spot-lights-v1.json
All acceptance boxes are complete, including the architecture milestone requiring both point and spot projections. GPU-only request scheduling and broader indirect specialization remain documented optional follow-ups rather than acceptance blockers.
Post-closure Bistro shadow follow-up is pushed in PR #147.
6da56aeselects scene-receiver cascades from the same positive view-space depth used by frustum fitting instead of spherical camera distance.06be86eapplies the same selection contract to shared/material receivers.- Both paths have source regressions, and the targeted
shadow_cascadetest set passes (3/3). - Fresh-start and same-process camera-rotation captures converge to the same world-space shadow footprint, ruling out stale cascade-cache state.
- A CPU path-traced render reproduces the broad building-shadow silhouette. Forcing all receivers to cascade 1 leaves the raised paving-island region visually unchanged (mean A/B delta below one 8-bit RGB level), confirming that region is authored material/ambient contrast rather than missing cascade coverage.
- Transmitted shadows were disabled in the opaque CSM A/B, so they are not masking the result.
The clean production-shader Bistro binary was rebuilt after removing the diagnostic override. CI cleanup
e7bd0c1also restores the PR contracts/lint lanes locally.Movement/rotation follow-up landed in
8d85719(fix(render): preserve shadows across cascade handoff).Root cause addressed: cross-cascade blending could sample the next cascade even when its retained translation-slack fit did not cover the receiver. The core path then read a clamped edge texel, while the material path treated the miss as fully lit; either path could erase part of an otherwise valid shadow as the camera crossed a split.
The fix makes the handoff coverage-aware in both paths: an out-of-fit next cascade preserves the current cascade result until the next fit actually covers the receiver. It adds no shadow texture samples and leaves the valid overlap path unchanged. Added regression guards for both shader families.
Verification complete: 4 targeted cascade tests green; native/folded WGSL parse green; quick contracts/FFI parity green; strict rendering lint green; macOS release build and Bistro binary rebuilt. Translation-only cache-vs-always-fresh A/B is effectively identical away from overlay text, ruling out stale depth for that leg.
Still pending: human traversal approval for the reported combined move+turn path. The prior open Bistro window predated
8d85719. Machine free space fell below 10 GiB due other active builds, so further compiler-driven oracle expansion is paused rather than risking the workspace.Final moving-camera cascade validation
The post-closure cascade-handoff regression is now resolved at
67661eeand independently human-validated in the interactive Bistro scene.Validation covered the reported failure mode rather than only the launch pose: moving and turning around the restaurant/street receiver while the directional shadow crossed cascade-fit boundaries. The previously missing rectangular/slab regions remained filled, and the reporter confirmed: “perfect, shadows work now.”
Regression evidence:
- selected-cascade projection misses now hand off to the next valid cascade instead of returning fully lit;
- the ordinary in-fit path still performs one cascade depth sample;
- all 55 shadow-focused library tests pass;
- deterministic ordinary in-fit captures remain pixel-identical apart from the timing overlay;
- follow-up
cc4b663moves byte-identical 2D WGSL out ofcore.rs(matching SHA-256) and restores the PR file-size gate without changing shader output or runtime work.
This completes the human movement-validation boundary called out in the prior follow-up.
Parent: #126
Problem
Bloom's cascaded sun shadows and per-light approaches cannot scale to UE-class geometric detail and many shadowed lights without fixed-resolution waste, aliasing, cascade transitions, or a shadow-map-per-light explosion. #11 already identifies virtual shadow maps as a major visual gap.
Outcome
Page-cached virtual shadow maps (VSMs) for directional, spot, and point lights on capable tiers, with stable filtering, explicit invalidation, bounded memory, and a correct cascaded/conventional fallback.
Architecture contract
Address space and cache
Page demand and rendering
Sampling and quality
Capability fallback
Milestones
Acceptance criteria
Likely files
native/shared/src/shadows.rsnative/shared/src/renderer/graph.rs,scene_pass.rs,hiz.rs, shader modulesnative/shared/src/src/Verification
Run the quality corpus with camera cuts, rapid movement, dynamic casters, alpha-tested assets, and a forced small page budget. Attach uncapped GPU timings, cache statistics, memory, and fallback captures.
Non-goals
Dependencies