Skip to content

[Renderer] Add virtual shadow maps with cached pages and scalable many-light shadows #132

Description

@proggeramlug

Parent: #126

Problem

Bloom's cascaded sun shadows and per-light approaches cannot scale to UE-class geometric detail and many shadowed lights without fixed-resolution waste, aliasing, cascade transitions, or a shadow-map-per-light explosion. #11 already identifies virtual shadow maps as a major visual gap.

Outcome

Page-cached virtual shadow maps (VSMs) for directional, spot, and point lights on capable tiers, with stable filtering, explicit invalidation, bounded memory, and a correct cascaded/conventional fallback.

Architecture contract

Address space and cache

  • Per-light virtual address space mapped through page tables to a shared physical page pool.
  • Directional light uses camera-centered clip levels with stable texel/page snapping.
  • Spot lights use a 2D projection; point lights use a documented cube/octahedral representation.
  • Page metadata tracks owner, virtual coordinate, last use, dirty state, and generation.
  • Fixed memory budget with deterministic allocation/eviction and coarse fallback coverage.

Page demand and rendering

  • Mark pages from the camera-visible receiver set; optionally use depth/visibility buffer to avoid requesting empty screen regions.
  • Compact requested/dirty pages on GPU.
  • Cull shadow casters against each page and issue GPU-driven page renders where supported.
  • Static pages persist. Dynamic caster/light/transform changes invalidate only affected pages, conservatively.
  • Missing pages sample a coarser resident level or conventional fallback—never unshadowed by accident.

Sampling and quality

  • Cross-page filtering must not reveal seams; include borders/gutters or neighbor-aware sampling.
  • Receiver-plane/depth bias and normal bias are scene-scale aware and exposed through quality presets, not arbitrary per-scene hacks.
  • Directional clip-level transitions blend without visible rings.
  • Alpha-tested casters preserve coverage. Translucent colored shadows are out of scope unless explicitly added later.

Capability fallback

  • High tier: VSM page marking/caching/GPU-driven render.
  • Middle/baseline tier: improved CSM/conventional per-light maps with cascade blending and caching.
  • Capability report states which path is active and why.

Milestones

  1. Directional-light address/page table with debug visualization.
  2. Receiver demand, physical cache, and coarse fallback.
  3. Static/dynamic invalidation and GPU-driven caster submission.
  4. Seam-safe filtering and clip-level blending.
  5. Spot and point light support plus shared budget arbitration.
  6. Quality-tier integration and fallback qualification.

Acceptance criteria

  • Fixed Sponza/Bistro camera paths show no page seams, clipmap rings, cascade pops, missing-page flashes, or persistent stale shadows.
  • A moving caster and moving light invalidate affected pages within one frame; unrelated static pages retain a high cache-hit rate.
  • A geometric-detail stress scene preserves contact detail demonstrably better than current CSM at an equal or measured memory budget.
  • At least 100 shadow-requesting local lights can be submitted; actual shaded/page cost is visibility/budget bounded and reported.
  • Page pool memory never exceeds the configured budget plus documented metadata/staging overhead.
  • Alpha-tested foliage coverage and skinned casters have automated cases.
  • Forced fallback path remains correct on web/mobile/constrained adapters.
  • Debug views expose virtual pages, physical occupancy, misses, invalidations, clip levels, and per-light cost.

Likely files

  • native/shared/src/shadows.rs
  • native/shared/src/renderer/graph.rs, scene_pass.rs, hiz.rs, shader modules
  • scene/light data in native/shared/src/
  • public quality/capability APIs in src/

Verification

Run the quality corpus with camera cuts, rapid movement, dynamic casters, alpha-tested assets, and a forced small page budget. Attach uncapped GPU timings, cache statistics, memory, and fallback captures.

Non-goals

  • Ray-traced shadows replacing VSMs everywhere.
  • Requiring VSMs on low tiers.
  • Colored translucent shadow transport.

Dependencies

Activity

  1. proggeramlug commented on Jul 25, 2026

    @proggeramlug
    ContributorAuthor

    Directional VSM milestone implemented and qualified behind the explicit BLOOM_VSM=1 gate. The issue remains open for dynamic-page rendering, receiver-driven demand, local lights, and tier rollout.

    Implemented:

    • Fixed 3-level virtual address space: 32x32 pages per level, 128x128 interior texels, 2-texel gutters, 132x132 physical pages.
    • Deterministic fixed-budget LRU residency with current-frame eviction protection, missing-page CSM fallback, signature invalidation, eight-frame CSM-to-VSM residency fade, and a hard device/config capacity bound.
    • Center-first, interleaved level scheduling so an invalidating near level cannot starve mid/far pages.
    • Opt-in GPU resources only: R32Uint page table, Depth32Float physical array, bounded dedicated per-page uniforms, and a zero-work/default allocation path.
    • Physical page rendering reuses the existing opaque and alpha-cutout shadow pipelines, performs per-page caster culling, preserves gutters, and uses a dedicated uniform buffer so deferred queue writes cannot corrupt CSM matrices.
    • Built-in scene and material-ABI receiver sampling. Missing, dirty, over-budget, and dynamically unsafe pages resolve through CSM, never unshadowed.
    • Lazy fallback sampling: resident steady-state pixels pay the 4-tap VSM kernel, not both 16-tap CSM and VSM. This fixed the first A/B performance regression before landing.
    • Stable-cache fast path eliminates repeated 224-page cache walks and page-table/parameter uploads after residency/fade stabilizes.
    • Telemetry now reports active/fallback state, capacity, physical/overhead/total bytes, hit/miss/eviction/denial/invalidation/render counts, and per-level residency.
    • Quality captures now emit virtual-shadow-pages.png and virtual-shadow-physical.png occupancy views.

    Safety contract:

    • Default mode remains the canonical CSM shader/layout and allocates no VSM GPU resource or per-pixel branch.
    • Continuous skinned/dynamic casters conservatively set dynamic_fallback: true; the exact live CSM path remains authoritative until dynamic VSM overlays are implemented.
    • Alpha-tested static casters render through the cutout page pipeline.

    Evidence on Apple M1 Max / Metal:

    • Sponza static run: active: true, 224/224 resident, 224 hits, 0 dirty, 0 denied, 0 evictions; level residency 144/64/16.
    • Sponza VSM-vs-CSM final: SSIM 0.990153, luminance RMSE 0.011009, OKLab mean delta 0.002991, edge delta 0.002022. The visible change is the intended higher-resolution/sharper shadow boundary; the diff gate passes.
    • Latest same-machine Sponza main pass: 1.498 ms CSM reference vs 1.302 ms VSM, while steady shadow-pass CPU returned to about 0.040 ms after the cache fast path. Whole-frame GPU values remain scheduler/noise dominated and are not claimed as a hard gate yet.
    • Skinned/alpha motion: dynamic_fallback: true; VSM-gated vs CSM final is effectively identical: SSIM 1.00000, luminance RMSE 0.00003, zero pixels above tolerance.
    • Quality runner failure in both cases is only the repository-wide missing approved portable baseline, not capture, validation, or image-threshold failure.
    • cargo test --manifest-path native/shared/Cargo.toml --lib --no-fail-fast: 181 passed, 0 failed, 1 ignored.
    • Default golden_lit_primitives_3d: passed.
    • Focused VSM suite: 15 tests, covering coordinate bounds, LRU and eviction protection, memory bounds, page-table age/fallback, crop/gutter math, fair scheduling, fast-path accounting, debug output, and WGSL parsing for scene/material variants.

    Still required before closing #132:

    • Receiver/depth-driven page marking and GPU compaction/submission.
    • Page-granular dynamic/skinned overlays instead of whole-feature CSM fallback.
    • True snapped directional clipmaps independent of the CSM oracle.
    • Spot/point-light virtual projections and shared pool arbitration for 100+ submitted shadow lights.
    • Bistro/moving-light/small-pool stress qualification, seam/ring/pop gates, quality-tier integration, and high-tier default enablement only after those gates pass.
  2. proggeramlug commented on Jul 25, 2026

    @proggeramlug
    ContributorAuthor

    Receiver-driven demand and dynamic/skinned safety milestone completed and qualified. #132 remains open for true clipmaps, GPU compaction/submission, local-light projections, and tier rollout.

    Implemented in this milestone:

    • Camera-visible receiver AABBs are projected into every directional level and expanded by a one-page filter/jitter guard.
    • Requests are ranked deterministically by receiver coverage and center distance, capped at 144/64/16 pages, and interleaved across levels. Missing or omitted pages continue to sample CSM.
    • Receiver-to-page results are cached by receiver-bounds signature and level matrices. The first implementation rebuilt the coverage hash every frame and regressed Sponza shadow CPU from about 0.04 ms to 0.333 ms; that version was rejected. The cached implementation is back to 0.048-0.057 ms in steady static captures.
    • Moving/skinned receivers are excluded from the static-demand signature so animation does not rebuild static page ranking each pose.
    • Dynamic casters now have conservative light-space page masks with a two-page guard. Masked pages use live CSM and never sample stale static VSM depth.
    • Small receiver footprints automatically keep the proven whole-frame CSM fallback because page-mask construction and VSM indirection cannot repay their cost. Larger footprints can retain VSM on unrelated static pages. Unbounded casters preserve whole-frame fallback.
    • Dynamic fallback telemetry now reports whole-frame-csm, full-demand-csm, or per-page-csm, plus page count.
    • Disabled VSM telemetry now reports actual zero physical capacity, bytes, overhead, and render budget rather than the internal one-slot CPU sentinel.

    Evidence on Apple M1 Max / Metal:

    • Final static Sponza: active:true, demand_source:receiver-bounds, 224 resident/requested/hits, 0 misses, 0 denied, 0 evictions, 0 dirty; 144/64/16 pages by level.
    • Static Sponza main pass: 1.268 ms in the latest run. Shadow CPU 0.0566 ms. Output versus the preceding receiver-cached capture: SSIM 1.00000, luminance RMSE 0.00010, 0.00% above tolerance.
    • Skinned/alpha motion: 67 receiver-demand pages, adaptive whole-frame-csm, shadow CPU 0.0327 ms versus 0.0312 ms CSM reference. Output versus CSM: SSIM 1.00000, luminance RMSE 0.00003, 0.00% above tolerance.
    • Default golden_lit_primitives_3d: pass.
    • Full library suite: 188 passed, 0 failed, 1 ignored.
    • Focused VSM suite: 22 passed, including bounded/unique/fair receiver demand, deterministic signatures, guard coverage, offscreen rejection, per-page dynamic masking, and unbounded-caster fallback.
    • Quality runner status is fail only because the repository does not contain the approved portable baseline; capture and explicit image comparisons pass.

    Local-light audit / required nonbreaking contract:

    Bloom currently exposes only unshadowed point lights. There is no spot-light type and no public or native shadow-request flag. Existing content, including the 192-point-light stress case, depends on that behavior. Automatically treating current point lights as shadowed would be a visual compatibility break and an unbounded performance regression, so this milestone does not do that implicitly.

    The safe local-light implementation should therefore:

    1. Add an explicit, default-false shadow request and priority to point lights, plus a first-class spot-light API. Preserve existing addPointLight byte-for-byte and route new options through additive FFI entry points.
    2. Extend the virtual address key with projection kind and face. Use 2D spot projection and a documented six-face cube or octahedral point representation; do not overload directional clip levels.
    3. Keep one physical page allocator for directional, spot, and point owners. Arbitration must be deterministic and visibility/priority aware, with per-light minimum/coarse fallback and a hard global byte/page cap.
    4. Add per-light page-table metadata outside the existing lighting UBO so unshadowed lights do not grow or slow the canonical loop.
    5. Qualify two negative controls: the current 192 unshadowed lights remain image/performance equivalent, and 100 explicitly shadow-requesting lights stay bounded with reported admitted/deferred/page cost.
    6. Only then connect point/spot sampling in the clustered and material paths; missing/deferred pages must resolve to conventional coarse shadow or an explicitly documented no-shadow policy chosen by the caller, never silently change existing lights.

    This preserves the no-regression constraint while leaving an independently implementable local-light contract.

  3. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Dynamic/skinned directional page-overlay milestone completed and qualified in 6b4c256 with evidence/docs in 75cdfe0. #132 remains open for true snapped clipmaps, GPU compaction/submission, local-light projections/arbitration, and rollout qualification.

    Implemented:

    • Bounded dynamic/skinned caster AABBs are projected with a two-page PCF/jitter guard; pages nearest each caster core are selected first, including separated multi-caster footprints.
    • Up to 4 affected pages and 64 total page draws are rebuilt per frame with both static and current-frame dynamic geometry. Opaque, MASK/cutout, foliage-motion, and skinned pipelines reuse the live CSM geometry, cutout bindings, joint palette, and wind parameters.
    • Prior/current affected pages are invalidated before page-table upload. Only successfully rendered pages become sampleable; deferred, over-budget, missing, and dirty pages use current CSM. There is no stale animated VSM depth or unshadowed fallback.
    • Small receiver sets (<128 pages) and unbounded casters retain whole-frame CSM. BLOOM_VSM unset remains the zero-allocation, zero-pass canonical CSM path.
    • Hard telemetry now exposes overlay footprint, rendered pages, draws, deferred pages, and page/draw budgets.
    • Added opt-in quality-motion --vsm-dynamic oracle and operational documentation in docs/virtual-shadow-maps.md.

    Final Apple M1 Max / Metal evidence:

    • 224 demanded/resident pages; 151 guarded dynamic pages; 4 overlay pages rendered with 8 draws; 75 demanded guard pages left dirty for CSM; 0 denied and 0 evicted. Total VSM memory remains fixed at 19,951,632 bytes.
    • Repeated overlay capture: SSIM 1.000000000, luma RMSE 0.000005030, 0% above 0.02 tolerance.
    • Prior per-page-CSM control to overlay: SSIM 0.992135584, RMSE 0.008894581; reviewed differences are the intended shadow/foliage/contact-edge detail, without page rectangles or missing shadow regions.
    • VSM-disabled old/new control: SSIM 1.000000000, RMSE 0.000012179, 0% above tolerance, and telemetry remains zero VSM bytes/work.
    • Interleaved timing medians: wall 13.873404 ms control vs 13.760610 ms overlay; shadow GPU 4.008879 vs 4.100516 ms (0.091637 ms absolute); bounded shadow CPU 0.054876 vs 0.180561 ms. Four passes/64 draws are hard ceilings and excess work uses already-rendered CSM.
    • 331 shared unit tests + headless device test + 59 GPU goldens + 4 render-target tests passed; 24 focused VSM tests passed. Contracts, strict Clippy, file-size ratchet, native release, and wasm Web-feature checks pass.

    Evidence: docs/evidence/issue-132-dynamic-vsm-overlay-v1.md and .json.

    Acceptance checklist: only the already-proven fixed memory-budget item is newly checked. Moving-light behavior, independent clipmaps, 100 local lights, VSM-specific automated foliage/skinned corpus coverage, constrained-adapter qualification, and full camera-path seam/pop gates remain deliberately open.

  4. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Independent page-snapped directional clipmap milestone completed in 3861584, with exact-revision qualification/docs in 9851407.

    Implemented:

    • Replaced VSM reuse of fitted CSM matrices with three independent camera-centered orthographic clip levels. Light-space X/Y origins snap to one virtual-page footprint; sub-page camera motion keeps every matrix and cached address byte-stable.
    • Added split-derived coverage guard plus quantized scene-depth pancaking. Receiver demand, physical-page rendering, scene shading, and material shading now use the same clipmap matrices.
    • Preserved the independently fitted CSM matrices as live fallback data. Missing, dirty, deferred, or out-of-volume VSM samples use the original CSM projection; no stale or accidentally unshadowed sample is exposed.
    • Expanded the sampling ABI from 16 to 208 bytes for three matrices plus policy words, with exact Rust/WGSL size/alignment tests and all scene/material bind layouts updated.
    • Added telemetry projection identity: camera-centered-page-snapped-clipmap.
    • Kept the entire path opt-in. With BLOOM_VSM unset, clipmaps are not computed, VSM resources/passes remain zero, and canonical shaders remain uninjected.

    Exact Apple M1 Max / Metal evidence:

    • Dynamic fixture: 224 demanded/resident pages, 153 guarded dynamic pages, 4 overlay pages / 8 draws, 119 deliberately dirty demand pages using CSM, 0 denials, 0 evictions, and fixed 19,951,824-byte VSM memory.
    • Static Sponza: 224 resident pages, zero dirty pages, denials, evictions, or steady-state page renders.
    • Three repeat captures: worst SSIM 0.999999940, luma RMSE 0.000021062, and 0% above the 0.02 tolerance.
    • Fitted-control to clipmap: dynamic SSIM 0.994245768 / RMSE 0.006658406; static Sponza SSIM 0.996063292 / RMSE 0.005528215. Full-image and heatmap review showed localized shadow-edge/foliage-detail changes with no page rectangles, missing shadows, or new unshadowed geometry.
    • VSM-disabled control: SSIM 1.000000000, RMSE 0.000013567, 0% above tolerance, and zero VSM bytes/work.
    • Median versus the preceding fitted-overlay qualification: wall 13.760610 -> 13.025648 ms, shadow CPU 0.180561 -> 0.172756 ms, shadow GPU 4.100516 -> 4.029285 ms. No measured steady-state regression.
    • Quick CI lane passed: all-platform FFI/schema parity, strict Clippy, file-size ratchet, WASM Web-feature check, 336 shared tests + headless device + 59 GPU goldens + 4 render-target tests. The focused VSM suite is 29/29.

    Evidence: docs/evidence/issue-132-directional-clipmap-v1.md and .json.

    No additional acceptance checkbox is being marked yet. A snapped-origin crossing currently invalidates the affected level and safely falls back to CSM while it rebuilds. Rolling/toroidal page-table preservation plus moving-camera/light transition qualification remain the next directional milestone; GPU request compaction/submission, local-light projections/arbitration, constrained-adapter qualification, and quality-tier rollout also remain open.

  5. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Directional clipmap camera-rebase milestone complete

    Implemented and qualified in draft PR #147 through 5abdf70.

    What landed:

    • camera-centered clip levels now carry integer planar page origins plus exact stable projection keys;
    • an origin-only rebase remaps virtual owners while preserving physical layers, depth, age, and content signatures;
    • pages leaving the 32x32 address space are the only static pages freed;
    • prior dynamic-overlay pages are invalidated before remap;
    • light-basis, scale, depth, content, or unexpected matrix changes still invalidate the affected level and resolve through CSM;
    • new telemetry reports level rebases and preserved/dropped page counts;
    • stable frames bypass all transition classification (f4da021), retaining the original fast path;
    • quality-motion --vsm-dynamic --vsm-scroll is an opt-in exact light-plane boundary oracle.

    Exact Apple M1 Max / Metal transition result, repeated three times:

    • 1 level rebase, 152 pages preserved, 0 occupied pages dropped;
    • 224 requests / 224 hits / 0 misses;
    • 0 denials and 0 evictions;
    • 13 safe invalidations, 8 page renders, 8 pending pages;
    • fixed allocation remains 19,951,824 bytes;
    • repeat RMSE <= 0.000009965, SSIM 1.0, 0% above tolerance;
    • visual review: no page rectangles, missing shadow bands, newly unshadowed regions, or foliage discontinuities.

    Stable-camera parity against the previous clipmap checkpoint was RMSE 0.000009865 / SSIM 1.0. The VSM-disabled control was RMSE 0.000005068 / SSIM 1.0 and reported zero allocation/work.

    Five-run stationary medians versus the previous checkpoint:

    • wall 13.025648 -> 12.766289 ms;
    • shadow CPU 0.172756 -> 0.171354 ms;
    • shadow GPU 4.029285 -> 3.899865 ms;
    • page CPU 0.031248 -> 0.032132 ms (0.000884 ms noise inside lower total shadow CPU);
    • page GPU 3.696149 -> 3.605031 ms.

    Final quick lane passed FFI parity including Linux/Web, strict Clippy, 340 shared tests + 1 ignored, headless device construction, 59 GPU goldens + 2 policy ignores, 4 render-target tests, WASM, quality/cooker tools, and 20 examples. Focused VSM suite: 33/33.

    Evidence: issue-132-clipmap-scroll-v1.md and machine-readable JSON.

    No acceptance checkbox is being claimed yet: fixed Sponza/Bistro motion and moving-light qualification remain. Next directional step is the conservative moving-light transition path, followed by GPU request/caster work.

  6. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Moving-light transition milestone complete

    Qualified in PR #147 at 2f73135; evidence committed at df9dbdb.

    The new opt-in quality-motion --vsm-dynamic --vsm-light-motion oracle changes the primary directional basis every 30 frames and captures frame 240 as it returns to the ordinary light. A basis change cannot satisfy the clipmap scroll key, so all directional levels take the conservative invalidation path before page-table upload.

    All three Apple M1 Max / Metal repeats reported identical safety counters:

    • 108 clean pages invalidated in the transition frame;
    • 236 resident owners, 228 dirty/unsampleable pages after the bounded 8 renders;
    • 224 requests / 224 ownership hits / 0 misses;
    • 0 rebases, 0 preserved pages, 0 denials, 0 evictions;
    • fixed 19,951,824-byte allocation.

    Repeat envelope: RMSE <= 0.000023677, SSIM >= 0.999999821, 0% above tolerance. Transition versus settled same-final-light: RMSE 0.006145390 / SSIM 0.999100208. Same-history CSM control: RMSE 0.008730293 / SSIM 0.992735803. Full-image and heatmap review found no page boundaries, missing bands, newly unshadowed regions, or stale old-direction shadows.

    Evidence: issue-132-moving-light-v1.md and JSON.

    The moving-caster/light acceptance checkbox is now complete: bounded dynamic overlays cover the moving-caster half, and this basis-change oracle covers moving light. Next implementation milestone is GPU-driven receiver marking, request compaction, caster culling, and page submission.

  7. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Fixed-address receiver request compaction complete

    Implemented at b4e96f0; evidence committed at 4392dd0 in PR #147.

    Directional receiver ranking now uses one fixed 1,024-entry R32-compatible coverage domain per level plus a sparse first-touched address list. This removes per-level hash tables while avoiding a dense scan for sparse scenes. The caps and complete request order remain 144/64/16 and 224 total.

    Safety/equivalence:

    • the previous hash implementation remains as a test oracle;
    • complete ordered output matches across overlapping, outside, invalid, duplicated, and tied coverage;
    • no GPU pass, readback, persistent allocation, residency change, or fallback change was introduced;
    • stationary, moving-light, and disabled images all stayed at SSIM 1.0 with 0% above tolerance;
    • VSM counters matched exactly and default-off remained zero bytes/work.

    The moving-light workload recomputes coverage every 30 frames. Three-run medians all improved:

    • wall 14.385408 -> 12.873031 ms;
    • shadow CPU 0.218197 -> 0.181323 ms;
    • page CPU 0.044228 -> 0.034917 ms;
    • shadow GPU 2.969117 -> 2.864068 ms;
    • page GPU 2.541226 -> 2.505326 ms.

    Full quick lane passed: Linux/Web/all-platform FFI parity, strict Clippy, 342 shared tests + 1 ignored, headless device, 59 GPU goldens + 2 policy ignores, 4 render-target tests, WASM, quality/cooker tooling, and 20 examples. Focused VSM: 35/35.

    Evidence: issue-132-request-compaction-v1.md and JSON.

    This is the bounded CPU oracle and storage ABI, not a claim that GPU marking is complete. Next: capability/workload-gated compute marking and compaction without same-frame blocking readback, then GPU page-caster culling/submission.

  8. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Async GPU receiver-marking checkpoint is implemented and pushed in PR #147 (ad6e7c6, evidence/docs db61815).

    Delivered boundary:

    • capable native adapters use compute marking only for 1,024–4,096 camera-visible receiver bounds;
    • smaller scenes retain the fixed CPU oracle with zero marker resources, dispatches, copies, or readbacks;
    • resources and pipeline are lazy;
    • two 12,288-byte readbacks are mapped asynchronously; production uses Poll, never a same-frame Wait;
    • first GPU output must exactly equal the complete ordered CPU demand or the backend disables itself and retains CPU demand;
    • projection changes stay exact synchronous CPU transitions; continuous same-projection motion consumes the newest completed result, with missing addresses safely using CSM.

    Qualification:

    • real Metal dispatch test: complete 1,024-receiver GPU result exactly equals the CPU oracle;
    • 1,140 moving receivers: 36 dispatches, 35 completions, validated=true, 0 validation failures, 102,896 lazy bytes;
    • GPU vs CPU image: RMSE 0.000018280, SSIM 1.0, 0% above tolerance; cache/fallback counters match;
    • paired large-workload CPU frame 18.615338→18.306413 ms, GPU frame 40.886617→37.220795 ms, shadow CPU 0.637508→0.386642 ms;
    • the experimental 256 threshold was rejected after GPU timestamp inflation at ~285 receivers, so that workload remains CPU-only;
    • default-off and ordinary two-receiver fixtures remain effectively identical with zero marker work;
    • complete quick lane passed: all-platform FFI parity, strict lint, 345 shared unit tests + headless device, 59 GPU goldens, 4 render-target tests, WASM, quality/cooker, and 20 examples.

    Detailed evidence: https://github.com/Bloom-Engine/engine/blob/codex/rendering-wip-integrated-20260727/docs/evidence/issue-132-async-gpu-receiver-v1.md

    Architecture status (kept deliberately precise):

    • bounded, capability/workload-gated GPU coverage marking without blocking readback
    • GPU-resident request compaction and page-residency scheduling (next; current dense result is compacted by the exact CPU oracle after async readback)
    • GPU page-caster culling and submission
  9. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Correction: GPU receiver backend retained explicit opt-in

    Follow-up direct pass instrumentation invalidated the automatic performance qualification in the preceding comment. The correction is pushed in PR #147 at 9fe0371 (default behavior) and b05c8b8 (evidence/docs).

    What changed:

    • production now defaults to the fixed CPU receiver oracle;
    • GPU marking requires explicit BLOOM_VSM_GPU_RECEIVER=1 in addition to the existing capability and 1,024–4,096-bound gates;
    • the failed GPU rank/compact prototype was removed before commit;
    • ordinary/default runs allocate 0 marker bytes and issue 0 marker dispatches;
    • the exact GPU-vs-CPU validation and safe async fallback remain available as an experimental architecture oracle.

    Why the earlier total-frame conclusion was rejected:

    • the earlier paired run showed a 0.250866 ms shadow-CPU reduction, but its window-server-throttled total GPU/wall values could not attribute marker cost;
    • direct timestamps measured the unchanged dense GPU mark at 0.485908 ms;
    • experimental rank and compact dispatches added 0.436717 ms and 0.504675 ms; combining the three in one compute pass still measured 1.510075 ms;
    • that is not an across-the-board net win, so there is no qualified automatic workload crossover.

    Verification after the correction:

    • focused VSM suite: 38/38, including exact real-Metal GPU demand equality;
    • production example with the 1,140-bound stress fixture and no GPU env opt-in reports receiver_marking_backend: fixed-cpu, gpu_receiver_enabled: false, 0 dispatches, and 0 bytes.

    Corrected evidence: https://github.com/Bloom-Engine/engine/blob/b05c8b8/docs/evidence/issue-132-async-gpu-receiver-v1.md

    Issue status remains precise: fixed CPU request compaction is complete; automatic GPU request compaction/residency and GPU page-caster submission remain open and must pass direct pass-level timing before activation.

  10. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Implemented and qualified the next milestone in draft PR #147 at code revision 7962bee (evidence/docs at 87b9d2e).

    Completed slice: bounded rigid-opaque page indirect submission

    • Directional VSM pages with at least 48 visible eligible draws now submit compact multi_draw_indexed_indirect ranges.
    • Classification remains the exact CPU page-frustum oracle; the implementation performs a single partition pass and precomputes one clip matrix per caster.
    • Only static rigid opaque geometry in the shared geometry arena is eligible. Cutout, skinned, foliage-motion, dynamic-overlay, dedicated-buffer, overflow, small, unsupported, Web, and GPU-driven-disabled paths remain on the compatibility renderer.
    • Lazy allocation: 51,200 bytes in the 512-caster fixture; small VSM, VSM-off, GPU-driven-off, and BLOOM_VSM_GPU_CASTERS=0 controls allocate zero bytes.
    • Telemetry is exposed at renderer_paths.vsm_gpu_casters.

    Qualification

    Five interleaved 60-warmup / 180-measured-frame runs on Apple M1 Max / Metal:

    • wall-frame mean median: 11.893637 -> 11.457747 ms (-3.66%)
    • CPU-frame mean median: 9.824165 -> 8.851101 ms (-9.90%)
    • GPU-frame mean median: 20.292364 -> 19.262754 ms (-5.07%)
    • VSM-page CPU: 0.106694 -> 0.094448 ms (-11.48%)
    • VSM-page GPU: 1.092336 -> 0.582765 ms (-46.65%)
    • render-total CPU: 4.921546 -> 4.435739 ms (-9.87%)

    CPU-control and indirect captures are byte-identical across 3,686,400 pixels: RMSE 0, SSIM 1, max error 0, OKLab delta 0, edge delta 0. The full scripts/ci-check.sh --quick lane passes, including FFI parity, Linux/Web builds/checks, 347 shared tests, 59 GPU goldens, WASM check, file-size ratchet, and quality/tooling gates.

    Evidence: https://github.com/Bloom-Engine/engine/blob/codex/rendering-wip-integrated-20260727/docs/evidence/issue-132-vsm-caster-indirect-v1.md

    Explicit remaining boundary

    This does not claim full GPU caster culling. A Metal fixed-count Cartesian prototype and a compute safety-cull prototype were rejected before commit. Remaining milestone-3 work is true GPU classification + indirect-count compaction on qualified backends, plus separate cutout/skinned/foliage/dynamic/instanced/dedicated-buffer qualification. Spot/point projections, shared arbitration, broader scenes, and tier rollout also remain open.

  11. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Completed and checked the geometric-contact-detail acceptance criterion in draft PR #147.

    Code/tool checkpoint: c1131d0; evidence/docs: 3050fd7.

    Automated VSM-vs-CSM oracle

    • New opt-in quality-motion --vsm-contact-detail fixture: 266 colored 0.055 m posts over a neutral receive-only ground, fixed camera/light/exposure, 120 warm-up frames so all 224 demanded VSM pages are resident and clean.
    • New dependency-free tools/quality/shadow_detail.py gate evaluates the same neutral-ground pixel intersection in both images and excludes colored geometry/sky/TAA chroma edges.
    • Negative controls prove sharp-vs-blurred passes, the reversed candidate fails, and colored geometry is excluded.

    Five deterministic captures per mode:

    • strong shadow-edge pixels: CSM 161, VSM 8,110 (50.37x)
    • edge p99: 0.071697 -> 0.116095 (1.619x)
    • 95th-to-5th percentile shadow contrast: 0.284005 -> 0.291565 (+0.007560)
    • all VSM captures shared one SHA-256; all CSM captures shared one SHA-256

    Measured incremental allocation is explicit: 19,951,824 VSM pool/metadata bytes + 51,200 caster-indirect bytes = 20,003,024 bytes. Five-run medians show GPU-frame mean effectively flat (+0.014120 ms / +0.03%), GPU p50 -2.90%, wall -2.71%, and CPU mean -5.82%. The VSM sampling cost itself is visible (main_hdr_pass +0.231049 ms); lower whole-frame totals are not claimed as an optimization.

    This checkpoint changes no production renderer behavior: it adds only an opt-in example branch, offline semantic gate, tests, and evidence. The full quick lane passes.

    Evidence: https://github.com/Bloom-Engine/engine/blob/codex/rendering-wip-integrated-20260727/docs/evidence/issue-132-contact-detail-v1.md

    Remaining #132 acceptance work is still Bistro/seam/ring qualification, alpha/skinned automation, constrained fallback qualification, debug-view completeness, and local-light VSM/arbitration.

  12. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Forced lower-tier VSM fallback complete

    Implemented in c34743e; evidence/docs are pushed at d306164 in draft PR #147.

    What changed:

    • Renderer startup records the accepted capability tier before VSM shader/layout/resource selection.
    • BLOOM_VSM=1 on forced baseline or modern now compiles canonical CSM and reports lower-tier-csm-fallback.
    • Telemetry separates requested, capability_eligible, enabled, active, and selection_reason.
    • The public capability report exposes the same state under runtime_support.virtual_shadows.
    • The capability table now states CSM fallback for baseline/modern and VSM-with-CSM-fallback for high-end.

    Qualification:

    • Before fix: forced baseline still activated 256 VSM pages and allocated 19,951,824 bytes.
    • After fix: 120/120-frame forced-baseline request is byte-identical to the CSM control at 2560x1440; both SHA-256 596a8b611496b56b9f2e5747f95b706effa2b85e136e84e481afb1d0bc987d6f.
    • Lower-tier VSM pool, page work, receiver work, and caster-indirect work are all zero.
    • Forced modern also exactly matches its CSM control.
    • High-tier VSM retains the previously qualified SHA-256 81c6d840a275ee4fe3ff9fbeb87a5a1e21b23bb1c087a4e24c024c2f778b0bde with all 224 pages resident and clean.
    • Five counterbalanced timing pairs show no shadow-path cost: shadow CPU 0.087858 -> 0.088146 ms (+0.000288 ms, noise); no timing improvement is claimed.
    • Full quick lane passes: 348 shared tests + 1 ignored, headless device, 59 GPU goldens + 2 policy ignores, 4 render-target tests, strict Clippy, FFI parity, file ratchet, WASM, quality tooling, and 20 examples.
    • Supported Web crate, iOS aarch64, and Android aarch64 checks pass.

    Evidence: https://github.com/Bloom-Engine/engine/blob/d306164/docs/evidence/issue-132-capability-fallback-v1.md
    JSON: https://github.com/Bloom-Engine/engine/blob/d306164/docs/evidence/issue-132-capability-fallback-v1.json

    The forced fallback acceptance box is now checked. Remaining acceptance work is fixed Sponza/Bistro motion quality, alpha/skinned automation, complete debug views, and 100 local shadow-requesting lights/shared arbitration.

  13. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Completed the alpha-tested foliage + skinned caster acceptance checkpoint in 862981f; evidence is recorded in 3015dcb.

    What this adds:

    • VSM telemetry now reports alpha-tested and skinned page draws independently, so an unloaded/missing caster cannot silently pass.
    • quality-motion --vsm-alpha-no-cast is a negative control that keeps Sponza's MASK curtain visible while removing only its shadow submission.
    • tools/quality/vsm_caster_coverage.py compares that control and an 8-frame Fox pose offset against fixed ground ROIs. It fails closed on missing/malformed telemetry, unchanged shadows, weak coverage, or opaque replacement coverage.

    Real Apple M1 Max / Metal result at 1600x900, 120 warm-up + 120 measured frames:

    • Alpha-tested shadow: 19,931 / 71,280 ROI pixels changed (27.962%), 4.646 segments/occupied row, p95 luminance delta 0.178709.
    • Skinned shadow: 4,465 / 51,840 ground-only ROI pixels changed (8.613%), p95 luminance delta 0.201384.
    • Full telemetry: 4 cutout + 4 skinned page draws, 4 rendered pages / 4-page budget, 8 draws / 64-draw budget, 119 deferred pages covered by live CSM fallback, 19,951,824 VSM bytes.
    • Repeat variance was bounded to a few 1-code-value pixels: max 1/255, 0% above 0.02, SSIM 1.0.

    No render pass, shader, resource, binding, page, or draw was added. The only production-path work is two integer classifications inside the already opt-in VSM page draw loop; default CSM behavior is unchanged.

    Regression result: complete quick lane passed (348 shared + 1 ignored, 59 GPU goldens + 2 policy ignores, 4 render-target tests, 25 quality tests, strict Clippy, FFI, ratchet, Web/WASM, and 20 examples).

    Evidence: https://github.com/Bloom-Engine/engine/blob/3015dcb/docs/evidence/issue-132-caster-coverage-v1.md

    I marked only this acceptance checkbox complete. The fixed Sponza/Bistro motion corpus, 100 shadow-requesting local lights, and complete debug views remain open.

  14. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Completed the fixed Sponza/Bistro camera-path acceptance checkpoint.

    Implementation:

    • 2ee8903 adds opt-in 30-frame light-plane camera paths to both established examples plus tools/quality/vsm_motion_corpus.py.
    • aa36327 makes all fixture bookkeeping/path math conditional on the opt-in flag.
    • 91e7843 records the exact captures and machine-readable evidence.

    The oracle uses four matched captures per scene (settled VSM/CSM + motion VSM/CSM) and evaluates (motion VSM - motion CSM) - (settled VSM - settled CSM). This removes shared auto-exposure/TAA/SSGI/SSR history and isolates only VSM transition behavior.

    Apple M1 Max / Metal results at 1600x900:

    • Sponza: residual RMSE 0.002120, p99 0.008394, 0.0052% pixels above 0.03, largest connected artifact 20 pixels, no seam/ring-like component.
    • Bistro: residual RMSE 0.004746, p99 0.018491, 0.4667% pixels above 0.03, largest connected artifact 225 pixels (0.0156% of frame), no seam/ring-like component.
    • Missing-shadow/stale-shadow ratios stayed below 0.1% in both scenes.
    • Sponza transition: 2 rebases, 228 pages preserved, zero dropped/dirty/rendered/denied/evicted pages.
    • Bistro transition: 3 rebases, 202 preserved + 54 safely dropped, 8 bounded renders, 26 dirty pages excluded from sampling, zero denials/evictions.
    • Both retained the fixed 19,951,824-byte allocation.
    • Independent repeat SSIM: 0.999997 Sponza, 0.999954 Bistro, with identical cache/work counters.

    Six seeded failures prove page seams, clipmap rings, missing-page flashes, stale/doubled shadows, missing rebases, and malformed telemetry fail closed. A seventh test proves matched controls remove unrelated temporal/exposure history.

    No renderer/pass/shader/resource/API behavior changed in this checkpoint. The paths and offline oracle are opt-in. The exact pushed state passed the full quick lane: FFI, strict Clippy, ratchet, Web/WASM, 348 shared tests, 59 GPU goldens, 4 render-target tests, 32 quality tests, and 20 examples.

    Evidence: https://github.com/Bloom-Engine/engine/blob/91e7843/docs/evidence/issue-132-full-scene-motion-v1.md

    Scope note: Sponza is the complete checked-in asset. Bistro is the governed 96 unique-mesh camera-visible subset from the pinned source, not a claim of full 2,909-instance streaming coverage.

    I marked only the fixed full-scene motion checkbox complete. The 100+ shadow-requesting local-light and complete debug-view criteria remain open.

  15. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Completed the full debug-view acceptance criterion in 840eb49 with evidence in 20ac21e.

    What is now exposed by an opt-in intermediate capture:

    • virtual-shadow-pages.png: the three 32×32 virtual clip levels, stacked near-to-far;
    • virtual-shadow-physical.png: physical pool occupancy in slot order;
    • distinct amber never-rendered misses and magenta previously-rendered invalidations;
    • stable near/middle/far clip-level colors plus virtual-shadow-legend.png;
    • a same-frame virtual-shadow-report.json containing the debug contract and a directional per-light cost row (requests/hits/misses/invalidations/renders/residency/dirty/rebases/draws/owned and shared bytes/budget).

    The fail-closed validator (tools/quality/vsm_debug_views.py) requires exact palette and geometry, whole cells, matching virtual/physical page counts, report-consistent occupancy/dirty/clip levels, and reconciled per-light counters and bytes.

    Real Apple M1 Max / Metal captures passed:

    • early fill: 212 visible never-rendered misses;
    • light transition: 102 visible invalidations plus 126 never-rendered misses, with 236 resident / 228 dirty pages;
    • settled static scene: 144/64/16 clean pages across levels 0/1/2 and zero dirty pages.

    This is capture-only and changes no shader, pass, GPU resource, binding, draw, page policy, sampling, or fallback. scripts/ci-check.sh --quick passes (349 shared unit tests + 1 ignored, device negotiation, 59 GPU goldens + 2 policy ignores, 4 render-target tests, Web/WASM, 39 quality tests, FFI parity, and 20 examples).

    Evidence: docs/evidence/issue-132-debug-views-v1.{md,json}.

    With this checkbox complete, #132 is 7/8; only the 100 shadow-requesting local-lights acceptance criterion remains.

  16. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Point-light portion of the final local-light acceptance criterion is now qualified on PR #147.

    Implemented and pushed:

    • a01cf8a — bounded shadow-required point-light API, deterministic visibility/admission, six-face shared VSM pages, retained/immediate/froxel sampling, fail-closed suppression, telemetry, 128-light fixture, and validator.
    • 630cb23 — uniform no-request early-out before dynamically indexed local metadata.
    • 93ebdc3 / 6a70ca7 — test-only source split required by the file-size gate.
    • e27793e — checked-in human- and machine-readable evidence.

    Measured on Apple M1 Max / Metal:

    • 128 submitted and visible; 5 admitted; 123 budget-suppressed.
    • 30 local pages maximum (5 x 6), sharing the existing 256-page pool and existing 8-page/frame render budget.
    • Cold: exactly 8 local pages rendered; pending lights remained zero-contribution; dirty directional pages used CSM fallback.
    • Warm: all 30 local pages clean, all 5 admitted lights active, zero page renders.
    • Controlled five-light shadowed/unshadowed A/B changed 0.858% of pixels, localized behind occluders (RMSE 0.01009, SSIM 0.99437); visual review found no cube seams, page blocks, missing bands, or frame-wide shift.
    • Both >=100-light telemetry reports pass the fail-closed validator.
    • Exact pushed head passed scripts/ci-check.sh --quick: 355 shared tests, 59 GPU goldens, 4 render-target tests, Web/WASM and all-platform FFI parity, 39 quality tests, visual metrics, asset cooker, and 20 canonical examples.

    Evidence:

    • docs/evidence/issue-132-local-lights-v1.md
    • docs/evidence/issue-132-local-lights-v1.json

    I marked the >=100 local-light criterion complete. I am leaving #132 open because its architecture/milestone text still explicitly includes a spot-light virtual projection; this delivery qualifies bounded point-light cube projections, not spot lights.

  17. proggeramlug commented on Jul 29, 2026

    @proggeramlug
    ContributorAuthor

    Completed the final spot-light milestone and closing #132.

    Pushed on PR #147:

    • 0bcd2c7 — public addShadowedSpotLight(...), one-page perspective projection, circular smooth cone sampling, deterministic shared-pool admission, fail-closed suppression, all-platform/Web ABI handling, telemetry, fixture, and validator.
    • d91d9c6 / 69b7403 — human/machine-readable evidence and follow-up-scope documentation.

    Apple M1 Max / Metal qualification:

    • 128 submitted and visible; 5 admitted; 123 suppressed before retained/immediate shading and froxel assignment.
    • Exactly 1 page per admitted spot and 5 local pages total, versus 6/30 for point cubes; the existing 256-page pool and 8-page render budget are reused.
    • Warm state: all 5 spots active, 5/5 pages clean, zero page renders, zero denials.
    • Same-scene/same-cone caster-on/off oracle: RMSE 0.01644, SSIM 0.98078, 2.027% pixels above 0.02. Differences are localized to occluder silhouettes; visual review found no page rectangles, projection seams, missing bands, cone discontinuities, or frame-wide shift.
    • No persistent GPU bytes or graph passes added over the point milestone. Directional-only requests retain the uniform early-out.
    • scripts/ci-check.sh --quick passes: 356 shared tests + 1 ignore, headless device construction, 59 GPU goldens + 2 hardware-policy ignores, 4 render-target tests, FFI/Web/Wasm, strict Clippy/format/ratchet, quality/visual/cooker suites, and all examples. The ordinary immediate and clustered many-point-light goldens pass exactly. The native Web crate check also passes.

    Evidence:

    • docs/evidence/issue-132-spot-lights-v1.md
    • docs/evidence/issue-132-spot-lights-v1.json

    All acceptance boxes are complete, including the architecture milestone requiring both point and spot projections. GPU-only request scheduling and broader indirect specialization remain documented optional follow-ups rather than acceptance blockers.

  18. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Post-closure Bistro shadow follow-up is pushed in PR #147.

    • 6da56ae selects scene-receiver cascades from the same positive view-space depth used by frustum fitting instead of spherical camera distance.
    • 06be86e applies the same selection contract to shared/material receivers.
    • Both paths have source regressions, and the targeted shadow_cascade test set passes (3/3).
    • Fresh-start and same-process camera-rotation captures converge to the same world-space shadow footprint, ruling out stale cascade-cache state.
    • A CPU path-traced render reproduces the broad building-shadow silhouette. Forcing all receivers to cascade 1 leaves the raised paving-island region visually unchanged (mean A/B delta below one 8-bit RGB level), confirming that region is authored material/ambient contrast rather than missing cascade coverage.
    • Transmitted shadows were disabled in the opaque CSM A/B, so they are not masking the result.

    The clean production-shader Bistro binary was rebuilt after removing the diagnostic override. CI cleanup e7bd0c1 also restores the PR contracts/lint lanes locally.

  19. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Movement/rotation follow-up landed in 8d85719 (fix(render): preserve shadows across cascade handoff).

    Root cause addressed: cross-cascade blending could sample the next cascade even when its retained translation-slack fit did not cover the receiver. The core path then read a clamped edge texel, while the material path treated the miss as fully lit; either path could erase part of an otherwise valid shadow as the camera crossed a split.

    The fix makes the handoff coverage-aware in both paths: an out-of-fit next cascade preserves the current cascade result until the next fit actually covers the receiver. It adds no shadow texture samples and leaves the valid overlap path unchanged. Added regression guards for both shader families.

    Verification complete: 4 targeted cascade tests green; native/folded WGSL parse green; quick contracts/FFI parity green; strict rendering lint green; macOS release build and Bistro binary rebuilt. Translation-only cache-vs-always-fresh A/B is effectively identical away from overlay text, ruling out stale depth for that leg.

    Still pending: human traversal approval for the reported combined move+turn path. The prior open Bistro window predated 8d85719. Machine free space fell below 10 GiB due other active builds, so further compiler-driven oracle expansion is paused rather than risking the workspace.

  20. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Final moving-camera cascade validation

    The post-closure cascade-handoff regression is now resolved at 67661ee and independently human-validated in the interactive Bistro scene.

    Validation covered the reported failure mode rather than only the launch pose: moving and turning around the restaurant/street receiver while the directional shadow crossed cascade-fit boundaries. The previously missing rectangular/slab regions remained filled, and the reporter confirmed: “perfect, shadows work now.”

    Regression evidence:

    • selected-cascade projection misses now hand off to the next valid cascade instead of returning fully lit;
    • the ordinary in-fit path still performs one cascade depth sample;
    • all 55 shadow-focused library tests pass;
    • deterministic ordinary in-fit captures remain pixel-identical apart from the timing overlay;
    • follow-up cc4b663 moves byte-identical 2D WGSL out of core.rs (matching SHA-256) and restores the PR file-size gate without changing shader output or runtime work.

    This completes the human movement-validation boundary called out in the prior follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions