Skip to content

perf ticket 008: visibility buffer (Nanite-style) replaces 4-MRT G-buffer #27

Description

@proggeramlug

Parent roadmap: #126
Design/qualification contract: docs/perf/008-visibility-buffer.md
Active implementation: PR #147

Corrected objective

Qualify a Nanite-style visibility-buffer composition path without assuming a bandwidth or overdraw win. Bloom already has an alpha-aware depth prepass, so the comparison must include depth, visibility raster, reconstructed PBR shading, compatibility rendering, all MRT writes, memory, and frame total.

The packed target is Rg32Uint (8 bytes/pixel), not Rgba32Uint (16 bytes/pixel). It stores a 32-bit draw ID plus a 31-bit primitive ID and front-face bit. Perspective-correct barycentrics are reconstructed from the referenced triangle rather than stored.

Landed architecture

  • Shared CPU/WGSL ID, face, barycentric, vertex/index, first-index, and base-vertex contract.
  • Exact 96-byte Vertex3D storage decoder using six packed vec4<u32> lanes.
  • Real depth-equal visibility raster against GPU-driven indirect geometry.
  • Full-screen shading reconstructs the production VertexOutputScene and calls the exact authoritative shade_main_scene; material, clustered lighting, shadows, IBL, velocity, and MRT logic are not forked.
  • Static fully opaque Tier-A shared-arena draws are eligible. Cutout/masked, blend, transmission, layered/custom, skinned, deforming, and unsupported draws remain explicitly forward-compatible.
  • BLOOM_VISIBILITY_BUFFER=validate|debug|shade selects qualification modes before device creation. Default/off requests no PRIMITIVE_INDEX, creates no visibility pipeline/resource, and records no work.
  • Shade mode owns one Rg32Uint target (8 bytes/pixel). Validate/debug also own an 8-byte diagnostic reconstruction target.
  • Capability telemetry exposes request state, activation reason, composition owner, eligible/compatibility counts, extent, bytes, and current-frame recording.

Verified evidence

  • Packed ID/front-face ABI and background sentinel tests.
  • CPU/GPU perspective reconstruction oracle, including non-zero first index and base vertex.
  • All 24 packed vertex lanes reconstructed from real storage buffers.
  • Real GPU visibility raster/reconstruction runtime test.
  • Real GPU full-PBR shading test where eligible forward fragments are deliberately suppressed.
  • Compatibility composition test includes a layered-PBR draw.
  • Deterministic off-vs-shade image test: 89 of 81,920 channels differ, maximum 1 LSB, mean delta 0.00108643.
  • Shade steady state creates no new texture or bind group after initialization.
  • Default/off path has zero allocation/work and preserves the shipping renderer.
  • Strict lint, WebAssembly check, contracts, 376 unit tests, 59 golden tests, render-target tests, and visibility GPU tests pass locally on Metal.

Remaining activation gates

  • Separate/compact eligible and compatibility indirect command streams so eligible geometry is not still dispatched through the forward compatibility shader. The current opt-in correctness path may cost more GPU time and must not ship as default.
  • Capture total uncapped GPU time for depth_prepass, visibility_raster_pass, visibility_pbr_pass, main_hdr_pass, and frame total against forward/off.
  • Prove all individual MRT outputs—not only the final screenshot—match for HDR, material, velocity, and albedo.
  • Run identical-camera Bistro and stress-scene corpora with SSR, SSGI, SSAO, TAA, VSM/cascades, planar probes, custom opaque materials, cutouts, skinning, refraction, and transparency.
  • Demonstrate no regression on low-overdraw scenes and a material improvement on a representative admitted workload.
  • Qualify at least one discrete and one integrated/tile-based adapter; retain fallback for adapters without PRIMITIVE_INDEX.
  • Decide whether masked/cutout geometry remains compatibility-only or gains an alpha-aware visibility raster without changing silhouettes/order.
  • Enable outside explicit qualification only after every visual, memory, and total-GPU gate passes.

Next implementation slice

  1. Extend GPU culling/compaction to route visibility-eligible and compatibility draws into distinct bounded indirect streams and counters.
  2. Make the main forward pass submit only compatibility commands in shade mode; preserve today’s single stream byte-for-byte in off/validate/debug modes.
  3. Add CPU-oracle tests for routing, zero-eligible/all-eligible/mixed counts, fixed-count fallback, and capacity/steady-state invariants.
  4. Add per-pass and frame-total A/B artifacts, followed by deterministic Bistro camera-corpus captures and individual-MRT readback gates.

Likely files

  • native/shared/src/renderer/gpu_driven.rs
  • native/shared/src/renderer/visibility_buffer.rs
  • native/shared/src/renderer/visibility_shading.rs
  • native/shared/src/renderer/scene_pass.rs
  • native/shared/tests/visibility_buffer_*.rs
  • tools/quality/ for governed Bistro/MRT/timing comparison

Non-goals for activation

  • Claiming a win from target byte size alone.
  • Replacing compatibility paths before they have equivalent semantics.
  • Enabling the path by default while total GPU time or any downstream buffer is unproven.

Activity

  1. proggeramlug commented on Jul 25, 2026

    @proggeramlug
    ContributorAuthor

    2026-07-25 dependency audit after #28 landed:

    The full visibility-buffer proposal is not currently safe to enable under the project no-regression rule. Two assumptions in the deferred design are now stale:

    1. Rgba32Uint is 16 bytes/pixel, not ~8 bytes/pixel. Replacing an 18-byte MRT raster with that target reduces the visibility raster write by only 11%, before the separate HDR/G-buffer shading pass is counted. An actually 8-byte encoding would require packed IDs/barycentrics and explicit mesh/primitive limits.
    2. EN-044 has since landed an alpha-aware depth prepass plus Equal-tested main shading. The expensive PBR/MRT fragment already runs only for the visible layer, so a visibility raster followed by random storage-buffer reconstruction adds a pass and geometry fetch cost rather than removing the overdraw assumed by this ticket.

    The safe intermediate described by the ticket is already present: lean_mrt drops material and albedo attachments on Android/constrained builds where their consumers are disabled. #28 also supplies shared STORAGE-capable vertex/index arenas and GPU submission, so the architectural prerequisite remains available when a measured target workload justifies a packed visibility experiment.

    Decision: leave #27 open/deferred; do not force-enable the proposed path without an A/B showing total depth + visibility + shading GPU time and bandwidth are net-positive. Reopen implementation on a bandwidth capture demonstrating that threshold, and use a packed 8-byte format plus compatibility routing for oversized/skinned/translucent/custom geometry.

  2. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Issue #134 now has a checked 63-row layered-PBR angular corpus plus forward/ray-query/oracle parity and clearcoat-normal motion gates in PR #147. The only #134 acceptance item still blocked is visibility/forward/path parity because Bloom remains forward-MRT today. When this visibility-buffer work adds its material evaluator, please make it consume tools/bloom-reference/reference/layered-pbr-angular-v1.json and satisfy the documented parity contract in docs/evidence/issue-134-layered-pbr-qualification.md; that makes #27 the explicit closure dependency rather than duplicating a new oracle.

  3. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Milestone: packed visibility ABI and no-regression activation contract

    Pushed in 0d2442d on PR #147.

    This corrects the stale baseline design before production integration:

    • Rgba32Uint is 16 bytes/pixel, not ~8; the v1 contract uses Rg32Uint at exactly 8 bytes/pixel.
    • Word 0 is a full draw ID. Word 1 is a 31-bit primitive ID plus one front-face bit; u32::MAX draw ID is the unambiguous background sentinel.
    • Perspective-correct barycentrics are reconstructed from the referenced triangle clip positions, so shared indexed geometry does not need expansion or a per-corner barycentric stream.
    • CPU and WGSL constants/math are checked together, including maximum-ID round trips, sentinel collision rejection, known perspective weights, parser validation, and exact allocation accounting (16,588,800 bytes at native 1080p before backend alignment).
    • Renderer diagnostics now report the contract as shipping_enabled:false and opt-in A/B qualification required.
    • The design document now requires total depth + visibility + shading timings and full transient memory—not format-size inference—before activation. Static opaque/masked Tier-A shared geometry is the candidate path; blend, transmission, layered/custom, skinned/deforming, and unsupported content remains explicit forward compatibility.

    Regression evidence:

    • 366 shared unit tests passed, 1 intentionally ignored;
    • negotiated production renderer construction passed;
    • 59 GPU goldens passed, 2 hardware-policy tests intentionally ignored;
    • all 4 render-target tests passed;
    • strict formatting/Clippy, FFI/schema parity, file ratchet, 39 quality-governance tests, 3 visual fault controls, and 29 cooker tests passed.

    Shipping impact is deliberately zero: no visibility texture, pipeline, pass, draw, binding, shader branch, allocation, or per-frame CPU work exists yet. The next checkpoint is an opt-in diagnostic raster/readback oracle that proves draw ID, primitive ID, front-face, clipping, and reconstructed attributes against a CPU oracle before any PBR composition is attempted.

  4. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Milestone: hardware visibility-raster oracle

    Pushed in 2dffe29 on PR #147.

    The diagnostic-only GPU test now rasterizes two triangles with opposite winding and non-uniform clip-space W into the packed Rg32Uint target, reads the result back, and proves:

    • draw ID and primitive ID survive rasterization exactly;
    • the packed front-face bit distinguishes opposite winding;
    • untouched pixels retain the background sentinel;
    • perspective-correct barycentrics reconstructed by the shared WGSL match an independent CPU oracle per covered pixel (2e-5 tolerance);
    • the required WebGPU PRIMITIVE_INDEX capability is explicitly requested and unsupported adapters skip the hardware oracle instead of failing unrelated tiers.

    Qualification is green: strict fmt/Clippy, contracts/file ratchet, 367 shared tests + negotiated-device construction, 59 GPU goldens, 4 render-target tests, and the quality governance/fault/cooker suites. Shipping impact remains zero: this adds no runtime texture, pass, allocation, shader branch, or frame work.

    Next checkpoint: validate vertex/index pulling and interpolated attributes against the existing shared GpuDrawRecord/geometry arenas. Only after that should an opt-in runtime A/B pass be wired; forward remains authoritative until image parity and total GPU time/memory are net-positive.

  5. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Milestone: shared geometry and attribute reconstruction oracle

    Pushed in cdc50c6 on PR #147.

    This closes the last ABI risk before an opt-in runtime pass:

    • Added a shared WGSL decoder for Bloom’s exact 96-byte Vertex3D layout. It uses six packed vec4<u32> lanes because native WGSL vec3 storage alignment would corrupt the existing 12-byte field offsets.
    • The Metal hardware oracle reads the actual Rust Vertex3D, GpuDrawRecord, and index-buffer bytes through storage bindings.
    • It deliberately exercises non-zero first_index and base_vertex, validates all 24 words of all three selected vertices byte-exactly, then checks all 24 perspective-interpolated lanes against an independent CPU oracle.
    • It also proves visibility draw/primitive/face decoding and material-ID lookup through the real 272-byte draw-record ABI.
    • Diagnostics now name primitive-index as the required capability and report the checked 96-byte vertex stride.
    • The design document no longer contains the stale “four MRT writes per overdrawn fragment” or deferred/reopen guidance; the current alpha-aware depth prepass and total-pass A/B gate are authoritative.

    Qualification: strict fmt/Clippy and contracts/file ratchet pass; 368 shared tests pass (1 ignored), negotiated-device construction passes, 59 GPU goldens pass (2 hardware-policy ignores), and all 4 render-target tests pass. Shipping remains unchanged: no runtime pass, texture, allocation, or frame work has been enabled.

    Next milestone: wire the same proven ABI into an explicit opt-in runtime visibility raster + reconstruction diagnostic, while keeping forward rendering authoritative and measuring full image/performance/memory deltas before any activation.

  6. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Milestone: opt-in production visibility raster and reconstruction runtime

    Pushed in 232dece on PR #147.

    This wires the proven packed ABI into the production renderer without changing the default image:

    • BLOOM_VISIBILITY_BUFFER=validate records a private depth-equal Rg32Uint visibility raster and compute reconstruction against the real GPU-driven draw, index, and vertex buffers while forward remains authoritative.
    • BLOOM_VISIBILITY_BUFFER=debug additionally overlays reconstructed normals only on admitted pixels, so compatibility holes and routing mistakes are visible.
    • The unset/off path requests no PRIMITIVE_INDEX feature, creates no visibility pipeline, texture, or bind group, and records no visibility GPU work.
    • Static opaque and masked shared geometry is admitted. Wind/deforming, skinned, layered, blend, transmission, custom, and unsupported content remains on the explicit forward compatibility path.
    • The raster uses the exact production GPU-scene vertex shader and Equal-tests the alpha-aware prepass depth, preserving clipping, transforms, mask coverage, winding, and the existing draw identity.
    • Public capability telemetry reports requested mode, activation reason, eligible and compatibility draw counts, extent, exact owned bytes, debug state, and whether the current frame actually recorded visibility work.
    • Retained GPU draw flags now preserve double-sided material semantics instead of silently treating every retained node as single-sided.

    The new real-GPU integration test constructs the renderer through production device negotiation, submits 32 retained shared-geometry draws, reads back the composed frame, and proves meaningful reconstructed-normal coverage. A second steady frame proves zero visibility texture and bind-group creations. A third empty frame proves the previous reconstruction is never replayed as a stale overlay.

    Qualification is green:

    • strict formatting and Clippy policy;
    • FFI/schema parity on macOS, Linux, Windows, Android, iOS, tvOS, watchOS, and web;
    • file-size ratchet;
    • wasm32 web compilation;
    • 370 shared unit tests passed, 1 intentionally ignored;
    • production device negotiation passed with no primitive-index request by default;
    • 59 golden-render tests passed, 2 hardware-policy tests intentionally ignored;
    • 4 render-target tests and the new Metal runtime integration passed;
    • 39 quality-governance tests, 3 visual fault controls, and 29 cooker tests passed.

    This remains a diagnostic milestone, not the shipping visibility renderer. The next milestone is the full material and lighting evaluator for eligible pixels, followed by Bistro/reference A/B image parity and total depth + visibility + shading time/memory qualification before any default activation.

  7. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Pushed the full-PBR visibility composition checkpoint in PR #147 as f7422a2.

    Key result: the new opt-in shade mode reconstructs production fragment inputs and invokes the exact existing PBR evaluator. A process-isolated off-vs-shade GPU capture differs in only 89 of 81,920 channels, all by one LSB (mean_delta=0.00108643). A layered-PBR compatibility draw remains present and forward-rendered, and steady state creates no new resources.

    The full local gate is green: strict lint, contracts, WebAssembly, quality governance, 376 unit tests, 59 goldens, four render-target tests, plus runtime/reconstruction/shading/parity GPU tests.

    I rewrote the issue body with the corrected 8-byte format, landed architecture, evidence, and remaining activation work. The next slice is command-stream separation/compaction plus total GPU and Bistro/MRT A/B qualification; the path remains explicitly off by default until those no-regression gates pass.

  8. proggeramlug commented on Aug 8, 2026

    @proggeramlug
    ContributorAuthor

    Visibility routing checkpoint pushed in c7b2259.\n\nWhat changed:\n- shade mode now keeps the original all-draw indirect stream for the shared depth prepass;\n- the cull compute pass also writes slot-preserving visibility-eligible and forward-compatibility streams;\n- visibility raster receives only eligible draws, and the main forward pass receives only compatibility draws;\n- off, validate, and debug modes allocate neither additional stream, so the shipping default remains unchanged;\n- exact routed memory is exposed in capability telemetry.\n\nCorrectness and regression evidence:\n- real-GPU routing oracle proves eligible, compatibility, and culled instance partitioning while preserving draw slots / first_instance;\n- forward-vs-visibility screenshot parity remains max 1 LSB, mean channel delta 0.00108643 (89 / 81,920 changed channels);\n- release suite: 378 unit tests passed (1 ignored), 59 golden tests passed (2 hardware-only ignored), plus all visibility runtime/parity tests;\n- strict release Clippy, wasm32 check, FFI/schema parity, quality governance, and example inventory all pass.\n\nThis removes duplicate cross-route vertex/fragment execution in the opt-in path. It does not yet justify default activation: Bistro total-frame GPU timings and per-MRT equivalence remain the next gates.

  9. proggeramlug commented on Aug 9, 2026

    @proggeramlug
    ContributorAuthor

    Compatibility ownership and MASK correctness checkpoint

    Pushed in 45a2cb1 on PR #147.

    The Bistro A/B exposed a real composition defect at the manhole/road boundary: visibility IDs were recorded against opaque prepass depth, then a nearer forward-only MASK surface updated final scene depth, but the later fullscreen visibility pass painted the stale road owner back over it.

    This checkpoint fixes the general ownership contract:

    • visibility PBR now samples final scene depth and rejects any reconstructed ID hidden by a later MASK, layered/custom, or unsupported compatibility draw;
    • the rejection happens before material/lighting evaluation, so hidden IDs avoid the expensive shading path;
    • surviving glTF MASK texels now output full coverage as required by MASK semantics, removing order-dependent alpha blending while preserving authored fractional-opacity and BLEND behavior;
    • the process-isolated GPU parity corpus now deliberately overlaps eligible geometry with both a nearer layered surface and a partially cut-out MASK surface, then compares screenshot, HDR, material, velocity, and albedo outputs.

    Deterministic GPU parity remains green:

    • final screenshot: 250 changed channels / 81,920, max 1 LSB, mean 0.00305176;
    • HDR: max absolute 0.000854492, mean 0.000004546065;
    • material/albedo: max 1 UNORM code;
    • velocity: max 0.000003815, mean 0.000000010617.

    Fresh fixed-camera Bistro off-vs-shade comparison improved from the preceding checkpoint:

    • luminance RMSE: 0.01204 -> 0.00699;
    • SSIM: 0.98184 -> 0.98907;
    • pixels above 0.025 tolerance: 2.10% -> 1.12%;
    • OKLab mean delta: 0.00231; edge mean delta: 0.00263.

    Scene depth, all three shadow cascades, SSGI, SSGI rejection/confidence, TAA motion, and TAA reprojected UV are byte-identical. The corrected manhole remains intact in the composed Bistro frame.

    Current Apple M1 Max sample still keeps activation opt-in: shade reduced mean GPU frame time from 32.27 ms to 28.17 ms, but the one-pair CPU/wall sample regressed and needs alternating repeated trials before it can satisfy the no-regression gate. No baseline was installed or changed.

    Validation: strict lint, wasm check, contracts/file ratchet, 39 quality-governance tests, visibility runtime/parity tests, 58 golden tests (2 hardware ignores; the backend live-object test passed on isolated rerun), and all 4 render-target tests.

  10. proggeramlug commented on Aug 9, 2026

    @proggeramlug
    ContributorAuthor

    PR #147 update — commit 9cebeef is pushed and supersedes the separate visibility raster/PBR pass layout.

    What changed:

    • Shade mode now writes alpha-aware packed visibility IDs during the existing depth prepass, so eligible geometry is traversed once for depth + ID rather than once for depth and again depth-equal for IDs.
    • Full-screen reconstructed PBR shading runs inside the existing main HDR MRT pass, before immediate/forward compatibility geometry. Normal depth/alpha composition therefore remains authoritative for layered/custom, skinned, deforming, translucent, and unsupported draws.
    • The compatibility route retains one bounded indirect stream; the redundant visibility-only indirect stream is gone. Off/default still requests no primitive-index feature and allocates/records no visibility work.
    • Validate/debug retain their independently timed diagnostic raster/reconstruction route.

    Correctness evidence on Apple M1 Max / Metal:

    • Strict process-isolated forward-vs-shade oracle remains bounded to 250 changed display channels out of 81,920, max 1 LSB, mean 0.00305176.
    • HDR MRT max absolute delta 0.000854492, mean 0.000004546065.
    • Material/albedo max 1 code; velocity max 0.000003815, mean 0.000000010617.
    • Governed Bistro off-vs-shade at 800x450: RMSE 0.00623, SSIM 0.99040, 0.94% above 0.025, OKLab 0.00221, edge delta 0.00249; all pass the scene thresholds.

    Performance evidence, 3 interleaved pixel-exact runs, 180 warmup + 240 measured frames each, medians:

    • Forward/off: wall 10.6831 ms, CPU mean 2.5563 ms, CPU p95 3.1023 ms, GPU mean 32.2235 ms, GPU p95 35.8045 ms.
    • Visibility/shade: wall 10.5793 ms, CPU mean 2.7369 ms, CPU p95 3.2174 ms, GPU mean 26.8703 ms, GPU p95 32.6437 ms.
    • Delta: wall -0.97%, GPU mean -16.61%, GPU p95 -8.83%; CPU mean +0.181 ms / +7.07% and p95 +0.115 ms / +3.71%, both inside the governed 0.35 ms / 15% and 1.0 ms / 25% envelopes. The earlier visibility implementation was materially slower on CPU and wall; inlining removed both extra render passes and made total frame time net-positive.

    Final local gates after the refactor:

    • 379 unit tests passed (1 ignored), 59 golden-render tests passed (2 hardware-only ignores), 4 render-target tests passed.
    • Real-GPU visibility runtime, shade runtime, and process-isolated MRT parity tests passed.
    • Strict lint, contracts/file-size ratchet, wasm check, FFI parity, and the 39-test quality governance suite passed.

    Remaining before default activation: repeat on a discrete Vulkan adapter, run the broader moving/cutout/skinned/refraction/transparency and low-overdraw corpora, and keep the compatibility fallback. The path remains explicit opt-in; no default renderer behavior changed.

  11. proggeramlug commented on Aug 9, 2026

    @proggeramlug
    ContributorAuthor

    Dependency update from #131 / PR #147: virtual clusters now carry the exact temporal and material state needed by visibility shading without changing the current visibility path.

    Qualified at code revision 1b4d828: dense instance addressing, current/previous model transforms, inverse-transpose normals, tint, and generation-safe material IDs. Unbound material maps fail before dispatch. The real Metal oracle covers 8 selected clusters and 24 decoded corners; the full quick lane and existing visibility parity/runtime gates pass.

    This does not satisfy a #27 activation gate by itself. The next integration must namespace virtual draw IDs in Rg32Uint, pull cooked vertices in the opt-in raster/shading path, reproduce HDR/material/velocity/albedo exactly, preserve compatibility routing, and remain disabled until bounded submission and total-GPU qualification pass.

  12. proggeramlug commented on Aug 9, 2026

    @proggeramlug
    ContributorAuthor

    The #131 virtual-geometry producer can now write collision-free IDs into #27’s shared Rg32Uint ABI (6d1183d, 688db93, qualified by 1fce846). Compatibility IDs use draw bit 31 = 0; virtual IDs use bit 31 = 1; the existing primitive/front-face word remains unchanged. A real Metal readback proves raw ID/depth rasterization.

    It is deliberately not registered with the visibility runtime yet, so #27’s current compatibility pixels and performance are unchanged. The next slice is a disjoint virtual fullscreen PBR/MRT consumer plus an explicit guard preventing the existing compatibility shader from interpreting virtual IDs. No #27 box changes until full four-MRT parity and composition are proven.

  13. proggeramlug commented on Aug 9, 2026

    @proggeramlug
    ContributorAuthor

    Virtual-geometry visibility now has an exact authoritative PBR reconstruction bridge at 9b3130f with evidence in docs/evidence/issue-131-virtual-visibility-pbr-v1.{md,json}.

    The virtual ID consumer reconstructs the same temporal/material inputs and invokes the same specialized shade_main_scene evaluator as the compatibility path, returning all four established MRTs. Namespace ownership is disjoint in both directions. A real production-layout Metal test creates the pipeline successfully at the 8-storage-buffer fragment contract, and the governed quick lane passes.

    This does not change #27 acceptance boxes yet: the path is deliberately unattached and ordinary visibility pixels are unchanged. Remaining #27-facing work is renderer pass registration, attachment load/store composition, and scene routing without overlap or holes.

  14. proggeramlug commented on Aug 26, 2026

    @proggeramlug
    ContributorAuthor

    A related #131 production-integration checkpoint now exercises the visibility-buffer composition path with registered virtual geometry:

    Virtual clusters now rasterize into the existing packed Rg32Uint target after ordinary/compatibility depth. The established fullscreen consumer discards virtual IDs; the virtual consumer discards ordinary IDs and writes the same authoritative four MRTs before forward compatibility composition. Target recreation invalidates/rebuilds the virtual bind group, and shading is gated on successful current-frame virtual raster so stale content cannot leak into composition.

    The production Metal runtime oracle proves that an enabled but empty virtual batch is byte-exact with the ordinary visibility/forward result. A registered cooked triangle then changes only its expected region while unrelated ordinary and compatibility pixels remain exact. The default/no-models visibility-preparation source remains identical to the parent, and paired uncapped performance windows were effectively flat within host noise.

    This does not change #27 activation gates or enable visibility shading by default. Distinct ordinary eligible/compatibility stream qualification, complete MRT/Bistro corpora, total GPU timing, and discrete/integrated cross-backend coverage remain required.

  15. proggeramlug commented on Aug 28, 2026

    @proggeramlug
    ContributorAuthor

    Current-head requalification confirms the direct four-MRT parity gate remains green at f7abef7 on Apple M1 Max / Metal. The process-isolated forward-vs-visibility test captures the production attachments before post-processing and exercises nonzero retained-object motion plus textured eligible geometry and layered/cutout compatibility composition. Metrics: final output 187 changed channels, max 1 LSB, mean 0.00228271; HDR max/mean absolute error 0.000854492 / 0.000004231930; material 36 changed components, max 1 LSB; velocity 55 changed components, max 0.000003815 with >256 moving components in the reference; albedo 76 changed components, max 1 LSB. All values pass the existing hard gates and every attachment is finite. This acceptance item was implemented by 0cd366d but remained unchecked, so the issue record is now corrected.

  16. proggeramlug commented on Aug 29, 2026

    @proggeramlug
    ContributorAuthor

    Current-head requalification also confirms the previously implemented stream-routing item at f7abef7. c7b2259 introduced a dedicated slot-preserving compatibility indirect stream while retaining the all-draw stream for the shared depth prepass; visibility raster consumes only DRAW_FLAG_VISIBILITY_ELIGIBLE commands and the forward MRT pass consumes only the complementary stream. cargo test --release --test visibility_buffer_shading_runtime -- --nocapture passes on Apple M1 Max / Metal, including the deliberate forward-shader fault that would overwrite admitted pixels if eligible geometry were still dispatched through compatibility shading. Runtime telemetry reports visibility_routed_indirect_streams=true. The checkbox is now corrected; activation remains blocked on the explicit uncapped total-pass A/B and broader corpus/hardware gates.

  17. proggeramlug commented on Aug 29, 2026

    @proggeramlug
    ContributorAuthor

    Visibility performance qualification is now reproducible and pushed in 0cce864 (workloads) and 6d083db (evidence). The tool now has process-isolated retained visibility-low-overdraw (576 draws) and visibility-layered-overdraw (4,608 draws/eight layers) workloads, records the workload in JSON, and emits the production capability/routing snapshot plus uncapped GPU timestamp totals. Shade mode measures inline ID raster in depth_prepass and inline visibility PBR in main_hdr_pass; no artificial extra pass was added for easier timing.

    Exact clean-revision ABBA on Apple M1 Max / Metal, 1600x900 native, 180 warmup + 240 measured frames per process:

    • low overdraw: forward 0.927790 ms mean / 0.955480 ms p95; shade 1.024298 / 1.049417 ms (+10.40% mean);
    • eight-layer overdraw: forward 2.535183 / 2.868125 ms; shade 2.484490 / 2.663354 ms (-2.00% mean);
    • layered depth_prepass improved 5.84%, but main_hdr_pass regressed 6.80% from full-screen reconstruction/PBR.

    The timing-capture acceptance item is now complete, but activation is explicitly rejected: low-overdraw exceeds the 5% regression guard and the layered case misses the required 5% material mean improvement. The path remains opt-in. This is consistent with Bloom already rejecting hidden fragments in its alpha-aware depth prepass.

    Evidence: Markdown, JSON.

    Exact-tree gates: 471 shared tests passed / 1 existing ignored; visibility routing and four-MRT parity real-GPU tests passed; 79 GPU goldens passed / 2 hardware diagnostics ignored; render-perf release/unit/strict-Clippy/format/diff checks passed. The only emitted shared warning remains the existing src/drs.rs unused mut.

  18. proggeramlug commented on Aug 29, 2026

    @proggeramlug
    ContributorAuthor

    MASK/cutout compatibility ownership is now explicitly decided, regression-gated, and pushed.

    Checkpoints:

    • f651840 — parity telemetry must retain six layered opaque draws plus the overlapping MASK draw as at least seven forward compatibility draws
    • 36d8b08 — permanent decision and real-GPU evidence

    V1 decision: glTF MASK/cutout remains on Bloom's alpha-aware forward compatibility path. The existing depth prepass already owns exact cutoff/coverage and rejects hidden PBR work; admitting MASK to visibility would need independent texture/sampler/UV-transform/coverage-mip/derivative/order qualification without a demonstrated total-GPU benefit.

    The process-isolated Metal oracle renders 32 moving visibility draws behind six moving layered draws and one moving 2x2 MASK draw. The mask has two opaque and two transparent texels, so surviving texels must replace visibility HDR/MRT output while discarded texels reveal the visibility-owned surface.

    Clean revision results:

    • final: 187 / 81,920 changed channels, max 1 LSB, mean 0.00228271
    • HDR max/mean: 0.000854492 / 0.000004231930
    • material and albedo: max 1 code
    • velocity max/mean: 0.000003815 / 0.000000005122 with >256 nonzero reference components
    • release target: 2 passed, 0 failed

    Evidence: Markdown and JSON.

    I checked only the MASK/cutout decision item. Broader Bistro/effects, performance, discrete-plus-integrated hardware, and default-activation gates remain open.

  19. proggeramlug commented on Aug 29, 2026

    @proggeramlug
    ContributorAuthor

    Full-Bistro effects checkpoint

    Pushed the large-scene visibility fix in 24d8ba4 and exact evidence in 6677c28.

    The clean-revision Metal run now loads all 2,909 bistrox.gltf placements and compares process-isolated forward/off vs visibility/shade across a 30-step camera route with TAA, SSAO, SSR, hardware-ray-query SSGI, bloom, compatibility rendering, and both cascaded and virtual shadows active.

    The first full-scene run found and fixed a real production limit: Bistro's 1,738,262 96-byte vertices occupy 159.143 MiB, exceeding Apple M1 Max Metal's 128 MiB maximum single storage binding. Visibility reconstruction now uses three offset-aligned vertex windows while preserving the global draw/index/base-vertex namespace and staying within the adapter's exact nine-storage-buffer fragment limit.

    All 14 captured comparisons pass. Worst CSM sample: mean 0.06482 code, RMS 0.46908, p99 2, max 15, SSIM 0.99943. Worst VSM sample: mean 0.06885, RMS 0.48673, p99 2, max 22, minimum SSIM 0.99876. Final-camera telemetry reports 2,404 visibility-owned and 164 forward-compatibility draws. The existing final/MRT parity target and all nine visibility unit/GPU tests remain green.

    Evidence: issue-27-bistro-effects-v1.md

    Scope is intentionally partial. The oracle disables asynchronous occlusion because workload-dependent previous-frame readback timing can change draw admission between A/B processes; occlusion remains a separate qualifier. Custom opaque/planar probes, the complete skinning/refraction/transparency stress corpus, a discrete adapter, and the already-failing performance activation gates also remain open. Visibility stays opt-in; I am not checking the broad corpus/default-activation items yet.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions