Skip to content

[Renderer] Make the steady-state frame allocation-stable and upload-efficient #139

Description

@proggeramlug

Parent: #126
Corresponds to EN-056 in docs/tickets.md
Related: #30

Problem

The renderer still performs avoidable steady-state work: repeated full lighting-buffer uploads, per-frame bloom/pass bind-group allocation, frame-graph construction, and other allocations/uploads documented in EN-056. These costs become significant at 4K, with many materials/lights, and as new quality passes arrive. Previous eager bind-group caching produced a black frame because cache keys missed frame-order/resource-validity state (#30).

Outcome

An allocation-stable steady-state frame with dirty/range-based uploads, persistent pass resources, compiled frame plans, and correctness-backed cache keys.

Required instrumentation

Before optimizing, add per-frame counters/bytes/timers for:

  • CPU allocations in renderer prepare/submit where practical;
  • Queue::write_buffer/staging writes by resource and byte count;
  • bind group, pipeline, encoder, and transient physical-resource creation;
  • graph compile/build count;
  • lighting/material/instance records dirtied vs uploaded;
  • cache hit/miss/rebuild reason.

Expose counters through existing profiler output and reset them per frame.

Work breakdown

  1. Lighting uploads: separate per-frame/view data from light arrays; upload once after light collection or update dirty ranges, not the whole ~8–9 KB block repeatedly.
  2. Material/instance uploads: use stable GPU buffers, dirty ranges, and growth strategy; avoid rebuilding unchanged records.
  3. Frame plan: consume the compiled render-graph cache; configuration/uniform changes must not rebuild topology.
  4. Bind groups: cache only when all bound resource identities/generations and validity state are represented. Land one pass at a time with screenshot/intermediate-buffer diff.
  5. Transient allocations: reuse graph-planned physical resources across frames/resizes.
  6. Shader/pipeline creation: prove all steady-state variants are prewarmed or cached; report first-use compilation separately.

Do not cache views of rotating history textures under a key that omits history index/generation/write-validity.

Acceptance criteria

  • After warm-up in an unchanged Ultra Sponza frame, graph compiles, pipeline creations, and new physical texture/buffer allocations are zero.
  • Bind-group creation is zero or an explicitly justified bounded count; each remaining site is named.
  • Lighting upload count/bytes scale with actual dirty data and are no longer repeated 8–9 times per unchanged frame.
  • A static-scene 1,000-frame run has stable renderer-owned memory after caches warm.
  • Toggle/resize/history-index stress tests produce no black frames, stale resource reads, or validation errors.
  • Quality corpus remains within image tolerance after each cache migration.
  • Before/after P50/P95/P99 CPU render-submit time and total upload bytes are attached at 1080p and 4K, vsync off.

Likely files

  • native/shared/src/renderer/mod.rs, graph.rs, transient.rs, postfx_chain.rs
  • material/light/scene pass modules
  • native/shared/src/profiler.rs, staging.rs

Non-goals

Dependencies

Activity

  1. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    First steady-state upload slice is implemented in draft PR #147 at fc7521a.

    Lighting upload architecture:

    • Renderer lighting setters now mutate one CPU LightingUniforms snapshot and perform no immediate GPU upload.
    • After camera, frame, and shadow-cascade state is finalized, the renderer compares that snapshot byte-for-byte with the last submitted snapshot and uploads only changed aligned ranges.
    • The unchanged shader ABI, bind group, and ~9 KiB GPU buffer are split into three logical dirty regions: fixed/directional fields, point-light records, and view/shadow/frame data.
    • Planning uses fixed arrays and byte slices: no steady-state heap allocation, readback, new bind group, pass, buffer, or texture. There are at most three queue writes per frame; unchanged regions produce none.
    • renderer_paths.steady_state_uploads.lighting now reports exact last-frame write_count, byte_count, and full_buffer_bytes in every native quality artifact.

    Correctness/performance guards:

    • Pure tests cover zero writes for an unchanged snapshot, three bounded/non-overlapping regions, and coalescing repeated mutations of one point-light record into one 32-byte upload.
    • The real-GPU 40-point-light golden now hard-gates its sixth unchanged frame to <= 3 writes and <= 512 bytes, instead of the former 8–9 complete ~8.7 KiB uploads. The same test still compares the rendered image to the checked-in golden.
    • Full local quick lane passed: 311/312 unit tests (one pre-existing ignored), negotiated headless renderer, 44/46 renderer goldens (two hardware-only ignored), render-target tests, shader/quality negative controls, strict correctness/performance clippy, and WASM compilation.
    • Real Chrome WebGPU known-frame smoke passed with unchanged RGB means (99, 177, 241).

    This directly addresses the lighting-upload work item. I am leaving the issue acceptance box unchecked until the pushed commit completes hosted qualification; broader material/instance uploads, remaining bind groups, memory stability, and before/after timing evidence remain open.

  2. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Implemented and pushed the next two #139 slices on PR #147:

    • ca5151d — allocation-free per-frame instrumentation for the 12 recurring core-frame bind-group creation sites. Native quality telemetry now exposes renderer_paths.steady_state_resources.bind_group_creations.{total,sites}, and the quality artifact contract rejects missing, unknown, negative, or internally inconsistent counters.
    • 068383c — final-composite bind groups are cached across the complete identity set: 8 possible HDR source views × 2 exposure-history slots. resize() invalidates all 16 entries before recreating referenced views.

    Regression/performance evidence at 068383c:

    • warmed real-GPU many-light golden: unchanged image gate passes and final_composite is hard-gated to 0 bind-group creations in the steady frame (previously 1 every frame)
    • full quick lane: 313/314 unit tests passed (1 existing ignored), negotiated headless device passed, 44/46 goldens passed (2 hardware-only ignored), render-target tests passed, strict correctness/performance clippy passed, wasm WebGPU check passed, quality governance passed, 20 canonical TypeScript examples inventoried
    • transition stress passed: TAA toggle/history, exposure enable/ping-pong, render-scale changes, and resize; resize recovery SSIM 0.9996128966, max channel diff 26, zero outlier pixels
    • FFI parity and file-size ratchet remain green

    This is a strict steady-state reduction with unchanged shaders, uniform layout, texture allocation, bind-group descriptors, render-pass order, and draw calls. Hosted checks for the pushed head are now the remaining qualification layer before checking broader #139 acceptance items.

  3. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Additional pushed slice: 897e5b5 caches scene-compose bind groups with one complete slot for each valid SSR input identity (cleared fallback, history 0, history 1). PT ownership and SSR toggles select an already complete binding; resize drops the cache before replacing any referenced view.

    Evidence on the exact commit:

    • warmed real-GPU many-light golden remains green and now hard-gates both unconditional sites to scene_compose: 0 and final_composite: 0
    • SSR history-lifetime test passed
    • hardware ray-query PT ownership transition test passed on Apple M1 Max / Metal
    • render-scale + resize recovery passed with SSIM 0.9996128966, max diff 26, zero outlier pixels
    • full quick lane passed: 314/315 unit tests (1 existing ignored), negotiated renderer, 44/46 goldens (2 hardware-only ignored), render targets, strict clippy, wasm, quality contract, FFI parity, file-size ratchet, and example inventory

    The steady-state reduction changes no shaders, uniform contents/layout, textures, passes, or draw calls.

  4. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Completed and pushed the measured core bind-group elimination through 5983186.

    Steady-state cache commits since the prior update:

    • 827af4f ordinary TAA: two history identities
    • 652ef22 reactive TAA: history identity + compiled plan ID + transient rebuild epoch
    • 8c111ab non-TAA upscale: persistent composed input
    • 2449e59 DoF/motion-blur/SSS/CAS: exact upstream color-source identity
    • 5983186 auto exposure (8 sources × 2 exposure slots) and custom post-pass A/B parity; quality contract now requires total named bind-group creation to be zero after warmup

    Exact-head evidence:

    • full quick lane green: 316/317 unit tests (1 existing ignored), negotiated renderer, 48/50 GPU goldens/integration tests (2 hardware-only ignored), render targets, strict clippy, wasm WebGPU, quality governance, FFI parity, file-size ratchet, 20-example inventory
    • forced GPU coverage added for half-resolution upscale, TAA-history DoF, full optional DoF→motion blur→SSS→CAS+auto-exposure chain, and a two-pass custom WGSL stack
    • reactive transparent-motion ghosting/flicker corpus remains within its established bounds and returns to taa_reactive: 0 after resize/rebuild
    • custom post-pass stack preserves scene geometry and returns to custom_post_pass: 0 after resize/rebuild
    • unchanged many-light golden hard-gates scene compose, final composite, SSR temporal, and ordinary TAA to zero steady-state creations

    Checked the four acceptance items now directly proven: bind-group creation, dirty lighting uploads, toggle/resize/history safety, and corpus image tolerance. Pipeline/physical allocation instrumentation, 1,000-frame memory stability, and before/after percentile timings remain open and are next.

  5. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Pushed 6874170 — steady-state graph/resource creation is now measured at the actual creation paths.

    New per-frame native quality telemetry under renderer_paths.steady_state_resources:

    • graph_compiles
    • command_encoder_creations.{total,sites.frame_submission}
    • transient_physical_creations.{textures,buffers}

    The compiled transient pool returns exact texture/buffer creation deltas; cache hits return zero. Graph compile count is a per-frame delta, not the pre-existing lifetime total. The normal frame-submission encoder is named and hard-gated to exactly one because command encoders are intentionally per-submission objects. Diagnostic screenshot/quality-capture topology and readback allocations are excluded from the steady-state snapshot, so qualification measures the last ordinary measured frame rather than its post-measurement capture frame.

    Contract gates now fail if a warmed unchanged frame recompiles its topology, creates any graph-owned physical texture/buffer, or creates other than one named submission encoder. Negative controls cover every new failure. Real-GPU gates prove zero graph compiles/physical creations on the unchanged many-light frame, on warmed reactive transparency, and after reactive resize recovery.

    Exact-head validation: full quick lane passed (316/317 unit tests, 1 existing ignored; negotiated headless renderer; 48/50 GPU goldens/integration tests, 2 hardware-only ignored; 3 render-target tests; strict clippy; wasm WebGPU check; quality governance; FFI parity; file-size ratchet; 20 canonical examples). No shader, render pass, descriptor, resource lifetime, cache key, draw, or upload behavior changed.

    The first acceptance checkbox remains open because its pipeline-creation clause is not yet instrumented; that is the next slice.

  6. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    Sponza steady-state creation criterion is now proven on the production Metal/headless path at commit 9823788 (the instrumentation itself landed in 6874170 and ae2e1c2).

    Qualification run:

    python3 tools/quality/run.py run full \
      --case sponza-interior \
      --report-only \
      --out tools/quality/out/local-sponza-headless-mode
    

    Observed after 180 warm-up frames across the 300-frame measurement window:

    {
      "present_mode": "auto-no-vsync",
      "uncapped": true,
      "steady_state_resources": {
        "graph_compiles": 0,
        "pipeline_creations": { "first_use": 0 },
        "transient_physical_creations": {
          "textures": 0,
          "buffers": 0
        },
        "bind_group_creations": { "total": 0 },
        "command_encoder_creations": {
          "total": 1,
          "sites": { "frame_submission": 1 }
        }
      },
      "steady_state_uploads": {
        "lighting": {
          "write_count": 1,
          "byte_count": 24,
          "full_buffer_bytes": 8864
        }
      }
    }

    The quality contract accepted every telemetry/resource gate. The case's only reported failure is:

    approved baseline missing: tools/quality/baselines/portable/sponza-interior.png
    

    That missing human-review baseline remains work for the separate visual-corpus criterion; it does not invalidate the now-checked creation-count criterion. CPU P95 was 2.669 ms and GPU P95 was 48.724 ms at this case's 800×450 resolution, but those values are not being used to satisfy the separate 1080p/4K before/after acceptance item.

    During this proof, the run exposed a qualification-only reporting bug: headless rendering correctly had no swapchain and no FPS cap, but rejected the AutoNoVsync metadata request and serialized the constructor's placeholder FIFO value. Commit 9823788 now retains the requested mode for headless telemetry without changing any surface-backed presentation or rendered output. A production headless-device regression test pins uncapped mode reporting and invalid-mode rejection.

  7. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    The static 1,000-frame renderer-owned memory criterion is proven at commit 396659c.

    The new real-GPU test static_ultra_scene_has_stable_renderer_owned_memory_for_1000_frames:

    1. Creates a production headless renderer.
    2. Enables Ultra preset 4.
    3. Renders a static scene with geometry and 40 colored point lights.
    4. Warms 16 frames to settle temporal histories, rotating caches, queue staging, and the bounded three-frame headless submission window.
    5. Waits for the GPU and snapshots live backend GPU objects, backend-reported allocation bytes where available, renderer growable frame-container capacity, compiled graph-plan count, and transient physical-slot count.
    6. Renders 1,000 unchanged frames, waits again, and requires exact equality for every snapshot field.

    Metal result:

    buffers=97
    textures=69
    texture_views=131
    bind_groups=1090
    bind_group_layouts=47
    render_pipelines=32
    compute_pipelines=18
    pipeline_layouts=43
    samplers=19
    shader_modules=44
    query_sets=1
    fences=1
    tracked_frame_cpu_capacity_bytes=1,918,560
    cached_graph_plans=1
    physical_transient_slots=0
    

    Every value above was identical before and after the 1,000-frame window. The backend also reported no change in its buffer/texture/acceleration-structure byte counters and allocation count; Metal currently exposes those byte counters as zero, so the test separately hard-gates every live object category and the renderer's tracked CPU capacity instead of treating zero byte telemetry as proof.

    The final frame additionally required:

    graph_compiles=0
    pipeline_creations.first_use=0
    transient_physical_creations.textures=0
    transient_physical_creations.buffers=0
    bind_group_creations.total=0
    

    wgpu's internal counter atomics are activated through a dev-dependency only. Production engine builds retain the default counter-free wgpu configuration, so this qualification adds no shipped per-object or per-frame overhead.

    Validation:

    ./scripts/ci-check.sh --quick
    PASS
    
    cargo test --manifest-path native/shared/Cargo.toml --release \
      --test golden_render \
      static_ultra_scene_has_stable_renderer_owned_memory_for_1000_frames \
      -- --nocapture
    PASS (2.54 s on Apple M1 Max / Metal)
    
  8. proggeramlug commented on Jul 28, 2026

    @proggeramlug
    ContributorAuthor

    The final performance acceptance criterion is now proven at commit 9d949ef.

    Controlled comparison:

    • before engine revision: 0bc6fabc7bb868f2ec8cb25469853ae2a3b74e4e (the parent of the [Renderer] Make the steady-state frame allocation-stable and upload-efficient #139 optimization series)
    • after engine revision: 93dfd942373024d4899071c9e3b8e561647906ef
    • Apple M1 Max / Metal, release build, production headless renderer, Ultra preset 4, vsync/FPS cap disabled
    • fixed plane + cube + 40 point lights
    • 300 warm-up frames, 900 measured frames, three runs per revision/resolution
    • reported P50/P95/P99 are medians of the three per-run percentiles
    • the primary clock covers the complete renderer submission frame: begin_frame, fixed draw/light submission, and end_frame
    Resolution Metric Before After Change
    1920×1080 P50 2.654 ms 2.465 ms −7.1%
    1920×1080 P95 4.069 ms 3.352 ms −17.6%
    1920×1080 P99 12.179 ms 3.507 ms −71.2%
    3840×2160 P50 6.301 ms 6.100 ms −3.2%
    3840×2160 P95 6.558 ms 6.273 ms −4.3%
    3840×2160 P99 7.688 ms 6.387 ms −16.9%

    The complete render-submit metric improved at every percentile and resolution.

    A separate trace-only qualification summed every Queue::write_buffer/texture payload between 32 measured submits. Trace-mode timings were explicitly invalidated and were not used in the timing table:

    Resolution Before After Reduction
    1920×1080 422,120 B/frame 23,264 B/frame 94.5%
    3840×2160 422,120 B/frame 23,264 B/frame 94.5%

    That is 398,856 fewer bytes per unchanged frame (18.14× less traffic), with zero steady texture-upload bytes.

    Reproducible evidence and raw aggregates:

    Validation at the evidence commit:

    • cargo fmt --manifest-path tools/render-perf/Cargo.toml -- --check
    • cargo check --manifest-path tools/render-perf/Cargo.toml
    • ./scripts/ci-check.sh --quick

    The quick lane passed: 316 unit tests, 49 golden/integration tests, render-target tests, strict clippy, wasm, FFI parity, quality governance/negative controls, file-size ratchet, and example inventory.

  9. proggeramlug commented on Aug 1, 2026

    @proggeramlug
    ContributorAuthor

    Closing as complete in draft PR #147. All seven acceptance criteria are checked and backed by exact-commit evidence: zero warmed graph/pipeline/transient creation, zero named steady-state bind-group creation, dirty lighting uploads, stable renderer-owned state across 1,000 frames, toggle/resize/history stress coverage, unchanged quality gates, and controlled 1080p/4K P50/P95/P99 plus upload-byte measurements. The final evidence is recorded at 9d949ef in docs/evidence/issue-139-steady-state-performance.{md,json}.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions