Repository navigation
[Renderer] Make the steady-state frame allocation-stable and upload-efficient #139
Description
Activity
First steady-state upload slice is implemented in draft PR #147 at
fc7521a.Lighting upload architecture:
- Renderer lighting setters now mutate one CPU
LightingUniformssnapshot and perform no immediate GPU upload. - After camera, frame, and shadow-cascade state is finalized, the renderer compares that snapshot byte-for-byte with the last submitted snapshot and uploads only changed aligned ranges.
- The unchanged shader ABI, bind group, and ~9 KiB GPU buffer are split into three logical dirty regions: fixed/directional fields, point-light records, and view/shadow/frame data.
- Planning uses fixed arrays and byte slices: no steady-state heap allocation, readback, new bind group, pass, buffer, or texture. There are at most three queue writes per frame; unchanged regions produce none.
renderer_paths.steady_state_uploads.lightingnow reports exact last-framewrite_count,byte_count, andfull_buffer_bytesin every native quality artifact.
Correctness/performance guards:
- Pure tests cover zero writes for an unchanged snapshot, three bounded/non-overlapping regions, and coalescing repeated mutations of one point-light record into one 32-byte upload.
- The real-GPU 40-point-light golden now hard-gates its sixth unchanged frame to
<= 3writes and<= 512bytes, instead of the former 8–9 complete ~8.7 KiB uploads. The same test still compares the rendered image to the checked-in golden. - Full local quick lane passed: 311/312 unit tests (one pre-existing ignored), negotiated headless renderer, 44/46 renderer goldens (two hardware-only ignored), render-target tests, shader/quality negative controls, strict correctness/performance clippy, and WASM compilation.
- Real Chrome WebGPU known-frame smoke passed with unchanged RGB means
(99, 177, 241).
This directly addresses the lighting-upload work item. I am leaving the issue acceptance box unchecked until the pushed commit completes hosted qualification; broader material/instance uploads, remaining bind groups, memory stability, and before/after timing evidence remain open.
- Renderer lighting setters now mutate one CPU
Implemented and pushed the next two #139 slices on PR #147:
ca5151d— allocation-free per-frame instrumentation for the 12 recurring core-frame bind-group creation sites. Native quality telemetry now exposesrenderer_paths.steady_state_resources.bind_group_creations.{total,sites}, and the quality artifact contract rejects missing, unknown, negative, or internally inconsistent counters.068383c— final-composite bind groups are cached across the complete identity set: 8 possible HDR source views × 2 exposure-history slots.resize()invalidates all 16 entries before recreating referenced views.
Regression/performance evidence at
068383c:- warmed real-GPU many-light golden: unchanged image gate passes and
final_compositeis hard-gated to0bind-group creations in the steady frame (previously1every frame) - full quick lane: 313/314 unit tests passed (1 existing ignored), negotiated headless device passed, 44/46 goldens passed (2 hardware-only ignored), render-target tests passed, strict correctness/performance clippy passed, wasm WebGPU check passed, quality governance passed, 20 canonical TypeScript examples inventoried
- transition stress passed: TAA toggle/history, exposure enable/ping-pong, render-scale changes, and resize; resize recovery SSIM
0.9996128966, max channel diff26, zero outlier pixels - FFI parity and file-size ratchet remain green
This is a strict steady-state reduction with unchanged shaders, uniform layout, texture allocation, bind-group descriptors, render-pass order, and draw calls. Hosted checks for the pushed head are now the remaining qualification layer before checking broader #139 acceptance items.
Additional pushed slice:
897e5b5caches scene-compose bind groups with one complete slot for each valid SSR input identity (cleared fallback, history 0, history 1). PT ownership and SSR toggles select an already complete binding; resize drops the cache before replacing any referenced view.Evidence on the exact commit:
- warmed real-GPU many-light golden remains green and now hard-gates both unconditional sites to
scene_compose: 0andfinal_composite: 0 - SSR history-lifetime test passed
- hardware ray-query PT ownership transition test passed on Apple M1 Max / Metal
- render-scale + resize recovery passed with SSIM
0.9996128966, max diff26, zero outlier pixels - full quick lane passed: 314/315 unit tests (1 existing ignored), negotiated renderer, 44/46 goldens (2 hardware-only ignored), render targets, strict clippy, wasm, quality contract, FFI parity, file-size ratchet, and example inventory
The steady-state reduction changes no shaders, uniform contents/layout, textures, passes, or draw calls.
- warmed real-GPU many-light golden remains green and now hard-gates both unconditional sites to
Completed and pushed the measured core bind-group elimination through
5983186.Steady-state cache commits since the prior update:
827af4fordinary TAA: two history identities652ef22reactive TAA: history identity + compiled plan ID + transient rebuild epoch8c111abnon-TAA upscale: persistent composed input2449e59DoF/motion-blur/SSS/CAS: exact upstream color-source identity5983186auto exposure (8 sources × 2 exposure slots) and custom post-pass A/B parity; quality contract now requires total named bind-group creation to be zero after warmup
Exact-head evidence:
- full quick lane green: 316/317 unit tests (1 existing ignored), negotiated renderer, 48/50 GPU goldens/integration tests (2 hardware-only ignored), render targets, strict clippy, wasm WebGPU, quality governance, FFI parity, file-size ratchet, 20-example inventory
- forced GPU coverage added for half-resolution upscale, TAA-history DoF, full optional DoF→motion blur→SSS→CAS+auto-exposure chain, and a two-pass custom WGSL stack
- reactive transparent-motion ghosting/flicker corpus remains within its established bounds and returns to
taa_reactive: 0after resize/rebuild - custom post-pass stack preserves scene geometry and returns to
custom_post_pass: 0after resize/rebuild - unchanged many-light golden hard-gates scene compose, final composite, SSR temporal, and ordinary TAA to zero steady-state creations
Checked the four acceptance items now directly proven: bind-group creation, dirty lighting uploads, toggle/resize/history safety, and corpus image tolerance. Pipeline/physical allocation instrumentation, 1,000-frame memory stability, and before/after percentile timings remain open and are next.
Pushed
6874170— steady-state graph/resource creation is now measured at the actual creation paths.New per-frame native quality telemetry under
renderer_paths.steady_state_resources:graph_compilescommand_encoder_creations.{total,sites.frame_submission}transient_physical_creations.{textures,buffers}
The compiled transient pool returns exact texture/buffer creation deltas; cache hits return zero. Graph compile count is a per-frame delta, not the pre-existing lifetime total. The normal frame-submission encoder is named and hard-gated to exactly one because command encoders are intentionally per-submission objects. Diagnostic screenshot/quality-capture topology and readback allocations are excluded from the steady-state snapshot, so qualification measures the last ordinary measured frame rather than its post-measurement capture frame.
Contract gates now fail if a warmed unchanged frame recompiles its topology, creates any graph-owned physical texture/buffer, or creates other than one named submission encoder. Negative controls cover every new failure. Real-GPU gates prove zero graph compiles/physical creations on the unchanged many-light frame, on warmed reactive transparency, and after reactive resize recovery.
Exact-head validation: full quick lane passed (316/317 unit tests, 1 existing ignored; negotiated headless renderer; 48/50 GPU goldens/integration tests, 2 hardware-only ignored; 3 render-target tests; strict clippy; wasm WebGPU check; quality governance; FFI parity; file-size ratchet; 20 canonical examples). No shader, render pass, descriptor, resource lifetime, cache key, draw, or upload behavior changed.
The first acceptance checkbox remains open because its pipeline-creation clause is not yet instrumented; that is the next slice.
Sponza steady-state creation criterion is now proven on the production Metal/headless path at commit
9823788(the instrumentation itself landed in6874170andae2e1c2).Qualification run:
python3 tools/quality/run.py run full \ --case sponza-interior \ --report-only \ --out tools/quality/out/local-sponza-headless-modeObserved after 180 warm-up frames across the 300-frame measurement window:
{ "present_mode": "auto-no-vsync", "uncapped": true, "steady_state_resources": { "graph_compiles": 0, "pipeline_creations": { "first_use": 0 }, "transient_physical_creations": { "textures": 0, "buffers": 0 }, "bind_group_creations": { "total": 0 }, "command_encoder_creations": { "total": 1, "sites": { "frame_submission": 1 } } }, "steady_state_uploads": { "lighting": { "write_count": 1, "byte_count": 24, "full_buffer_bytes": 8864 } } }The quality contract accepted every telemetry/resource gate. The case's only reported failure is:
approved baseline missing: tools/quality/baselines/portable/sponza-interior.pngThat missing human-review baseline remains work for the separate visual-corpus criterion; it does not invalidate the now-checked creation-count criterion. CPU P95 was 2.669 ms and GPU P95 was 48.724 ms at this case's 800×450 resolution, but those values are not being used to satisfy the separate 1080p/4K before/after acceptance item.
During this proof, the run exposed a qualification-only reporting bug: headless rendering correctly had no swapchain and no FPS cap, but rejected the
AutoNoVsyncmetadata request and serialized the constructor's placeholder FIFO value. Commit9823788now retains the requested mode for headless telemetry without changing any surface-backed presentation or rendered output. A production headless-device regression test pins uncapped mode reporting and invalid-mode rejection.The static 1,000-frame renderer-owned memory criterion is proven at commit
396659c.The new real-GPU test
static_ultra_scene_has_stable_renderer_owned_memory_for_1000_frames:- Creates a production headless renderer.
- Enables Ultra preset 4.
- Renders a static scene with geometry and 40 colored point lights.
- Warms 16 frames to settle temporal histories, rotating caches, queue staging, and the bounded three-frame headless submission window.
- Waits for the GPU and snapshots live backend GPU objects, backend-reported allocation bytes where available, renderer growable frame-container capacity, compiled graph-plan count, and transient physical-slot count.
- Renders 1,000 unchanged frames, waits again, and requires exact equality for every snapshot field.
Metal result:
buffers=97 textures=69 texture_views=131 bind_groups=1090 bind_group_layouts=47 render_pipelines=32 compute_pipelines=18 pipeline_layouts=43 samplers=19 shader_modules=44 query_sets=1 fences=1 tracked_frame_cpu_capacity_bytes=1,918,560 cached_graph_plans=1 physical_transient_slots=0Every value above was identical before and after the 1,000-frame window. The backend also reported no change in its buffer/texture/acceleration-structure byte counters and allocation count; Metal currently exposes those byte counters as zero, so the test separately hard-gates every live object category and the renderer's tracked CPU capacity instead of treating zero byte telemetry as proof.
The final frame additionally required:
graph_compiles=0 pipeline_creations.first_use=0 transient_physical_creations.textures=0 transient_physical_creations.buffers=0 bind_group_creations.total=0wgpu's internal counter atomics are activated through a dev-dependency only. Production engine builds retain the default counter-free wgpu configuration, so this qualification adds no shipped per-object or per-frame overhead.
Validation:
./scripts/ci-check.sh --quick PASS cargo test --manifest-path native/shared/Cargo.toml --release \ --test golden_render \ static_ultra_scene_has_stable_renderer_owned_memory_for_1000_frames \ -- --nocapture PASS (2.54 s on Apple M1 Max / Metal)The final performance acceptance criterion is now proven at commit
9d949ef.Controlled comparison:
- before engine revision:
0bc6fabc7bb868f2ec8cb25469853ae2a3b74e4e(the parent of the [Renderer] Make the steady-state frame allocation-stable and upload-efficient #139 optimization series) - after engine revision:
93dfd942373024d4899071c9e3b8e561647906ef - Apple M1 Max / Metal, release build, production headless renderer, Ultra preset 4, vsync/FPS cap disabled
- fixed plane + cube + 40 point lights
- 300 warm-up frames, 900 measured frames, three runs per revision/resolution
- reported P50/P95/P99 are medians of the three per-run percentiles
- the primary clock covers the complete renderer submission frame:
begin_frame, fixed draw/light submission, andend_frame
Resolution Metric Before After Change 1920×1080 P50 2.654 ms 2.465 ms −7.1% 1920×1080 P95 4.069 ms 3.352 ms −17.6% 1920×1080 P99 12.179 ms 3.507 ms −71.2% 3840×2160 P50 6.301 ms 6.100 ms −3.2% 3840×2160 P95 6.558 ms 6.273 ms −4.3% 3840×2160 P99 7.688 ms 6.387 ms −16.9% The complete render-submit metric improved at every percentile and resolution.
A separate trace-only qualification summed every
Queue::write_buffer/texture payload between 32 measured submits. Trace-mode timings were explicitly invalidated and were not used in the timing table:Resolution Before After Reduction 1920×1080 422,120 B/frame 23,264 B/frame 94.5% 3840×2160 422,120 B/frame 23,264 B/frame 94.5% That is 398,856 fewer bytes per unchanged frame (18.14× less traffic), with zero steady texture-upload bytes.
Reproducible evidence and raw aggregates:
Validation at the evidence commit:
cargo fmt --manifest-path tools/render-perf/Cargo.toml -- --checkcargo check --manifest-path tools/render-perf/Cargo.toml./scripts/ci-check.sh --quick
The quick lane passed: 316 unit tests, 49 golden/integration tests, render-target tests, strict clippy, wasm, FFI parity, quality governance/negative controls, file-size ratchet, and example inventory.
- before engine revision:
Closing as complete in draft PR #147. All seven acceptance criteria are checked and backed by exact-commit evidence: zero warmed graph/pipeline/transient creation, zero named steady-state bind-group creation, dirty lighting uploads, stable renderer-owned state across 1,000 frames, toggle/resize/history stress coverage, unchanged quality gates, and controlled 1080p/4K P50/P95/P99 plus upload-byte measurements. The final evidence is recorded at
9d949efindocs/evidence/issue-139-steady-state-performance.{md,json}.
Parent: #126
Corresponds to EN-056 in
docs/tickets.mdRelated: #30
Problem
The renderer still performs avoidable steady-state work: repeated full lighting-buffer uploads, per-frame bloom/pass bind-group allocation, frame-graph construction, and other allocations/uploads documented in EN-056. These costs become significant at 4K, with many materials/lights, and as new quality passes arrive. Previous eager bind-group caching produced a black frame because cache keys missed frame-order/resource-validity state (#30).
Outcome
An allocation-stable steady-state frame with dirty/range-based uploads, persistent pass resources, compiled frame plans, and correctness-backed cache keys.
Required instrumentation
Before optimizing, add per-frame counters/bytes/timers for:
Queue::write_buffer/staging writes by resource and byte count;Expose counters through existing profiler output and reset them per frame.
Work breakdown
Do not cache views of rotating history textures under a key that omits history index/generation/write-validity.
Acceptance criteria
Likely files
native/shared/src/renderer/mod.rs,graph.rs,transient.rs,postfx_chain.rsnative/shared/src/profiler.rs,staging.rsNon-goals
Dependencies