Skip to content

Price a mesh cage by the control points that were dragged - #655

Merged
leonardoaraujosantos merged 1 commit into
mainfrom
perf/mesh-lattice-dragged-points
Sep 29, 2026
Merged

leonardoaraujosantos merged 1 commit into
mainfrom
perf/mesh-lattice-dragged-points

Conversation

@leonardoaraujosantos

Copy link
Copy Markdown
Contributor

Why

mesh::Lattice::displacement summed every control point at every vertex: nx * ny * nz multiply-adds whatever had been dragged. At the 32³ ceiling that is 32,768 terms per vertex. ClaySpace previews a mesh cage drag by laying the cage over the mesh on every pointer move, and one frame with a single corner in hand measured:

cage per frame, before (ClaySpace, 62k vertices)
3³ ~10 ms
8³ ~35 ms
32³ ~1.7 s (4.7–5.2 s on the 80k-triangle fixture in CyberdyneCorp/ClaySpaceDesktop#176)

The offset field is linear in the offsets, so a control point at rest adds exactly nothing. That whole cost was work with no result.

What changes

  • Lattice keeps the set of control points with a non-zero offset, maintained by set_offset (a point set back to zero leaves it). displacement sums over that set only; is_identity() is that set being empty (O(1)); dragged_count() reports its size.
  • Each axis's Bernstein basis is built in O(n) from powers of t and 1 − t with precomputed binomials, in double, rather than de Casteljau's O(n²) recurrence in float — negligible beside the old n³ sum, not beside the new one.
  • BM_MeshLatticeDrag times one preview frame (apply with a record, revert) at 3³, 8³ and 32³ with one point dragged over ~100k triangles.
  • OpenSpec change price-a-mesh-cage-by-its-dragged-points adds the meshing requirement; docs/07 states the cost.

No C ABI, Python or file-format change. The SDF lattice deformer (clattice_point) is a separate evaluator and is untouched.

Evidence

  • New unit tests in test_mesh_lattice.cpp: the sparse sum matches the full trivariate Bernstein sum (reference written the slow way, de Casteljau in double) at 2, 3, 8 and 32 divisions, inside, on and outside the box; the corner of a 32³ cage still interpolates exactly; the dragged count follows writes including a point returned to rest. Existing lattice tests unchanged and passing.
  • ctest --preset cpu-only: all unit suites pass locally (the one local failure, clay_device_bench_selftest, is the host's Python 3.9 rejecting float | None in tools/check_device_bench.py; unrelated).
  • BM_MeshLatticeDrag (CPU time, loaded machine): 3³ 2.9 ms, 8³ 3.7 ms, 32³ 8.8 ms.
  • End to end in ClaySpace with this engine: a single-point drag frame at 32³ went from ~1.7 s to ~11 ms on a 62k-vertex mesh (3³ and 8³ ~9–10 ms, the remainder being normals and the host's own work).

Needed by CyberdyneCorp/ClaySpaceDesktop#176; ClaySpace consumes it once it ships in a release.

Lattice::displacement summed every control point at every vertex, so one
evaluation cost nx*ny*nz multiply-adds whatever had been dragged: 32,768
terms per vertex at the 32^3 ceiling, about 1.7 s per application over a
62k-vertex mesh with a single corner in hand. A host previewing a cage drag
pays that on every pointer move.

The offset field is linear in the offsets, so a point at rest adds nothing.
The cage now keeps the set of dragged points (maintained by set_offset) and
sums over it alone; is_identity is that set being empty. Each axis basis is
built in O(n) from powers of t and 1-t with precomputed binomials, in double.

Results agree with the full Bernstein sum to float resolution; no ABI,
Python or format change. BM_MeshLatticeDrag times one preview frame at 3^3,
8^3 and 32^3.
@leonardoaraujosantos
leonardoaraujosantos merged commit 7ac9f6d into main Sep 29, 2026
16 checks passed
leonardoaraujosantos added a commit that referenced this pull request Oct 4, 2026
Covers the twenty PRs since v0.120.1 (#655, #667-#669, #673-#688), ABI
minors 0.121.0 through 0.126.0, and the scene format minor 19 -> 20.
The device gate block is a placeholder until the iPad run is recorded.

Carries the four manual hardware gates forward as waivers at 4401b35:
the only kernel-relevant change since 59e42cc is #680's per-brick cull
band in include/clay/eval/bake_volume.h, which changes what a brick
compiles, not kernel arithmetic, and the parity corpus is unchanged.
leonardoaraujosantos added a commit that referenced this pull request Oct 5, 2026
Covers the twenty PRs since v0.120.1 (#655, #667-#669, #673-#688), ABI
minors 0.121.0 through 0.126.0, and the scene format minor 19 -> 20.
The device gate block is a placeholder until the iPad run is recorded.

Carries the four manual hardware gates forward as waivers at 4401b35:
the only kernel-relevant change since 59e42cc is #680's per-brick cull
band in include/clay/eval/bake_volume.h, which changes what a brick
compiles, not kernel arithmetic, and the parity corpus is unchanged.
leonardoaraujosantos added a commit that referenced this pull request Oct 6, 2026
Covers the twenty PRs since v0.120.1 (#655, #667-#669, #673-#688), ABI
minors 0.121.0 through 0.126.0, and the scene format minor 19 -> 20.
The device gate block is a placeholder until the iPad run is recorded.

Carries the four manual hardware gates forward as waivers at 4401b35:
the only kernel-relevant change since 59e42cc is #680's per-brick cull
band in include/clay/eval/bake_volume.h, which changes what a brick
compiles, not kernel arithmetic, and the parity corpus is unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant