perf: add deterministic synthetic benchmark corpus - #379
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Fresh packed-artifact acquisition completed for unchanged exact head Acquisition identity:
Observed p50 / p95 milliseconds:
The suite completed successfully from a clean detached checkout and exact packed artifact. This is a local reference snapshot, not an accepted support budget, cross-machine claim, protected-main activation, or release proof. Product hot-path changes remain with their existing source-owner PRs rather than being duplicated here. |
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
|
Autoresearch experiment 5 kept at |
Preserve the no-op baseline and add an explicit before/after corpus, packed suite lane, and scenario correctness gates. Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Exact-head local evidence: changed-transition probe and prerequisite integrationSource:
Separate synthetic scenario baselinesOne serial local acquisition on Node 24.19.0; 25 measured samples per operation/profile, one existing warm-up, all four full packed-suite profiles. Units: milliseconds. These are distinct workloads, not an A/B optimization comparison.
The new scenario preserves the old metric's meaning and starts its own baseline. It does not establish the 20 ms target, realistic buyer-document latency, a rich-document-tree support envelope, physical-device IME performance, leak freedom, or hosted acceptance. Hosted current-head checks/reviews and protected prerequisite integration remain separate gates. Raw local receipts and the hardware fingerprint remain in the local evidence bundle; do not transfer predecessor results to a later head. |
…euse Stack performance PR #379 after canonical owner #176 at 94b5ca8. The sole conflicted source file now exactly matches the owner, retaining original encoder reuse, output bounds and hostile-option validation. All benchmark and chunking deltas remain intact. Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
First-invocation accounting and generation isolation — 2026-09-06Candidate The earlier latency producers executed one preliminary operation that did not New JavaScript latency receipts use Exact-candidate local verification:
The earlier 108 files / 2,700 timed samples remain unchanged. Their additional Keep Draft and integrate the canonical parents first. These local results do |
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
… limits Signed-off-by: Seongho Bae <me@seonghobae.me>
Latest output-consistency repairSource-browser acquisition (not packed): first current Current Office verification: same Current packed-browser acquisition: at Current package and new authored-input baseline: full build, independent Current candidate is RED The first full acquisition is retained: 1,159 passed, 15 failed in 228.52 seconds. Harness README, TRD and the existing performance gap baseline distinguish Predecessor input identity repair evidenceExact candidate: The configured full coverage run is terminal: 1,167 passed, 3 failed, An unchanged one-worker run of the affected two files then produced 6 passed, An additional unchanged full coverage run, serialized without our other heavy The current exact-head build and full All local logs remain under Actual authored-document baseline on the same headThe unchanged canonical README, PRD, TRD and CONTRACTS documents were measured Archive SHA-256: Full extracted-package browser validation subsequently passed all 70 tests |
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Scope and ownership
Refs #375. This branch is the established single-writer implementation lane for Inkspan's deterministic large-document performance evidence. Protected
mainremains the only shipped implementation authority. No SHA, merge-base, check result, review count, or workflow state recorded in this PR body is lifecycle authority; refetch those values at decision time.The lane remains standalone, deterministic, provider-neutral, and network/model independent. It contains deterministic synthetic Markdown/Office fixtures, strict summary/comparison primitives, bounded Markdown and revision-evidence measurement harnesses, retained-memory settling analysis, demo-vendor chunking, an Office render duration/peak-RSS measurement primitive bound to the canonical synthetic fixture lock, and a one-command suite that composes Markdown + revision/transition-evidence/envelope-canonicalization + autosave enqueue/coalescing/commit measurements. Packed mode measures exact modules extracted from one verified npm artifact, rejects Markdown, HTML, or envelope inputs that do not match the committed profile lock, and records package/runtime/source/reference-hardware provenance.
This work does not by itself establish production performance budgets, a supported large-document envelope, complete packed-editor/browser/IME/Yjs/Office matrix acquisition, leak freedom, or release acceptance.
Transition scenario fidelity and prerequisite stack
This Draft branch inherits #402 through canonical serialization owner #176 and is stacked on
fix/public-markdown-resource-options-175. Parent updates use ordinary merge history. The sole source conflict in that integration was resolved to match the owner exactly, preserving existing encoder reuse plus the owner's output bounds and hostile-option validation. Every benchmark, Office and demo-chunking delta remains intact. The prerequisite's CI, package and dependency changes remain owned by #402. None is shipped before protected integration.The original
transition-evidence-<profile>metric remains a no-op transition.--operation transition-changed --resulting-input <file>adds a distincttransition-changed-evidence-<profile>metric and enforces changed=true with unequal revision digests. The old scenario now requires changed=false with equal digests. Invalid/mismatched results, identical changed inputs, unsafe second-input paths, and output aliases fail before sample publication. The summarizer accepts the new ID while the comparator rejects cross-scenario comparisons.The packed suite accepts the explicit resulting input with or without HTML serialization, validates its committed byte count/digest, and CI acquires both transition scenarios. Earlier fixture bytes/hashes are preserved. The added fixture appends one plain paragraph; envelope fixtures do not become rich editor list/table/image trees merely because their text contains Markdown. See
benchmarks/README.mdfor the probe contract and limitations. This is measurement correctness and a separate synthetic baseline, not a speedup or a buyer-workload support claim.Executable contract
Benchmark producers and fixture generators fail closed on unsafe symlink/non-directory ancestors, unsafe leaf targets, pre-existing hard-linked outputs, bounded-input violations, and unverifiable package/module provenance. Packed mode requires the exact package tarball digest and package identity, the active Node runtime identity, the source checkout identity, and a reference-hardware identifier.
benchmarks/run-current-suite.mjsadditionally fails closed unless the benchmark checkout cleanliness guard succeeds before delegating tobenchmarks/run-current-suite-core.mjs. This closes the false-provenance class where modified or untracked source could otherwise produce evidence labeled with an unchangedHEADSHA. The failure is bounded and does not disclose dirty file paths.office/benchmarks/measure_render.pyaccepts only a committed synthetic Office fixture whose exact byte count and SHA-256 matchbenchmarks/office-fixtures.lock.json. It requires a clean checkout and verified source revision, rejects unbounded iteration counts before inspecting caller-selected input, uses a fresh Python child process for each render sample, and records render duration plus process peak RSS with p50/p75/p95/max summaries and runtime/reference-hardware provenance. Ordinary evidence contains fixture identity/hash/size and measurements, never the document body or caller path. It performs no network, credential, service, database, or model operation.Failure-contract / TDD lineage
A direct reproduction against the predecessor implementation established that
git rev-parse HEADalone cannot distinguish clean source from tracked or untracked worktree mutations. The narrow repair added a clean-checkout guard plus isolated temporary-repository contract tests for clean acceptance and dirty-state rejection; packed-suite tests exercise the delegated clean path. The test fixture is isolated from the repository checkout so parallel benchmark tests are not contaminated by a temporary dirty worktree.The Office measurement contract was added test-first on the canonical performance branch: it requires lock-bound synthetic input, stable privacy-safe rejection of arbitrary/private content, bounded iteration work, isolated repeated samples, duration and peak-RSS evidence, and source/runtime/reference-hardware provenance. The immediately superseded test-only generation was cancelled before terminal hosted RED evidence, so it is lineage rather than passing evidence; current-head verification must be read live and predecessor/cancelled runs never transfer.
Earlier RED/GREEN iterations established output-symlink, hard-link, measurement privacy/resource, suite ordering, failure-atomicity, packed-artifact identity, runtime identity, source-SHA, summary/comparison, and retained-memory-analysis contracts. Historical workflow results document lineage only; they never transfer to a later head or base.
First-invocation measurement generation
The revision, Markdown/HTML and autosave producers record every requested operation invocation, starting with the first. They no longer execute a preliminary operation without recording its duration. Existing result validation, setup/timer boundaries and privacy-safe failure publication remain intact. This measures operations, not whole-process startup.
Version 2 introduced first-invocation accounting. New JavaScript latency samples now use
contractVersion: 3, additionally derivinginputSha256from the same bounded bytes captured for measurement; changed transitions retain a distinct orderedresultingInputSha256. No-file autosave identifies its prepared synthetic revision-evidence payload. Summaries preserve these identities in JSON and text. Legacy versions 1 and 2 retain their original shapes and meanings; the comparator rejects cross-version or mismatched-input pairs before computing a verdict, even when profile labels match. No historical evidence is backfilled. Office, corpus-lock and suite-inventory contracts retain their independent versions. See the measurement accounting contract. No editor API or supported-performance promise changes.The first-invocation correction reproduced nine capped-call failures before the fix. A separate negative control removed only the generation-comparison guard and made both mixed-generation cases incorrectly succeed; restored tests reject both directions. These are measurement-correctness results, not latency improvement claims. Keep the historical raw evidence unchanged and start a fresh baseline for the new method.
Captured-input identity repair — 2026-09-06
The direct producer boundary could previously give different input documents comparable profile labels. This repair derives byte identities before the timer, preserves UTF-8 byte-order-mark identity, and verifies that later file replacement cannot relabel already captured input. Missing/malformed/extra/legacy-backfilled input identities and equal changed-transition identities fail closed. Hashes are not anonymization or authenticity proofs; no document content, path, new dependency, editor API change or timer change is introduced.
The source and failure lineage is recorded in captured-input identity research. Focused predecessor evidence is explicitly scoped: 86 checks passed at
c4fba276, and 14 captured/prepared-input checks plus TypeScript passed atc7d829d9. The broaderc7d829d9performance run failed: 158 passed, 8 failed, 1 worker RPC error across 41 files. Its unchanged timeout and null-child-exit failures remain retained, not relabeled as success.The current
75f195defull configured coverage run is terminal: 1,167 passed, 3 failed, 212 files, 674.14 seconds. The failures are unchanged deadlines in packed-artifact path stability and two corpus/source-provenance rejection tests; this is not green full coverage. The unchanged isolated run of the affected two files produced 6 passed and 2 different deadline failures; neither a timeout cause nor green full coverage is established. Current-head build and full package verification passed. Local Office verification passed 179 tests with configured statement/branch coverage and docstrings at 100% on one interpreter. The canonical source browser run passed all 70 tests across Chromium, Firefox and WebKit, including consensus, rune3377c8f-5f00-4f7c-932d-386269eefaad(package digest null). Packed-editor acceptance remains open; no predecessor result transfers. The Performance Evidence workflow is manually disabled and regular CI admission skips Draft PRs; neither was changed to obtain green checks. This remains Draft/Proposed and does not close #375 or establish the 20 ms target.Historical source reconciliation and local acquisition — 2026-09-05
Later exact-head validation: the unchanged
75f195defull coverage run serialized without our other heavy validation jobs passed all 1,170 tests in 212 files, 154.25 seconds, with all configured source coverage metrics at 100%. Earlier failed full and isolated runs are retained. This single repeat neither proves their cause nor establishes a product speedup. Standalone benchmark subprocess code remains outside configured source coverage. New captured-input measurements and current packed-browser acceptance remain separate work.On
a51cd5f8220b1100d515cc8eb05685ce89a942a6:7887c94822fc27ebd590627700c0c20c9b5c7f79d3b6c13b491fdb96d37b0321. Its Markdown/revision/autosave modules were byte-compared with the post-browser build. Source cleanliness and exact revision were checked before and after acquisition.refhw-sha256-5e7a3cf2c807457c88b1f5e87e401548c7496b50c0917217dae566ab8cf9493e. This is a shared development host, not accepted release hardware; its earlier load snapshot was elevated. These measurements begin a new source-lineage diagnostic baseline, not a causal optimization comparison.Changed-transition p95 values, milliseconds:
The stress scenario does not meet 20 ms. All slower runs remain in the denominator; neither a best run nor a changed/no-op scenario substitution is acceptance. The historical encoder-reuse 8.7% no-op result remains historical only. See the inherited canonical reconciliation record for retained-delta provenance and safety tests.
Configured TypeScript coverage excludes standalone benchmark subprocess code; its contract tests are not a claim of measured 100% subprocess coverage. Actual packed editor input/IME/mount/retained-memory acquisition, benchmark-script coverage, accepted support budgets and protected reference-hardware evidence remain open. Draft-admission skipped hosted jobs are non-passing; local results are not independent review or protected integration.
Remaining #375 acceptance work
This PR does not close #375 until protected evidence covers the applicable acceptance boundary. Remaining product work includes:
Browser/IME work, reference-host integration, release workflow, and organization-required workflow behavior remain with their established owners. Do not create competing source writers merely to make this performance lane appear complete.
Decision-time acceptance rule
Before any readiness, merge, release, closure, ownership, or support-envelope decision, independently refetch and reconcile at least:
main, this PR's exact head, its independently resolved live base, ancestry/divergence, mergeability, and changed paths;Pending, queued, in-progress, skipped-required, cancelled, absent, neutral, failed, stale, predecessor, wrong-checkout, synthetic-source-only, status-only, model-only, or vacuous evidence is non-passing. Automated comments/reviews are technical input, not qualifying independent approval. Any material head/base/ruleset movement invalidates the corresponding decision evidence.
Do not self-approve, weaken gates, transfer predecessor evidence, fabricate release identity, create a competing CI/security writer, or represent branch behavior as protected-main shipped truth.