Repository navigation
Assess JVM warm-up and fair cohort admission before an official Java-inclusive benchmark #73
Description
Activity
- added a commit that references this issue
on Sep 30, 2026 tappe9 commented
on Sep 30, 2026 OwnerAuthorMore actionsBounded diagnostic completed and audited — 2026-09-30
The one approved JVM diagnostic run 36686348909, attempt 1 completed successfully in about 62 minutes. No rerun or extra measurement was performed.
Provenance and original evidence
- Measured source:
cee473091477cca328871b1349a2090b793cc588(PR feat: add bounded JVM warm-up diagnostic #80); tree6914510a90f70bf9000b62abac185ce4f99a3587 - Raw artifact 11086656632,
jvm-diagnostic-36686348909-1, 581,024 bytes; current API expiry: 2026-12-29T07:53:35Z - ZIP SHA-256:
07c206f911d52f6cae35e4470844f428ff91913c31b72741626273021d210073, matching GitHub's digest - Independent audit verified all 703 manifest hashes and the exact 704-file study inventory; source/tree, all 12 recorded lock hashes, versions and image identities agree
Completeness
All nine rotated clean-start attempts completed and reported successful scoped cleanup. All 27 warm-ups and 108 measured windows were retained. Strict re-parsing of every oha file found 26,330,926 HTTP 200 responses, no transport/non-200 errors. All 1,750 Docker samples agree with their nine distinct API container/project identities and recorded resource configuration. All 27 endpoint assessments recompute exactly.Observed outcome under the unchanged, predeclared rules
- Java JSON: first 30-second window after the five-second warm-up = 4,856–5,328 req/s, versus later-three medians of 14,190–14,737 req/s; first-window deficit 63.65–66.00%, repeated across all three starts
- Java DB: first window = 3,694–4,533 req/s, versus later-three medians of 6,203–6,339 req/s; deficit 28.41–41.73%
- All Node/Go first-window throughput differences were below 0.90%. However, overall support is false for all three implementations under the conservative stability rules: Java 0/9, Node 3/9, Go 3/9 supported endpoint/attempts. Fifteen of 27 endpoint/attempts fail only the strict late-trend condition; this must not be relabeled as proof that the controls need longer warm-up
The Java differences are descriptive comparisons against later windows, not proven steady-state values. Contract checks precede warm-up, later endpoints inherit earlier activity, requests drain between oha invocations, and sampling/orchestration adds wall time (up to 2.09 seconds beyond oha elapsed per window). These are successive measured windows, not exact process-age intervals or gapless load. No JIT/GC events were collected, so the cause cannot be assigned to JIT or GC. Checksums are not signatures; raw repeated state checks and PostgreSQL identity/fixture initialization are not independently retained, although the successful harness checks them.
Recommendation and remaining decision
Keep Java registered-only for now. These repeatable early-window deficits do not support adding Java to a warmed-state comparison under the unchanged five-second budget, and they do not establish any specific replacement duration. The owner still needs to choose whether to stop here or authorize a separately predeclared common-profile investigation, eventually covering all nine implementations. Do not retrospectively relax these criteria or mix these diagnostic rows into old official results.This issue stays open pending that decision. Existing official profiles/cohorts, applications, pins, README/results/history and main source remain unchanged by execution; no official benchmark or publication occurred.
- Measured source:
- added a commit that references this issue
on Oct 1, 2026 tappe9 commented
on Oct 1, 2026 OwnerAuthorMore actionsCommon warm-up screen v1 completed and audited — 2026-10-01
Outcome: inconclusive. None of the predeclared 5/35/65-second candidates qualifies across all 18 traces. Keep Java registered-only.
The one approved additional screen run 36804467161, attempt 1 completed successfully in 82 minutes 20 seconds. No diagnostic rerun, adaptive extension or additional measurement was performed.
Provenance and completeness
- Source:
ea6a7962d74b0b608d948898f09b78eae9559799(PR #81); tree:0487c15ec9f10a5aa6e06d1053aeb031a3034ec8 - Predeclared protocol:
common-warmup-screen-v1 - Raw artifact 11139592774:
common-warmup-screen-36804467161-1, 919,857 bytes; current API expiry: 2026-12-30T02:09:12Z - ZIP SHA-256:
2a228ae7e58d80309cebc2cfcf688fb474e8cb60cfb9c1457a19763ab8362191, matching GitHub's digest
An independent audit found no discrepancies: exactly 1,005 files, all 1,003 manifest-covered hashes, all 12 recorded lock hashes, source/tree and image identities verified. Three sacrificial validations retained 42 contract checks. All 18 endpoint-isolated traces completed, with 18 warm-ups and 144 measured windows. All 21 distinct API/PostgreSQL pairs reported successful scoped cleanup; measured API images match their sacrificially validated images.
Strict independent re-parsing of all 162 load outputs found 35,162,131 HTTP 200 responses and zero transport/non-200 load errors. All 2,308 Docker stats samples and retained state/fixture/readiness/timing evidence were checked. Separately, startup readiness polling retained 141 transient transport failures before 21 successful
/healthresponses; these are not measured-load failures or trace retries. All per-trace and aggregate assessments recompute.Results under the fixed rules
Counts below are passing traces / total traces. Each individual endpoint has two clean starts.
Implementation/endpoints 5s 35s 65s Java /json 0/2 2/2 2/2 Java /db/42 0/2 0/2 0/2 Java /cpu 0/2 0/2 1/2 Node: all three endpoints 6/6 6/6 6/6 Go: /json and /cpu 4/4 4/4 4/4 Go /db/42 0/2 0/2 0/2 Total 10/18 12/18 13/18 The common screen requires 18/18, not a majority.
- Java JSON supports the 35s/65s candidates within this segmented-load screen in both starts.
- Java DB remains changing within the candidate windows: at the 65s candidate, window 3 throughput is still 39.62% and 30.68% below the corresponding final-three reference median.
- Java CPU passes 65s in only one start; the other has window 3 throughput 5.65% below its reference, exceeding the fixed 5% bound.
- Go DB reference p95 spreads are 13.92% and 12.56%, exceeding the fixed 10% reference-stability bound. This does not establish that Go needs more warm-up or that a longer duration alone would resolve the variation.
Interpretation and stopping decision
These candidates are cumulative sending-time budgets, not exact process ages or independently tested uninterrupted warm-up durations. Each oha invocation drains requests; readiness, state checks, sampling and orchestration introduce gaps. There is one host and only two starts per pair, with host/build caches retained and observer overhead not quantified. Passing a screen is not a steady-state or statistical-equivalence claim. No JIT/GC events were collected, so no causal attribution is justified.
The initial JVM study and its original strict-trend interpretation remain intact. This new protocol was fixed before collection; its different rule must not be used to relabel earlier results. Cleanup is journaled and enforced by the source, but separate raw post-cleanup resource inventories are not retained. Dependency/runtime versions are source-derived; PostgreSQL's server version is directly queried.
Stop at this approved budget. No common replacement duration is established, and no nine-implementation confirmation or official rollout is authorized by this result. Existing official profiles/cohorts, applications, pins, README/results/history are unchanged. Issue #73 remains open for the owner's next decision.
CI status note
Exact-head PR CI passed. Post-merge CI 36802189456, attempt 3 has all 13 jobs, including
required, completed successfully. As checked at 03:46 UTC, GitHub still reports the overall run asin_progress; a subsequent Pages run has not appeared. Earlier CI attempts retain unrelated dashboard-navigation and temporary-directory-cleanup failures. This status inconsistency is separate from the successfully completed diagnostic; no additional CI or diagnostic rerun is being used to repair the display state.- Source:
tappe9 commented
on Oct 1, 2026 OwnerAuthorMore actionsOwner decision and completion — 2026-10-01
Keep Java registered and CI-covered only. Retain the current eight-stack official cohort. No further warm-up investigation is planned for this round.
The owner approved closing this diagnostic/decision issue and proceeding with the bilingual README cleanup in #74. This decision does not admit Java to an official comparison, choose a replacement warm-up duration, or change any existing profile, cohort, runtime pin or published result.
Evidence and acceptance
- The initial diagnostic audit and common warm-up screen audit retain source/image/version identities, all attempts and predeclared interpretation rules. The initial study and the later screen remain separate evidence.
- Warm-up adequacy remains inconclusive. None of the 5/35/65-second screen candidates passes all 18 traces. Selected passing samples do not establish a common duration, steady state or a causal JIT/GC explanation.
- Concurrency and instrumentation limits distinguish one JVM process from request-handling threads and document the shared resource/pool caps without runtime tuning.
- The completed decision is the explicitly allowed registered-only outcome. Both frozen cohort definitions and all historical measurements remain intact. Diagnostic success is not an official performance result.
- No Java-inclusive rollout is justified by this evidence. Any future reconsideration needs a separately authorized plan and budget: a common profile for every selected member, the full cohort measured sequentially in one job, all three samples per endpoint retained, exact-source publication and the producer-to-Pages verification described by Verify Pages handoff on the next justified manual official benchmark #59. This closure neither performs nor substitutes for that work.
The bounded diagnostics and the owner's decision resolve this issue's acceptance criteria. Closing as completed refers to the investigation and decision, not to Java admission. The implementation, protocols and existing diagnostic evidence remain available for future reconsideration, subject to the artifacts' documented retention periods. No new measurement, rerun or publication is part of this closure.
Completion — 2026-10-01
The owner chose to keep Java registered/CI-covered only and retain the current eight-stack official cohort. No further investigation is planned for this round. The diagnostic evidence remains inconclusive; no common replacement warm-up or Java-inclusive rollout is established. See the decision and acceptance review. Completion refers to this diagnostic/decision issue, not Java admission.
Priority and purpose
P2 — diagnostic and decision issue, not approval to activate a cohort or run an official benchmark. Determine whether and under which explicit common conditions Java / Spring Boot should join the next official comparison.
Verified baseline
Rechecked main on 2026-09-18:
9b92c74cba34bd5a4bfc389075b9bef135131034. Refresh current repository state before work.PR #69 registered Java as the ninth implementation without adding it to either frozen official cohort. The Java guide explicitly requires a separate cohort/warm-up decision. The methodology uses a five-second common warm-up, 50 connections and three 30-second measured runs per endpoint, with 1 CPU, 512 MiB and pool maximum 10.
There is no existing official Java result and no evidence here proving that five seconds is sufficient or insufficient. Correctness CI is not evidence of steady-state performance. #72 separately tracks exact Java-version compatibility in report/site validation.
Questions to answer
Scope
Acceptance criteria
Dependencies and boundaries
#72 must be completed before any Java-inclusive official report can be admitted. Diagnostic work need not wait for #72 when it uses an explicitly non-publishable format. #70 is the recommended earlier investigation, not an artificial hard dependency on collecting Java evidence. #52 remains deferred and is not a prerequisite.
No official workflow dispatch, automatic result publication, cohort activation, recurring benchmark or permission/credential change is authorized here. Creating this issue is not authorization to close #59. Follow the local-only investigation and Superpowers artifact policy in
AGENTS.md.