Skip to content

Make cold-refresh performance snapshots statistically robust #509

Description

Problem

The performance snapshot currently runs 10 fresh PET processes against one shared cache directory. In practice, only iteration 1 exercises the cold cache; iterations 2-10 are warm. The interpolated P95 is therefore dominated by a single cold sample, which is not statistically robust.

PR #508 changes only the Python comparator's macOS budget tuple, tests, and documentation, yet unchanged Linux/Windows product paths produced these failures:

  • Run 31504572420, Linux: iteration 1 = 3,035ms, iterations 2-10 = 48-54ms, untimed diagnostics = 48-135ms; reported P95 = 1,694ms and failed.
  • Run 31505511672, Windows: iteration 1 = 16,567ms, iterations 2-10 = 133-150ms, untimed diagnostics = 140-176ms; reported P95 = 9,179ms and failed.

Inventories, medians, time-to-first, and all diagnostic locator/phase distributions remained healthy. This makes it difficult to distinguish a genuine PET cold-start regression from a one-off hosted-runner/cache/antivirus event.

Scope

  • Measure cold refresh with multiple independent cold-cache iterations rather than one sample.
  • Keep warm-cache P50/P95 separate from cold-refresh metrics.
  • Gate both cold and warm performance with budgets derived from repeated exact-base runs.
  • Preserve inventory checks, interpreter-timeout counts, and phase/locator diagnostics.
  • Associate diagnostics with the corresponding measured workload closely enough to attribute a cold failure.
  • Do not solve this by simply widening or disabling the existing budgets.

Acceptance criteria

  • Repeated unchanged-head hosted runs do not fail because of one uncorrelated cold sample.
  • An injected multi-second delay in every cold iteration fails a tested gate.
  • Warm median/tail regressions remain independently blocking.
  • Missing samples, inventory mismatches, and malformed metrics still fail closed.
  • Linux, Windows, and macOS workflows plus comparator unit tests pass.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestimportantIssue identified as high-priority

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions