Skip to content

test: add long-running runtime reliability tests #69

Description

@RentnerKev

Summary

Add bounded long-running reliability/soak coverage beyond the fast production smokes, using the
published v1.0.0-alpha.6 runtime as the baseline and the completed Beta feature set as the final
target.

Scenarios

Repeatedly exercise:

  • create, update, enable/disable, and delete Proxy Hosts and Redirect Hosts
  • certificate candidates, issuance/renewal/import, binding jobs, operation completion/failure,
    and retry recovery
  • Access Policy changes
  • CrowdSec allow/deny, degraded/recovery, and configuration updates if feat: integrate CrowdSec protection #64 ships
  • Forward Auth success/failure/recovery if feat: add Forward Auth access policies #65 ships
  • importer result-state durability/retry where meaningful if feat: add Nginx Proxy Manager importer #66 ships
  • rapid and repeated desired-state reconciliation
  • Caddy, controller, and web restart
  • transient database, controller, Caddy Admin API, upstream, and external-integration failures
  • backup/restore checkpoints where appropriate

Execution model

  • keep existing fast production smokes as the merge gate
  • provide representative short reliability checks for relevant pull requests
  • run the longer bounded harness on schedule, manual dispatch, and the final Beta release
    candidate
  • make duration, iteration count, concurrency, and resource capture explicit inputs
  • use deterministic seeds/fixtures and always clean up resources

Detect and assert

  • stuck or duplicate reconciliation
  • stale revision activation or state divergence
  • retry storms, unbounded queues, or requests that hang
  • process crashes and failed restart recovery
  • connection/file-descriptor growth and materially unbounded memory/CPU behavior where observable
  • loss/corruption of durable certificate, policy, or integration state

Do not invent a universal benchmark number. Record baselines, explain selected bounds, and gate on
correctness, recovery, and clearly justified resource limits.

Acceptance criteria

  • Deterministic, bounded soak harness with useful redacted diagnostics.
  • Scheduled/manual/release-candidate execution without multi-hour work on every PR.
  • Restart and transient-failure recovery leave desired and active state consistent.
  • No runaway retries, stuck reconciliation, crashes, or unbounded resource trend within the
    documented test envelope.
  • Shipped CrowdSec, Forward Auth, and importer result-state flows are represented.
  • No production secrets; cleanup succeeds after pass or failure.

Priority and sequencing

P0. Extends the existing reconciliation and production-smoke foundations rather than replacing
them. Final evidence is required by #75.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: ciContinuous integration and GitHub automation.area: runtimeCaddy and privileged controller runtime behavior.relatedRelated work that is not a confirmed duplicate.

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions