Follow-up from the hive review of #528 (at 8f6e895). #528 drops pull_request and merge_group runs from the success-rate sample (completed_json in .github/workflows/factory-health.yml, compute_pipeline_health in scripts/monitor_pipeline.py), but the back-to-back failure streak is still computed from the unfiltered run list (the consecutive_failures jq in factory-health.yml and _consecutive_failures() in monitor_pipeline.py).
On the total < MIN_RUNS branch that means:
Suggested fix: apply the same .event exclusion to both streak computations, keep the bats CLASSIFY_LOGIC copy in sync, and add a test for each direction. Kept out of #528 to hold that PR to its reviewed scope.
Follow-up from the hive review of #528 (at 8f6e895). #528 drops
pull_requestandmerge_groupruns from the success-rate sample (completed_jsonin.github/workflows/factory-health.yml,compute_pipeline_healthinscripts/monitor_pipeline.py), but the back-to-back failure streak is still computed from the unfiltered run list (theconsecutive_failuresjq in factory-health.yml and_consecutive_failures()in monitor_pipeline.py).On the
total < MIN_RUNSbranch that means:alert, the false-alert class fix(factory): exclude pre-merge runs from factory health success rate #528 removes;pushfailures (the fix(factory): [projectbluefin/bluefin] [E2E] success rate dropped to 0% (24h window) #479 case), downgradingalerttolow-sample.Suggested fix: apply the same
.eventexclusion to both streak computations, keep the bats CLASSIFY_LOGIC copy in sync, and add a test for each direction. Kept out of #528 to hold that PR to its reviewed scope.