Add plugin health monitor workflow with failure streak alerts - #38
MikeGarciaAGM wants to merge 16 commits into
Conversation
|
Updated this branch to include the notification path the issue asked for: the monitor now posts a tagging comment on the tracking issue when a streak reaches the threshold, and the workflow grants issues: write for that path. I also hardened the artifact upload so a missing report file doesn't hide the real monitor error.\n\nValidation: a mocked GitHub API smoke test confirmed the script writes the report, posts the alert comment, and exits non-zero when findings are present. |
|
Small hardening follow-up: checkout now disables persisted credentials, and the monitor helper has docstrings plus a one-retry rate-limit-aware GitHub API wrapper. I re-ran the mocked API smoke test after the change; the script still writes the report, posts the alert comment, and exits non-zero when findings exist. |
|
Small follow-up on the plugin health monitor: the alert path now skips posting an identical comment if the exact alert body already exists on the tracking issue, so daily runs won't spam repeated failure reports. I verified the duplicate path with a mocked run; the monitor still exits non-zero for findings, but it reuses the existing alert URL instead of re-posting. |
|
Small follow-up: the alert comment now includes the latest failed run URL plus failed jobs/failed step names, so maintainers can jump straight to the useful failure context from the issue thread. I also verified the new path locally with a mocked GitHub API smoke test. |
|
Small follow-up on the plugin health monitor: I added a GitHub Actions job summary so maintainers can see repositories scanned, findings count, threshold, and the alert comment link directly in the run view. It doesn't change behavior; it just makes review faster and less clicky. |
|
Small follow-up: I added usage notes to the profile README so maintainers can see the monitor's purpose, env vars, and local run command at a glance. This doesn't change behavior; it just makes the lane easier to review and operate. |
|
Small follow-up: I split the plugin health monitor helpers into a tiny module and added Node built-in tests for streak counting, alert formatting, and failure-context rendering. The workflow now runs those tests before the daily scan, so the lane is easier to verify without changing the bounty scope. |
|
Small portability follow-up: the workflow no longer depends on jq for the final threshold gate. It now uses Node to parse the report, which keeps the CI path self-contained and a bit less brittle on runners. |
|
Small follow-up: I extracted the duplicate-alert lookup into a tiny helper and added node:test coverage for the exact-match path. The branch is pushed, and the monitor path still passes locally with 5/5 tests. |
|
Small follow-up: duplicate-alert protection is now stable across reruns. The alert body carries a hidden stable key, and the lookup also normalizes legacy comments that predate the key, so the monitor can reuse an existing alert even when run URLs/log links change. Local validation still passes with 6/6 node:test cases. |
|
Small but important alignment tweak: the monitor now only counts workflow_dispatch runs, which matches the bounty's manual-trigger scope and reduces false positives from unrelated events. I also documented that scope in the profile README. The branch is pushed and the node:test suite still passes 7/7 locally. |
|
Rechecking this lane: PR #38 still looks mergeable-clean on my side, with the plugin health monitor scope kept tight around the failure-streak alert path. The branch includes the workflow, helper module, tests, and the alert-comment path that the issue asked for, so I believe this is still the intended review/merge surface for #5886. If there is any final wording or formatting change you want before merge, I can do that quickly; otherwise I’d appreciate a review pass when you have a slot. |
Summary
Resolves #12.
Implements a daily plugin health monitor that scans public repositories in
ubiquity-os-marketplace, counts onlyworkflow_dispatchruns so it stays aligned with the bounty's manual-trigger signal, detects workflows with>= 10consecutive failures, and posts an alert comment on the tracking issue when findings exist.What was added
.github/workflows/plugin-health-monitor.yml0 8 * * *) and on manual dispatch.GITHUB_TOKEN.issues: writeso findings can be posted back to the tracking issue.profile/plugin-health-report.jsonas an artifact.profile/scripts/plugin-health-monitor.mjsworkflow_dispatchruns before counting failures.@0x4007and@gentlementlegenwhen findings exist.profile/scripts/plugin-health-monitor-lib.mjsprofile/scripts/plugin-health-monitor.test.mjsprofile/README.mdWhy this matches the bounty ask
Notes