Skip to content

Add plugin health monitor workflow with failure streak alerts - #38

Open
MikeGarciaAGM wants to merge 16 commits into
ubiquity-os:mainfrom
MikeGarciaAGM:feat/plugin-health-monitor-5886
Open

MikeGarciaAGM wants to merge 16 commits into
ubiquity-os:mainfrom
MikeGarciaAGM:feat/plugin-health-monitor-5886

Conversation

@MikeGarciaAGM

@MikeGarciaAGM MikeGarciaAGM commented Jun 2, 2026 •

Copy link
Copy Markdown

Summary

Resolves #12.

Implements a daily plugin health monitor that scans public repositories in ubiquity-os-marketplace, counts only workflow_dispatch runs so it stays aligned with the bounty's manual-trigger signal, detects workflows with >= 10 consecutive failures, and posts an alert comment on the tracking issue when findings exist.

What was added

  1. .github/workflows/plugin-health-monitor.yml

    • Runs daily (0 8 * * *) and on manual dispatch.
    • Runs unit tests before the monitor.
    • Executes the monitor script in read-only mode using GITHUB_TOKEN.
    • Grants issues: write so findings can be posted back to the tracking issue.
    • Uploads profile/plugin-health-report.json as an artifact.
    • Publishes a concise job summary with scan counts and top findings.
    • Fails the job if findings exist.
  2. profile/scripts/plugin-health-monitor.mjs

    • Lists org repos from GitHub API.
    • Enumerates workflows and latest workflow runs.
    • Filters to workflow_dispatch runs before counting failures.
    • Calculates consecutive failure streaks.
    • Emits a structured JSON report.
    • Posts a concise alert comment tagging @0x4007 and @gentlementlegen when findings exist.
  3. profile/scripts/plugin-health-monitor-lib.mjs

    • Extracts pure helpers for streak counting and alert formatting.
    • Keeps alert bodies stable via a hidden key plus legacy-body normalization.
    • Keeps the failure-context rendering testable in isolation.
  4. profile/scripts/plugin-health-monitor.test.mjs

    • Covers streak counting.
    • Covers workflow_dispatch filtering.
    • Covers failure-context rendering from mocked API data.
    • Covers alert formatting, duplicate alert reuse, and report-path output.
  5. profile/README.md

    • Adds usage notes, environment variables, local test command, and the manual-trigger scope.

Why this matches the bounty ask

  • Checks plugin ecosystem health on a daily schedule.
  • Flags sustained failure streaks (10+) across plugin repos.
  • Stays scoped to the bounty's manual-trigger signal instead of unrelated events.
  • Keeps the implementation simple and non-config-heavy.
  • Includes tests and reviewer-facing output to reduce review friction.

Notes

  • The monitor is read-only against the marketplace repos themselves, but it does comment on the tracking issue when a failure streak threshold is exceeded.
  • Duplicate alert comments are de-duplicated using a stable hidden key, with fallback normalization for legacy comments.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Updated this branch to include the notification path the issue asked for: the monitor now posts a tagging comment on the tracking issue when a streak reaches the threshold, and the workflow grants issues: write for that path. I also hardened the artifact upload so a missing report file doesn't hide the real monitor error.\n\nValidation: a mocked GitHub API smoke test confirmed the script writes the report, posts the alert comment, and exits non-zero when findings are present.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small hardening follow-up: checkout now disables persisted credentials, and the monitor helper has docstrings plus a one-retry rate-limit-aware GitHub API wrapper. I re-ran the mocked API smoke test after the change; the script still writes the report, posts the alert comment, and exits non-zero when findings exist.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up on the plugin health monitor: the alert path now skips posting an identical comment if the exact alert body already exists on the tracking issue, so daily runs won't spam repeated failure reports. I verified the duplicate path with a mocked run; the monitor still exits non-zero for findings, but it reuses the existing alert URL instead of re-posting.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up: the alert comment now includes the latest failed run URL plus failed jobs/failed step names, so maintainers can jump straight to the useful failure context from the issue thread. I also verified the new path locally with a mocked GitHub API smoke test.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up on the plugin health monitor: I added a GitHub Actions job summary so maintainers can see repositories scanned, findings count, threshold, and the alert comment link directly in the run view. It doesn't change behavior; it just makes review faster and less clicky.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up: I added usage notes to the profile README so maintainers can see the monitor's purpose, env vars, and local run command at a glance. This doesn't change behavior; it just makes the lane easier to review and operate.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up: I split the plugin health monitor helpers into a tiny module and added Node built-in tests for streak counting, alert formatting, and failure-context rendering. The workflow now runs those tests before the daily scan, so the lane is easier to verify without changing the bounty scope.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small portability follow-up: the workflow no longer depends on jq for the final threshold gate. It now uses Node to parse the report, which keeps the CI path self-contained and a bit less brittle on runners.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up: I extracted the duplicate-alert lookup into a tiny helper and added node:test coverage for the exact-match path. The branch is pushed, and the monitor path still passes locally with 5/5 tests.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small follow-up: duplicate-alert protection is now stable across reruns. The alert body carries a hidden stable key, and the lookup also normalizes legacy comments that predate the key, so the monitor can reuse an existing alert even when run URLs/log links change. Local validation still passes with 6/6 node:test cases.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Small but important alignment tweak: the monitor now only counts workflow_dispatch runs, which matches the bounty's manual-trigger scope and reduces false positives from unrelated events. I also documented that scope in the profile README. The branch is pushed and the node:test suite still passes 7/7 locally.

@MikeGarciaAGM

Copy link
Copy Markdown
Author

Rechecking this lane: PR #38 still looks mergeable-clean on my side, with the plugin health monitor scope kept tight around the failure-streak alert path. The branch includes the workflow, helper module, tests, and the alert-comment path that the issue asked for, so I believe this is still the intended review/merge surface for #5886. If there is any final wording or formatting change you want before merge, I can do that quickly; otherwise I’d appreciate a review pass when you have a slot.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants