From 30d3b51e078122d7e9d3d180c04be49f8a006df3 Mon Sep 17 00:00:00 2001 From: operations Date: Mon, 5 Oct 2026 18:55:31 -0400 Subject: [PATCH] =?UTF-8?q?=F0=9F=93=96=20docs:=20add=20site=20rollback=20?= =?UTF-8?q?runbook?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Signed-off-by: operations --- changelog.d/added-site-rollback-runbook.md | 1 + runbooks/README.md | 7 +++ runbooks/site-rollback.md | 52 ++++++++++++++++++++++ 3 files changed, 60 insertions(+) create mode 100644 changelog.d/added-site-rollback-runbook.md create mode 100644 runbooks/README.md create mode 100644 runbooks/site-rollback.md diff --git a/changelog.d/added-site-rollback-runbook.md b/changelog.d/added-site-rollback-runbook.md new file mode 100644 index 0000000..d839cbc --- /dev/null +++ b/changelog.d/added-site-rollback-runbook.md @@ -0,0 +1 @@ +- Add a docs site rollback runbook under `runbooks/` covering Netlify deploy restore and fix-forward steps. diff --git a/runbooks/README.md b/runbooks/README.md new file mode 100644 index 0000000..342370d --- /dev/null +++ b/runbooks/README.md @@ -0,0 +1,7 @@ +# Runbooks + +Start here when the docs site (docs.hivecommons.dev) is broken after a deploy. + +| Symptom | Runbook | +| --- | --- | +| A deploy shipped broken pages, missing content, or a failed build is blocking updates | [site-rollback.md](site-rollback.md) | diff --git a/runbooks/site-rollback.md b/runbooks/site-rollback.md new file mode 100644 index 0000000..fbc91de --- /dev/null +++ b/runbooks/site-rollback.md @@ -0,0 +1,52 @@ +# Docs site rollback + +Use this when a production deploy of the docs site serves broken pages, wrong +or missing synced content, or broken navigation. The site is built by Netlify +(`netlify.toml`) from `main`; every deploy is kept and can be re-published. + +## 1. Triage (5 minutes) + +1. Confirm the symptom on the live site, not only on a deploy preview. +2. Find the first bad deploy in the Netlify deploy list and the commit it built. +3. Decide whether the cause is site code in this repo or content synced at + build time from a project repo (hive, hotshot, pluk, rationguard, + promptargs, dibs). The `prebuild` script falls back to committed content + when a sync is unavailable, so a sync failure alone does not fail the build. + +## 2. Mitigate: restore the last good deploy + +1. In Netlify, open the site's Deploys list and select the last deploy that + rendered correctly. +2. Choose **Publish deploy**. This switches production to that build without + a rebuild and also pauses auto-publishing, so a new push to `main` will not + overwrite the rollback. +3. Verify the live site: the home page, `/docs/hive/overview/introduction`, + a page from each synced project, and a 404 URL (the branded + `public/404.html` should render). + +## 3. Fix forward + +1. Revert the offending commit on a branch, or fix it, and open a PR + (DCO-signed: `git commit -s`). +2. Run the local gates from `CONTRIBUTING.md` (`npm run build`, + `npm run check-links`, `npx vitest run --coverage`). +3. After the PR merges and the new deploy is verified on its deploy preview, + choose **Start auto publishing** in Netlify to resume normal deploys. + +## 4. If bad content came from a project repo + +Fix it in the source repo (content is synced from there), then trigger a +rebuild of this site. Do not hand-edit synced content under `docs/`; the next +sync overwrites it. + +## 5. NPS relay function + +`netlify/functions/nps.mts` is deployed with the site, so rolling back the +site also rolls back the function. Hub pull routes answering 503 means the +`NPS_RELAY_HUB_SECRET_SHA256` site environment variable is unset; that is a +configuration issue, not a deploy regression. + +## 6. After the incident + +Record what happened, the time to restore, and the follow-up fixes. Use the +postmortem template in `hivecommons/hive` if the outage reached users.