From aa3404532399923052e1dd8e093c6c5f73651f3c Mon Sep 17 00:00:00 2001 From: Andrew Anderson Date: Mon, 5 Oct 2026 17:28:19 -0400 Subject: [PATCH] =?UTF-8?q?=F0=9F=93=96=20hive:=20resync=20docs=20from=20v?= =?UTF-8?q?5=20(Level=206=20guide:=20per-repo=20opt-out,=20issue=20claims,?= =?UTF-8?q?=20review=20notice=20removed)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pulls hivecommons/hive#10704 into the committed snapshot and triggers a production build; the site only rebuilds on pushes to this repo, so the live guide still showed the pre-#10704 text. Also refreshes 22 other synced pages that drifted since the last snapshot and adds ADR 0016/0017. Signed-off-by: Andrew Anderson --- docs/content/hive/acmm-policy-matrix.md | 56 +++++-- .../hive/adr/0005-forge-abstraction.md | 2 +- .../adr/0010-escalation-circuit-breaker.md | 23 ++- docs/content/hive/adr/readme.md | 1 + docs/content/hive/agent-configuration.md | 137 ++++++++++++++-- docs/content/hive/architecture.md | 25 ++- docs/content/hive/backup-dr.md | 19 ++- docs/content/hive/contributor-relay.md | 155 +++++++++++++++--- docs/content/hive/documentation-map.md | 31 +++- docs/content/hive/env-vars.md | 4 +- docs/content/hive/getting-started.md | 44 ++--- docs/content/hive/hivectl.md | 98 ++++++++--- docs/content/hive/integration-guide.md | 5 +- .../integrations/work-source-providers.md | 2 +- docs/content/hive/landscape.md | 105 +++++++++++- docs/content/hive/manual-provisioning.md | 92 +++++++++-- docs/content/hive/net-admin-requirement.md | 21 ++- docs/content/hive/release-channels.md | 16 +- docs/content/hive/roadmap.md | 2 +- docs/content/hive/running-at-level-6.md | 5 +- docs/content/hive/securing-your-hive.md | 11 +- docs/content/hive/troubleshooting.md | 14 ++ docs/content/hive/work-sources.md | 27 ++- 23 files changed, 760 insertions(+), 135 deletions(-) diff --git a/docs/content/hive/acmm-policy-matrix.md b/docs/content/hive/acmm-policy-matrix.md index 45114f8..06ab3b1 100644 --- a/docs/content/hive/acmm-policy-matrix.md +++ b/docs/content/hive/acmm-policy-matrix.md @@ -15,7 +15,7 @@ Each agent runs in one of four modes, controlling what actions it can take on Gi - **Advisory**: Agent observes and records findings as beads on the dashboard. No GitHub interaction. - **Measured**: Agent can file GitHub issues to make findings visible to the team. No code changes. -- **Holdgated**: Agent can write code and open PRs, but every PR gets a `hold` label. A human must review and remove `hold` before merge. Agent never merges. One exception: a hold the hive applied *for level reasons* is released automatically once the current level no longer calls for it — see the promotion note below. A hold **you** applied is never removed automatically. +- **Holdgated**: Agent can write code and open PRs, but every PR gets a `hold` label. A human must review and remove `hold` before merge. Agent never merges. Level-applied holds are not released automatically on promotion. A human removes them after review; a deliberate one-off `release_level_holds: true` request can release attributable Hive level holds — see the promotion note below. A hold **you** applied is never removed automatically. - **Full**: Agent operates autonomously — opens PRs and merges on green CI. Highest trust level. ## ACMM Levels @@ -68,9 +68,9 @@ Delivery agents open GitHub issues — bugs, docs gaps, CI problems, security vu | **sec-check** | **holdgated** | `sec-check-holdgated.md` | | brainstorm | advisory | `brainstorm-advisory.md` | -### L5 — Semi-Autonomous (Semi-Automated) (12 agents) +### L5 — Semi-Autonomous (Semi-Automated) (13 agents) -Agents open issues AND pull requests. All PRs get a hold label — humans batch-review and approve. Architect produces RFCs, strategist coordinates across agents, and reviewer works the hold-gated PR queue every 30 minutes. The system proposes; it does not merge autonomously. +Agents open issues AND pull requests. Agent PRs get literal `hold` from the level gate — humans batch-review and approve. The dashboard `hive-pause/` label is a separate manual hold, and `hive/` is provenance only. Architect produces RFCs, strategist coordinates across agents, and reviewer works the hold-gated PR queue every 30 minutes, and adjudicator works escalated (`needs-human`) hive PRs through the reviewer lane every 30 minutes (recommend-close only; closing is operator-only below L6). The system proposes; it does not merge autonomously. | Agent | Mode | Template | |-------|------|----------| @@ -83,13 +83,14 @@ Agents open issues AND pull requests. All PRs get a hold label — humans batch- | architect | holdgated | `architect-holdgated.md` | | strategist | holdgated | `strategist-holdgated.md` | | reviewer | converse | `reviewer-queue.md` | +| adjudicator | issues+prs | `reviewer-lane.md` (by role; no `kick_template`) | | telemetry (paused) | holdgated | `telemetry-holdgated.md` | | operations (paused) | holdgated | `operations-holdgated.md` | | brainstorm | advisory | `brainstorm-advisory.md` | -### L6 — Fully Autonomous (13 agents) +### L6 — Fully Autonomous (14 agents) -Existing autonomous lanes can open issues, create PRs, and auto-merge on green CI. No hold label. Outreach handles community engagement. Reviewer stays advisory even here — its `requires_human` verdict is what pulls a PR out of the auto-merge lane. Telemetry and operations remain paused and use `ISSUES_AND_PRS`, so they never merge their own PRs. +Existing autonomous lanes can open issues, create PRs, and auto-merge on green CI. Auto-merge is off below L6; switching to L6 turns it on for every active repository, after which owners can toggle repositories individually. Non-outreach L6 PRs do not get the level hold, but outreach PRs are still held for human review. Outreach handles community engagement. Reviewer stays advisory even here: `requires_human` alone does not stop the App auto-merge sweep. `review.require_approval: true` gates scanner eligibility on same-head approval, not the App sweep; use actual holds or GitHub branch rules for a universal merge stop. Adjudicator is the reviewer lane: it repairs, de-escalates, or recommends closing (and may close) escalated `needs-human` hive PRs, and never merges. Telemetry and operations remain paused and use `ISSUES_AND_PRS`, so they never merge their own PRs. | Agent | Mode | Template | |-------|------|----------| @@ -103,6 +104,7 @@ Existing autonomous lanes can open issues, create PRs, and auto-merge on green C | strategist | full | `strategist-full.md` | | outreach | full | `outreach-full.md` | | reviewer | converse | `reviewer-queue.md` | +| adjudicator | issues+prs | `reviewer-lane.md` (by role; no `kick_template`) | | telemetry (paused) | full | `telemetry-full.md` | | operations (paused) | full | `operations-full.md` | | brainstorm | advisory | `brainstorm-advisory.md` | @@ -113,13 +115,13 @@ Existing autonomous lanes can open issues, create PRs, and auto-merge on green C ## Key Rules -1. **All PRs are holdgated below L6.** No agent can auto-merge unless running at L6 (Fully Autonomous). +1. **Level holds use literal `hold`.** Agent PRs below L6 are hold-gated with `hold`, and outreach PRs are held at every level. The dashboard `hive-pause/` label is for manual item holds. See [Hive Labels and Control Signals](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md). 2. **Advisory agents never get GH auth.** The `${GH_AUTH}` template variable is only injected into measured, holdgated, full, and converse templates. The converse tier is the one place an agent writes to GitHub without sitting on the mode ladder: `reviewer` is `mode: ADVISORY` plus the orthogonal `converse` capability ([#4492](https://github.com/hivecommons/hive/issues/4492)), which grants comments and PR reviews and nothing else — no issue creation, no relabelling, no push, no merge. It needs the auth block because posting a review *is* a GitHub write. 3. **Supervisor uses no-GitHub advisory mode.** At every level, supervisor uses `supervisor-nogithub.md` in the built-in ACMM packs — it monitors agent health, not code. 4. **Mode escalation is per-agent.** At L4, some agents are measured (issues only) while others are holdgated (issues + PRs). The level defines the mix. 5. **Knowledge priming works at all levels.** The `${KNOWLEDGE}` template variable injects relevant facts from git sources and wiki layers regardless of the agent's mode. 6. **Brainstorm is always advisory.** It produces KB facts and beads, never GitHub issues or PRs. Its role evolves from inception (L1) to ongoing ideation (L2+), but its mode stays advisory at all levels. -7. **Reviewer is L5/L6-only by default and never merges.** It joined the L5 and L6 rosters in [#8023](https://github.com/hivecommons/hive/issues/8023) at a 30-minute cadence in every governor mode. Below L5 no pack lists it, so an operator who wants repo-grounded PR review creates it by hand and a pack apply leaves that agent's mode, model, backend, and pause state alone. Its mode stays `ADVISORY` at both levels, including L6: it reads the queue, comments, and returns a verdict, and it is that verdict — `requires_human` or `reject` — that pulls a PR out of the auto-merge lane. +7. **Reviewer is L5/L6-only by default and never merges.** It joined the L5 and L6 rosters in [#8023](https://github.com/hivecommons/hive/issues/8023) at a 30-minute cadence in every governor mode. Below L5 no pack lists it, so an operator who wants repo-grounded PR review creates it by hand and a pack apply leaves that agent's mode, model, backend, and pause state alone. Its mode stays `ADVISORY` at both levels, including L6: it reads the queue, comments, and returns a verdict. `requires_human` or `reject` stops scanner eligibility only when `review.require_approval: true` requires same-head approval; the App auto-merge sweep does not read those verdicts. Use a hold or GitHub branch rule when merging must stop. See [Running at Level 6](/docs/hive/running-at-level-6#does-the-hive-reviewer-stop-a-merge). 8. **Telemetry and operations are L5/L6-only opt-in agents.** Below L5 they are absent from the pack roster and dashboard, do not spawn panes, and cannot be kicked. At L5–L6 they use a paused cadence in every governor mode until an operator opts in; they may open issues and PRs but never merge. ## ioscan hardening defaults per level @@ -200,9 +202,15 @@ Promoting or demoting a running hive between levels is a single operation — th hive reconciles its agent roster and per-agent modes to match the target level. **From the dashboard:** open the Governor config and set the ACMM level. This is -the normal path. +the normal path. When a change makes the L6 self-authored merge sweep eligible, +Hive turns auto-merge on for every active repository, and the dashboard shows an informational modal listing held App-authored PRs and +the watched repositories' auto-merge switches so owners can switch off repos that should not participate; it does not release holds. **Over the API:** `PUT /api/packs/level` with `{"level": N}` where N is 1–6. +Level-applied `hold` labels are never released automatically on a level change. +For a deliberate one-off recovery, include `release_level_holds: true` on this +PUT; Hive removes only its own level-applied `hold` labels whose latest hold +event was by the App and comments on each PR with the operator and target level. What happens when the level changes (`handlePackSetLevel` → `ApplyPack`): @@ -214,19 +222,27 @@ What happens when the level changes (`handlePackSetLevel` → `ApplyPack`): roster** — adding every agent the level introduces (for example architect/strategist at higher levels) and applying each agent's mode for that level (advisory → measured → holdgated → full). +4. If the change crosses the L6 self-merge boundary, Hive restarts the request + relay generation so the self-authored auto-merge sweep starts or stops under + the new ACMM verdict without waiting for a pod restart or GitHub App re-save. + Existing level-applied holds remain held unless this `PUT /api/packs/level` + request explicitly included `release_level_holds: true`. When the sweep just + became eligible, the response includes `self_merge_sweep_active`, + `level_holds_pending`, `repos`, and `level_changed_at` so the dashboard can + inform the owner what remains held and which repos participate in L6 + auto-merge. Notes: - **Promotion adds agents and capability; demotion narrows it.** Moving up to L6 makes agents auto-merge on green CI; moving down returns them to holdgated or advisory. The per-level capability grid is the table at the top of this page. - Promotion also **releases the level holds the hive itself applied** to open App - PRs that the new level no longer requires, so you do not have to clean them up - by hand after a level bump. Release is fail-closed: it applies only to - App-authored PRs carrying the hive's own attributable level-hold notice, only - when the most recent `hold` label event was applied by the App, and never while - a self-authorization hold applies. A hold a human applied — or re-applied after - the hive removed one — is never touched. + Promotion does **not** release level holds the hive itself applied; a human + removes `hold`, or an operator makes the one-off API call above. The one-off + release is fail-closed: it applies only to PRs carrying the hive's own + attributable level-hold notice and only when the most recent `hold` label + event was applied by the App. A hold a human applied — or re-applied after the + hive removed one — is never touched. - **Operator-created agents are preserved.** `ApplyPack` reconciles pack agents; agents you created yourself are not removed by a level change (deletion is tombstoned separately — see agent configuration). @@ -271,3 +287,13 @@ entry plus a visible decision bead. The hive-wide `acmm_level` remains the ceiling. Until the per-repo ACMM RFC (#6111) lands, Hive keeps this repo-keyed seam and teaches the live proxy to apply the repo override on matching repository requests so enforcement observes the decision without a restart. + +For the narrow L6-except-one-repo case, `project.repo_policies[].auto_merge: +false` disables only auto-merge on that repository. Below L6, auto-merge is +effectively off for every repo regardless of stored values. The hive-wide ACMM +level still determines issue/PR creation authority, but merge authority is removed at +the `hive-merge` relay, App self-authored sweep, and proxy direct-merge seam. +The dashboard toggle (`POST /api/repos/auto-merge`) is gated asymmetrically: +switching auto-merge *off* needs the same tier as pausing the repo (verified +owner or GitHub repo write), but switching it back *on* restores Hive's merge +authority and requires a verified owner (#9070). diff --git a/docs/content/hive/adr/0005-forge-abstraction.md b/docs/content/hive/adr/0005-forge-abstraction.md index 351632a..131a437 100644 --- a/docs/content/hive/adr/0005-forge-abstraction.md +++ b/docs/content/hive/adr/0005-forge-abstraction.md @@ -30,7 +30,7 @@ styles are already neutral. Most Hive code can depend on the work-item model instead of GitHub-specific API types, which makes GitLab and Forgejo support additive rather than a fork of the -scheduler. The hold label gives the system one cross-forge merge gate primitive. +scheduler. The canonical `hive-pause/` label gives the dashboard one exact-match manual hold primitive; level-gated GitHub PRs use literal `hold`, and `hive/` is provenance only. The trade-off is an intentionally incomplete abstraction: callers that need a real merge still use a forge-specific client until Hive has enough semantics to standardize that operation honestly. diff --git a/docs/content/hive/adr/0010-escalation-circuit-breaker.md b/docs/content/hive/adr/0010-escalation-circuit-breaker.md index f872c95..fb12d5f 100644 --- a/docs/content/hive/adr/0010-escalation-circuit-breaker.md +++ b/docs/content/hive/adr/0010-escalation-circuit-breaker.md @@ -23,8 +23,27 @@ and raw failure evidence when available, and applies the `needs-human` label so future fix dispatch skips the PR. For unchanged red heads, track staleness separately and cap re-engagements at -three per current SHA. A branch that moves resets the re-engagement counter; a -permanently red, never-moving branch is not nudged forever. +`MaxReEngagements` (six) per current SHA. A branch that moves resets the +re-engagement counter; a permanently red, never-moving branch is not nudged +forever. + +A re-engagement is a DELIVERED kick, not a counter increment: the reaper and +the merge watcher resolve the PR's owning agent (falling back to +`review.fixer_agent`, then `scanner`, when the owner is paused or unreachable), +send it a targeted FIX-BEFORE-NEW kick for that one PR, and charge the budget +only once the kick is accepted. Spacing is at least the staleness window and at +least the owner's slowest cadence, so six attempts cannot be spent inside one +cadence window. The staleness clock starts when CI SETTLES — a head red on one +check while others are still running is not yet stuck. The escalation comment +quotes the number of kicks actually delivered. + +Shared CI breakage is not a fix attempt. A failing check red on at least three +other open PRs in the same pass is an incident (a broken base branch, a runner +outage), so the pass is treated as no-information: no attempt counted, no +staleness clock, no re-engagement. Attempts are counted per head TREE, so an +empty `ci: retrigger` commit does not consume one. An escalation is un-parked +automatically — `needs-human` removed, ledger reset, one comment saying why — +when the head goes green or when the shared breakage it escalated over clears. A PR escalating for the SECOND time — after the reviewer lane ([#5480]) already repaired or de-escalated it once — gets a structured hand-off note instead of diff --git a/docs/content/hive/adr/readme.md b/docs/content/hive/adr/readme.md index 040c104..a461e54 100644 --- a/docs/content/hive/adr/readme.md +++ b/docs/content/hive/adr/readme.md @@ -56,3 +56,4 @@ What becomes easier, harder, safer, or riskier because of this decision? - [ADR-0016: Scope `script-src` as two directives, close the element half with hashes](/docs/hive/adr/0016-csp-script-src-scope) - [ADR-0017: Quadlet `.container`/`.pod` units as the Podman persistent lifecycle](/docs/hive/adr/0017-podman-quadlet-lifecycle) - [ADR-0018: Shared dashboard design tokens and component layer](https://github.com/hivecommons/hive/blob/v5/src/docs/adr/0018-dashboard-design-tokens.md) +- [ADR-0019: Escalate to direction, spec, signal, or meta-issue instead of stalling](https://github.com/hivecommons/hive/blob/v5/src/docs/adr/0019-escalation-over-stalling.md) diff --git a/docs/content/hive/agent-configuration.md b/docs/content/hive/agent-configuration.md index 614be57..b251554 100644 --- a/docs/content/hive/agent-configuration.md +++ b/docs/content/hive/agent-configuration.md @@ -6,6 +6,8 @@ A hive **agent** is a long-running AI worker — a CLI session the hive keeps al Start with only a name, a method, and a model. Add the rest when the agent needs it. +For labels that route or gate agent work, see [Hive Labels and Control Signals](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md). + ## The smallest agent that works ```yaml @@ -71,7 +73,7 @@ Every field below exists in the config schema today. Grouped by what it does: agents: scanner: display_name: scanner # dashboard label (defaults to the YAML key) - description: "Triages issues and opens hold-gated fix PRs." + description: "Triages issues and opens PRs gated by `hold`." emoji: "🔍" # dashboard badge color: "#3498db" # dashboard accent color role: scanner # behavioral role; defaults to the agent name @@ -203,9 +205,21 @@ spec omits `launch_cmd`, Hive still builds the backend command normally. # discouraged. See docs/per-repo-agents.md. on_demand: false # true = never kicked by the governor timer; # only triggered explicitly (e.g. inception) + continuous: false # legacy shorthand: continuous in every + # non-quiet governor mode; prefer per-mode + # cadence: surge: continuous + continuous_cooldown: 60s # optional cool-down before that re-kick; + # default 60s, exponential backoff on + # undeliverable kicks + continuous_budget_pct: 80 # optional token-budget guard; default 80% + # of governor.budget.total_tokens clear_on_kick: true # default true; false keeps session context across kicks stale_timeout: 28800 # seconds of silence before the agent counts as stale — # must exceed its longest cadence + busy_no_activity_threshold: 30m + # surface Working/no-transcript activity after repeated + # undeliverable kicks; default 30m + max_turn_duration: 0s # optional visibility-only Working ceiling; 0 disables restart_strategy: immediate # how to bring a dead session back beads_dir: /data/beads/scanner # work-record (bead) storage; default per-agent replicas: 3 # materialize scanner, scanner-2, scanner-3 (max 5) @@ -307,10 +321,11 @@ Rounding out the schema — fields you will rarely touch: |---|---|---| | `id` | Stable identifier | agent name | | `acmm_levels` | ACMM levels this agent participates in | all | -| `caveman_mode` | Prompt-compression experiment: `lite`, `full`, `ultra`, `wenyan`; see below | off | +| `caveman_mode` | Prompt-compression experiment: `lite`, `full`, `ultra`, `wenyan`; see below | empty (disabled) | +| `jev_mode` | Give the agent the Jev typed-decision tool: `off`, `assist`; see below | empty (off — nothing installed, no Jev calls) | | `explain_mode` | Ask the agent to report why it made each tool call: `off`, `brief`, `full`; see below | inherit the hive default | | `metrics_collector` | Named metrics source for the stats panel | none | -| `stats_display` | Custom sidebar metrics (key, label, source, field, style). The `health` source (the primary repo's CI/coverage/release checks) is offered only to agents that can own CI — never to an `ADVISORY` or `on_demand` agent — and a `pct`/`pct-bar` stat with no measurement renders `—`, not `0%`. | none (an agent starts with no stats unless it is a built-in with defaults) | +| `stats_display` | Custom sidebar metrics (key, label, source, field, style, optional icon/target). The Diagnostics **Quality stats** card mirrors the `quality` agent's configured entries. A `pct`/`pct-bar` stat with no measurement renders `—`, not `0%`. | none (an agent starts with no stats unless it is a built-in with defaults) | | `hidden` (packs only) | Keep a pack agent out of the default roster view | false | ## Explain mode (debugging agent behaviour) @@ -391,7 +406,7 @@ Leave it off outside of debugging: the explanation is extra output tokens on eve | Mode | Dashboard description | When to use | | --- | --- | --- | | `lite` | Removes filler while preserving normal language. | Lowest-risk token reduction for routine agents. | -| `full` | Converts output toward terse "caveman-speak". | Default example mode when cost matters and operators accept rougher prose. | +| `full` | Converts output toward terse "caveman-speak". | Use when cost matters and operators accept rougher prose. | | `ultra` | Telegraphic compression. | High-volume lanes where compact summaries are more important than nuance. | | `wenyan` | Classical Chinese-style compression. | Specialized/experimental mode; use only when readers and downstream tools can tolerate it. | @@ -403,6 +418,75 @@ Implementation notes: - Unsupported backends log that caveman is not supported and continue without compression. - The UI describes the feature as roughly 65% output reduction, but exact savings vary by prompt, backend, and task. +## Jev typed decisions (`jev_mode`) + +`jev_mode: assist` gives an agent **Jev** (TypeSafe AI's confidence-scored +typed-decision model) as a tool, so quick yes/no, pick-one and rubric-score +judgments — duplicate checks, "which of these N issues is relevant", "does +this diff touch secrets", rubric scoring before posting a review, gating +low-value actions — are answered by a cheap model billed on input tokens only +instead of frontier-model reasoning. It is per agent and off by default, +mirroring `caveman_mode`. + +```yaml +agents: + scanner: + jev_mode: assist # off (default) | assist + +jev: # hive-wide client; all optional + provider: openrouter # openrouter (default) | typesafe + model: typesafe/jev-1.13 # default per provider (typesafe: jev-latest) + endpoint: "" # override the provider's systemone URL + api_key_env: JEV_API_KEY # falls back to the connected OpenRouter gateway key + timeout: 5s +``` + +Toggle it per agent from `hive.yaml`, the agent's General settings panel +("Jev Mode", next to Caveman Mode), or `hivectl agent jev-mode-set +assist`. The dashboard select stays disabled with a hint until the hive can +resolve a Jev key (`JEV_API_KEY`, or an OpenRouter gateway connected under +Settings → Governor → Model Gateways). +Because the skill and env are applied at launch, saving a change of +`jev_mode` from the dashboard or `hivectl` restarts the agent (as a model or +backend change does); editing `hive.yaml` by hand takes effect on the agent's +next start. + +What turning it on does: + +- **Skill.** Before the CLI starts, the manager writes the embedded + `jev-decide` `SKILL.md` into the agent's CLI home (`~/.claude/skills`, + `$CODEX_HOME/skills`, `~/.copilot/skills`, `~/.gemini/skills`, + `~/.config/goose/skills`), as the agent user. The skill tells the agent + when Jev is the right tool and that it must never be used to generate code + or prose. Backends without a known skills directory still get the CLI and + env below. +- **Tool.** `hive jev decide --type choice|score|probability --question … + [--option name[=description]]… [--level …]… [--state json]` prints + `{"answer","confidence","probabilities","model","input_tokens"}`. `choice` + picks one of 2–32 options; `score` returns a position along 2–10 ordered + rubric levels; `probability` returns P(yes) for a statement (TypeSafe's + *Noul*; confidence is derived as |2·P − 1| because the provider reports + none for it). +- **Proxying.** The CLI only ever talks to the hive's loopback decision + endpoint (`127.0.0.1:18446`). The hive identifies the caller from the + socket UID **only** (the unforgeable half of the egress proxy's check — the + self-asserted `Proxy-Authorization` fallback the proxy allows under + `HIVE_PROXY_ADVISORY_OK` is deliberately not honoured here), refuses any + agent whose live `jev_mode` is not `assist`, attaches the Jev key, and + forwards to the provider. Agents never see the key. Consequence: Jev + requires per-agent UID isolation (the entrypoint's `uid-map.json`); on a + shared-UID or advisory-only deployment every call is refused as + unidentified. +- **Budget and audit.** Input tokens are recorded against the agent through + the same inference token sink the governor budget reads, and every call is + written to the audit log as `jev_decision` with `question_type`, + `confidence`, `input_tokens` and `model` — never the question or state. +- **Env.** `HIVE_JEV_MODE=assist` and `HIVE_JEV_ENDPOINT` are exported to the + agent only when the mode is on. + +Config validation accepts only `off`, `assist`, or empty; the dashboard and +`hivectl` writes apply the same gate. + ## Methods: subscription CLIs vs self-hosted inference `backend:` picks one of two families. They differ in how you authenticate and where model lists come from: @@ -425,7 +509,7 @@ Two rules of thumb: Only `POST /v1/messages` is translated into an OpenAI `/v1/chat/completions` call. The Claude CLI also talks to its Anthropic host for housekeeping — telemetry batches (`/api/event_logging/...`), error reports, `POST /v1/messages/count_tokens` — and none of that has a meaning to an OpenAI-compatible gateway; forwarding it used to cost a gateway `400 Missing required parameter: 'messages'` per call, charged against the provider's request rate limit (roughly two failures per real completion in practice). The translator and the MITM reroute now answer those locally: `count_tokens` returns a chars-based estimate, anything under `/api/` returns `{}`, and any other path is a 404 in Anthropic error shape with a `WARN` log line naming the method and path, so a new CLI endpoint shows up in the hive log rather than as gateway noise. Inference-routed `claude` sessions are additionally launched with `DISABLE_TELEMETRY=1`, `DISABLE_ERROR_REPORTING=1`, and `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1`; subscription sessions are not. -Every Model Gateway (and the bob backend) also accepts an optional `key_name` — a human-chosen LABEL for the configured key, e.g. `key_name: openrouter-prod-key`. It is safe-to-show metadata, not a secret: the dashboard's gateway row displays it as "Using key: ``", or "(unnamed)" when no label is set, so operators can tell keys apart without ever seeing the value. See [`inference-backends.md`](https://github.com/hivecommons/hive/blob/v5/docs/inference-backends.md) for a full YAML example. +Every Model Gateway (and the bob backend) also accepts an optional `key_name` — a human-chosen LABEL for the configured key, e.g. `key_name: openrouter-prod-key`. It is safe-to-show metadata, not a secret: the dashboard's gateway row displays it as "Using key: ``", or "(unnamed)" when no label is set, so operators can tell keys apart without ever seeing the value. For gateway keys, the row and settings/status APIs also expose `keySHA256`, the lowercase SHA-256 of the exact effective key string Hive will inject through the inference proxy; this lets operators compare Hive with LiteLLM/OpenRouter key-hash displays without revealing the secret. Saving a replacement key refreshes live inference proxy routes for running agents, so the next kick uses the new key without an agent restart. See [`inference-backends.md`](https://github.com/hivecommons/hive/blob/v5/docs/inference-backends.md) for a full YAML example. Kubernetes manifests for deploying inference backends (vllm Deployment, EPP RBAC, kustomization) are in [`deploy/inference/`](https://github.com/hivecommons/hive/blob/v5/src/deploy/inference/). @@ -506,6 +590,7 @@ governor: - **Per-agent, per-mode intervals.** Anything Go's duration parser accepts works (`5m`, `2h`) for interval mode. - **Pausing.** The value `pause` (or `paused`) suspends an agent for that mode without disabling it. - **On-demand agents** (`on_demand: true`) are skipped by the governor timer entirely — they run only when explicitly triggered (the inception workflow drives `brainstorm` this way). +- **Continuous agents** are configured per governor mode by setting that mode's cadence value to the literal `continuous` (case-insensitive), for example `agents.scanner.cadence.surge: continuous` with `busy: 1h`, `quiet: 6h`, and `idle: off`. No agent name (including `scanner`) has special behavior. The legacy `continuous: true` bool remains a shorthand for continuous in every non-QUIET mode, but new configs should prefer per-mode cadence values. They still need an active cadence entry in the current mode, but cadence no longer decides the next kick while continuous mode is on. Instead, the manager waits until the kicked CLI is genuinely back at its input prompt, records that turn end, and the governor schedules the next kick at `ended_at + continuous_cooldown` (default `60s`). Enabling continuous on an idle agent, including by entering a mode whose cadence is `continuous`, schedules its first continuous kick after one cooldown; leaving a non-continuous mode clears any pending continuous kick and `/api/status` reports `continuousBlocked: "not_in_mode"` for agents that are continuous only in other modes. Continuous never interrupts a busy turn and still respects `enabled: false`, operator/fleet-breaker pause, governor-mode `pause`/`off` (including QUIET-mode pauses), on-demand, non-kick channels, budget/provider holds, provider rate-limit/quota backoff, and upgrade/restart holds; failed deliveries such as "CLI did not reach input prompt" use capped exponential backoff. `continuous_budget_pct` (default `80`) stops only the continuous re-kick loop when the current token-budget window reaches that percentage of `governor.budget.total_tokens`; normal cadence resumes, and `/api/status` reports `continuousBlocked: "budget"` plus per-agent `continuousModes`, `continuousKicks`, and `continuousTokens` counters. The dashboard exposes per-mode `Continuous` in the agent settings **Cadences** form; the Governor cadence table is read-only and shows an `∞ continuous` chip in each continuous mode cell. - **Budget.** When the weekly token budget is exhausted, kicks are suppressed hive-wide (exempt agents excepted) until the period rolls over. Set `stale_timeout` with your cadences in mind: an agent kicked every 4h with a 30-minute stale timeout will look dead between kicks. The shipped packs use "longest cadence × 2". @@ -558,10 +643,10 @@ You don't have to design a roster. Hive ships six **ACMM packs** (`level-1.yaml` |---|---|---| | L1 | Inception (Assisted) | inception: brainstorm + guide, everything conversational | | L2 | Advisory (Instructed) | advisory beads only; agents observe, humans act | -| L3 | Quality-Gated (Measured) | quality opens issues and hold-gated test PRs; the rest stay advisory | -| L4 | Security-Aware (Adaptive) | all agents open issues — no PRs yet | -| L5 | Semi-Autonomous (Semi-Automated) | issues **and** hold-gated PRs; humans batch-approve | -| L6 | Fully Autonomous | auto-merge on green CI, no hold label | +| L3 | Quality-Gated (Measured) | quality opens issues and test PRs gated by literal `hold`; the rest stay advisory | +| L4 | Security-Aware (Adaptive) | scanner/guide file issues; quality, ci-maintainer, and sec-check can open PRs gated by literal `hold` | +| L5 | Semi-Autonomous (Semi-Automated) | issues **and** PRs; PRs are gated by literal `hold`; humans batch-approve | +| L6 | Fully Autonomous | auto-merge on green CI; non-outreach PRs have no level hold, outreach PRs remain held | Applying a level **reconciles the whole roster**, not just the diff: missing agents are created (as overlay files in `/data/agent-configs/`), existing agents are merged — pack values fill blanks, but your explicit `backend:`, `model:`, and `enabled: false` always win — and the level's `kick_template`, `mode` and `on_demand` are updated so the agent's *policy* matches the level (an `on_demand` you toggled yourself in the agent's settings dialog is operator-owned and left alone; leaving on-demand starts the agent, entering it stops it). A failed agent doesn't abort the rest; the level is only recorded as cleanly applied when every agent reconciled. @@ -686,7 +771,7 @@ Both call sites build the same `defsrc.Resolver` (`main.go:1356`), gated by `fun Two merge rules to know before you rely on this: - **A blank field never clears a baked value.** For most fields, an empty string or empty slice in the fetched definition is skipped, so a minimal definition can't silently wipe presentation you set elsewhere. `ClearOnKick` and `IncludeRepos` are the deliberate exceptions — their zero value (`false`) is a legitimate setting, so the definition's value is taken as authoritative whenever the source resolves live (`defsrc.go:212-216`). -- **Everything else on the agent is preserved untouched**, explicitly including: `Enabled`/`Paused`/`Managed` (operator lifecycle state), `ID`, `BeadsDir`, `MetricsCollector`, `ACMMLevels`, `OnDemand`, `CavemanMode`, and — critically — the `definition_source`/`prompt_source` pointers themselves. A live definition cannot re-point the agent at a different repo (`ApplyToConfig` re-asserts this at `defsrc.go:437-440` even though the merge already excludes it). Nothing under the hive-level `variables.security` block is reachable either — it isn't part of `AgentConfig` at all. +- **Everything else on the agent is preserved untouched**, explicitly including: `Enabled`/`Paused`/`Managed` (operator lifecycle state), `ID`, `BeadsDir`, `MetricsCollector`, `ACMMLevels`, `OnDemand`, `CavemanMode`, `JevMode`, and — critically — the `definition_source`/`prompt_source` pointers themselves. A live definition cannot re-point the agent at a different repo (`ApplyToConfig` re-asserts this at `defsrc.go:437-440` even though the merge already excludes it). Nothing under the hive-level `variables.security` block is reachable either — it isn't part of `AgentConfig` at all. ### The trust boundary: allowlisted repos are seed-only @@ -749,13 +834,15 @@ One agent reaches its template by **role** rather than by `kick_template`: an ag So the worst an override can do is change the *wording* of a kick that was already going to be sent. The template-specific variables are `${REVIEWER_WORK_LIST}`, `${REVIEWER_MAX_PRS}`, `${REVIEWER_PASSED_LABEL}`, `${REVIEWER_RECOMMEND_CLOSE_LABEL}` and `${REVIEWER_CLOSE_AUTHORITY}`, alongside the usual built-ins. -Do not confuse it with `reviewer-queue.md`, which belongs to the pack-defined `reviewer` agent in the L5/L6 packs: that agent wakes on a 30-minute cadence, works the open PR queue in advisory mode and routes `requires_human` / `reject` verdicts to a maintainer via the triage label itself. +Do not confuse it with `reviewer-queue.md`, which belongs to the pack-defined `reviewer` agent in the L5/L6 packs: that agent wakes on a 30-minute cadence, works the open PR queue in advisory mode and routes `requires_human` / `reject` verdicts to a maintainer via the triage label itself. Because a `kick_template` shadows the role-based routing, that agent never reaches this lane; the L5/L6 packs therefore ship a separate template-less `adjudicator` agent (`role: reviewer`, `mode: ISSUES_AND_PRS`, 30-minute cadence) to work escalated PRs ([#9477](https://github.com/hivecommons/hive/issues/9477)). Its work list covers every escalated hive PR the hub lists — the escalated rows of `ci_failing` plus the `escalated` list in `ci-failing.json`, which carries conflicted, green, and pending escalations that have no failing check. Portable agents bundle everything — config plus a `promptTemplate` — in a single `AgentDefinition` YAML you can import from a URL in the dashboard. The reference schema is [`../AGENT-DEFINITION.md`](https://github.com/hivecommons/hive/blob/v5/src/AGENT-DEFINITION.md), and a worked example lives at [`../examples/agents/customized-agent.yaml`](https://github.com/hivecommons/hive/blob/v5/src/examples/agents/customized-agent.yaml). ### Writing guide: how issues and PRs should read (`project.writing_guide`) -Every default template that files an issue or PR carries the variable `${WRITING_GUIDE}` immediately before the body template it tells the agent to fill in (`--body "## Finding …"`, `--body "## Test Improvement …"`). It expands to the text of `project.writing_guide`, wrapped in a short header that says who set it and that it governs how the body *reads*, not what the policy requires it to contain. It is **empty by default**, and an empty guide renders nothing — a hive that never sets it gets byte-identical prompts. +Every default template that files an issue or PR carries the variable `${WRITING_GUIDE}` immediately before the body template it tells the agent to fill in (`--body "## Finding …"`, `--body "## Test Improvement …"`). It expands to the text of `project.writing_guide`, wrapped in a short header that says who set it and that it governs how the body or review comment *reads* — wording, plainness, length, tone — not what the policy requires it to contain. The body template stays authoritative: every section, field and piece of evidence appears in the template's order, and the guide never drops, renames or reorders a section. If the guide asks for a high-level summary, the agent writes it at the top, before the template's first section, as an addition ([#9747](https://github.com/hivecommons/hive/issues/9747)). That opening paragraph and the title are written for a newcomer who does not know the codebase — what the thing is, what the problem or change is, why it matters to a user, in plain words and with no unexplained internal names, paths or jargon; the technical detail follows unchanged in the template's sections below ([#9926](https://github.com/hivecommons/hive/issues/9926)). It is **empty by default**, and an empty guide renders nothing — a hive that never sets it gets byte-identical prompts. + +Owners can edit the same value from the dashboard at **Settings → Labels → Writing guide**. The editor writes `project.writing_guide` through the same config-save path as the required-labels policy, so the change takes effect on the next kick, survives restart, and appears in the downloaded `hive.yaml`. ```yaml project: @@ -769,7 +856,7 @@ project: Why a setting and not `AGENTS.md` ([#7667](https://github.com/hivecommons/hive/issues/7667)): a style rule in a repo's `AGENTS.md` reaches the agent as background knowledge, lower in the prompt than the policy's own body template, and when the two disagree the agent follows the template. The variable puts the owner's rule *next to* the template, which is the only position that changed anything when tried. The alternative — editing each template in the prompt editor — saves a full copy of that policy to `/data/policies/` that then shadows every upstream update to it, per agent, for a style preference. -Where you will see it: the agent's Prompt Template tab renders the guide where the kick will place it, so you can confirm the setting took. Templates whose prompts are built in Go rather than from a policy file do not all carry the variable: the **review swarm** still does not. The **contributor relay's task prompt** does, since [#8124](https://github.com/hivecommons/hive/issues/8124) — it is built in Go, so it takes the rendered guide as a parameter rather than expanding `${WRITING_GUIDE}`, and places it immediately before the instruction that tells the agent to open the PR. The guide travels with the *assigning* hive, so a relay subscribed to two hives gets each hive's guide on that hive's tasks; see [`contributor-relay.md`](/docs/hive/contributor-relay#the-assigning-hives-writing-guide-travels-with-the-task). Review comments (`reviewer-queue.md`) are deliberately outside it: the guide is about issue and PR bodies. +Where you will see it: the agent's Prompt Template tab renders the guide where the kick will place it, so you can confirm the setting took. Prompts built in Go take the rendered guide as a parameter rather than expanding `${WRITING_GUIDE}` themselves. That includes the **contributor relay's task prompt**, since [#8124](https://github.com/hivecommons/hive/issues/8124), immediately before the instruction that tells the agent to open the PR. The guide travels with the *assigning* hive, so a relay subscribed to two hives gets each hive's guide on that hive's tasks; see [`contributor-relay.md`](/docs/hive/contributor-relay#the-assigning-hives-writing-guide-travels-with-the-task). The **review swarm** also carries it in review prompts, so posted review comments use the same editorial voice as issue and PR bodies. ## Label policy: which issues agents may work @@ -797,7 +884,29 @@ Semantics: - **Exempt wins on conflict**: an issue carrying both an exempt label and a required label stays excluded. There is deliberately no separate `exclude_labels` field — the exempt list *is* the exclusion mechanism, applied first. - PRs and the Hold list are unaffected: open PRs are in-flight work, and held issues still appear under On Hold. -Both polarities are enforced at **enumeration** — the point where GitHub issues become the hive's actionable set — not in the prompt. A filtered issue never enters the queue, never appears in a kick, never triggers plan-from-label, and cannot be re-selected by a confused (or prompt-injected) agent re-listing the repo. Kick prompts additionally state the active require policy so agents know the list is intentionally short. Both lists are edited on the Labels tab; an active require gate is also noted read-only under **Repositories**, and hub-managed hives can receive `issue_filter` with their project config over the heartbeat. +Both polarities are enforced at **enumeration** — the point where GitHub issues become the hive's actionable set — not in the prompt. A filtered issue never enters the queue, never appears in a kick, never triggers plan-from-label, and cannot be re-selected by a confused (or prompt-injected) agent re-listing the repo. Kick prompts additionally state the active require policy so agents know the list is intentionally short. Both lists are edited on the Labels tab; an active require gate is also noted read-only under **Projects**, and hub-managed hives can receive `issue_filter` with their project config over the heartbeat. + +### Reporter trust: who filed it, not only what it is labelled + +The two polarities above look only at labels, so a maintainer's issue and a first-time stranger's issue are admitted by the same rule — and at ACMM L6 a stranger's request can be worked and merged on green CI with nobody looking. The require polarity fixes that only by making *everyone*, maintainers included, hand-label their own issues first. **Reporter trust** ([#9665](https://github.com/hivecommons/hive/issues/9665)) splits the two: + +```yaml +project: + issue_filter: + reporter_trust: + enabled: true # off by default — existing hives change nothing + trusted_associations: [OWNER, MEMBER, COLLABORATOR] # the default; CONTRIBUTOR is deliberately not in it + trusted_logins: [external-maintainer] # trusted whatever GitHub says about them + untrusted_require_labels: [triage/accepted] # the default + awaiting_label: needs-triage # default; set "" to skip the visible wait label + comment: true # default; posts one marked wait explanation +github: + reporter_trust_hold: true # optional; nil follows reporter_trust.enabled +``` + +- **Admission.** GitHub reports an `author_association` on every issue. A reporter in `trusted_associations` (or whose login is in `trusted_logins`) has their issues admitted by the ordinary rules. Anyone else's issue is not actionable until a maintainer adds one of `untrusted_require_labels`. On the first excluded enumeration, the core `github.Client.fetchIssues` poller posts one `` comment explaining the gate and applies `awaiting_label` (default `needs-triage`) so the wait is visible; later scans do not repeat the comment. This is a poller-side write, not an agent prompt, and is capped at 10 new notices per poll. Once a maintainer adds the required label, the same poll loop removes the awaiting label only if that marker records that Hive added it, then admits the issue. This runs after the hold/exempt checks and **before** `require_labels`, so the ordinary allow-list still applies to trusted reporters afterwards — "everyone needs an approval label" stays expressible exactly as before. Hive- and bot-filed issues are not judged here; the [#5117 self-authorization gate](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md) owns those on the PR side. An unknown reporter (no login, or no association in the payload) is treated as untrusted. +- **Merge.** A PR whose rationale (its closing or referencing links, or the request's declared issue list) traces to an untrusted reporter's issue receives `hold` at **every** ACMM level, including L6, together with a marked notice explaining who asked. Any one untrusted citation holds. A human removes the label; Hive never auto-releases a reporter-trust hold — the level-hold release path recognises the notice and leaves the label, and holdguard re-holds on a new head SHA. `github.reporter_trust_hold` follows `reporter_trust.enabled` unless set explicitly; `project.repo_policies[].reporter_trust_hold` overrides per repo; `HIVE_REPORTER_TRUST_HOLD` locks it from the environment, exactly like the #5117 knobs. +- **Where you see it.** Both halves are edited in the dashboard: **Settings → Labels → Reporter trust** (the switch, the association checkboxes, extra logins, triage labels) and **Settings → Repos** (the hold, with the same all-repos default / per-repo override / inherit rows as #5117). An active gate is stated in the read-only note under **Projects**, and each repo card counts issues **needs triage** separately from generic filter refusals. The issue itself carries `needs-triage` while waiting when Hive had to add it. Holds carry the reason in the PR notice and in the `agent_pr_created` audit entry (`reporter_trust_held`, `reporter_login`, `reporter_association`). **Not the same thing as the contribute filters.** `hub.contribute_labels_mode` + its label list gate which issues are *handed out to external contributors* over `/contribute` — they have never gated the hive's **own** agents, so an operator who allow-listed a queue label there (a common setup for routing labeled issues to contributors) still had a hive whose own scanner could work every other open issue. `project.issue_filter.require_labels` is the agent-side gate; configure both if you want the same label to govern both lanes. diff --git a/docs/content/hive/architecture.md b/docs/content/hive/architecture.md index 3312a71..d10422b 100644 --- a/docs/content/hive/architecture.md +++ b/docs/content/hive/architecture.md @@ -267,6 +267,19 @@ Agent **identity** is derived from the connection's owning UID (`/proc/net/tcp` repo outside the configured set. Only `api.github.com` is inspected; `github.com` (OAuth, git smart-HTTP) is tunneled opaquely. +### Structural dependency ratchets + +Some guardrails are intentionally mechanical ratchets rather than immediate +decomposition work. `internal/testutil/sleep_ratchet_test.go` prevents fixed +test sleeps from regrowing, and +`internal/testutil/dashboard_import_ratchet_test.go` does the same for +`pkg/dashboard` coupling: it scans the top-level dashboard package's non-test Go +files, collapses internal imports to `github.com/hivecommons/hive/pkg/`, and +compares them to `internal/testutil/dashboard_import_allowlist.txt`. New +top-level `pkg/` dependencies must route through an existing seam or be added to +the allowlist with PR justification, while stale allowlist entries must be +removed so the dependency surface only shrinks. + --- ## 6. ACMM — controlling agent autonomy @@ -281,8 +294,8 @@ flowchart LR L2["L2 Advisory
observe + report"] L3["L3 Quality-Gated
quality opens PRs"] L4["L4 Security-Aware
more agents file/PR"] - L5["L5 Semi-Autonomous
all PRs, hold-gated"] - L6["L6 Fully Autonomous
auto-merge on green"] + L5["L5 Semi-Autonomous
all PRs, hold gated"] + L6["L6 Fully Autonomous
auto-merge switches on"] L1 --> L2 --> L3 --> L4 --> L5 --> L6 ``` @@ -290,10 +303,10 @@ flowchart LR |-------|------|------------------| | L1 | Inception (Assisted) | Advisory beads + project inception only | | L2 | Advisory (Instructed) | Observe and report findings as beads; no GitHub writes | -| L3 | Quality-Gated (Measured) | `quality` opens hold-gated PRs about testing gaps, coverage, and CI health; others advisory. This measurement foundation is what earns automation at higher levels | -| L4 | Security-Aware (Adaptive) | All agents file issues (bugs, docs, workflows, vulns); still no PRs | -| L5 | Semi-Autonomous (Semi-Automated) | All agents open PRs — every PR carries a `hold` label for human review | -| L6 | Fully Autonomous | Agents open PRs and **auto-merge on green CI**; no hold required | +| L3 | Quality-Gated (Measured) | `quality` opens PRs gated by literal `hold` about testing gaps, coverage, and CI health; others advisory. This measurement foundation is what earns automation at higher levels | +| L4 | Security-Aware (Adaptive) | All agents file issues; quality, ci-maintainer, and sec-check can open PRs gated by literal `hold` | +| L5 | Semi-Autonomous (Semi-Automated) | All agents open PRs — every PR carries literal `hold` for human review | +| L6 | Fully Autonomous | Agents open PRs and **auto-merge on green CI**; switching to L6 enables auto-merge for every active repo, owners may toggle repos afterward, non-outreach PRs have no level hold, outreach PRs remain held | `supervisor` is always advisory. The full matrix is in [`acmm-policy-matrix.md`](/docs/hive/acmm-policy-matrix). diff --git a/docs/content/hive/backup-dr.md b/docs/content/hive/backup-dr.md index 7912d78..cfcec31 100644 --- a/docs/content/hive/backup-dr.md +++ b/docs/content/hive/backup-dr.md @@ -6,6 +6,21 @@ Hive has two backup paths with different scopes: nightly encrypted hub disaster- > **See also:** [Hub disaster recovery](https://github.com/hivecommons/hive/blob/v4/docs/HUB_DISASTER_RECOVERY.md) — the full hub-level runbook (key escrow, spoke fleet recovery, Slack blast, rebuild from zero) that the `hive-backup` archives described here feed into. For moving a live hive to a **different** host or cluster (as opposed to backing it up in place), see [Moving a Hive between hosts, same runtime](https://github.com/hivecommons/hive/blob/v5/src/docs/move-host.md), [Self-hosted Kubernetes cluster move](https://github.com/hivecommons/hive/blob/v5/src/docs/move-kubernetes.md) and [Hub-registered hive cutover](https://github.com/hivecommons/hive/blob/v5/src/docs/move-hub-registered-cutover.md); [Cross-runtime moves](https://github.com/hivecommons/hive/blob/v5/src/docs/move-cross-runtime.md) covers Podman ↔ Docker cross-host and Compose/Quadlet ↔ Kubernetes, using the same-host Docker → Podman migration below as its reference case. +## Config export is not a backup + +The dashboard avatar menu has two separate owner actions: + +- **Export effective config (JSON)** calls `GET /api/config/export`. It is a + human-readable, deterministic snapshot for review, drift detection, and bug + reports. It includes the effective config plus layer and side files, with + secret-looking values redacted by default. +- **Back up this hive** calls `POST /api/backup`. It creates an encrypted + disaster-recovery archive that includes operational data such as beads and + requires the hive backup key to restore. + +Use the JSON export to answer "how is this hive configured right now?" Use the +encrypted backup to recover data after a failure. + ## Hub disaster recovery: `hive-backup` `src/cmd/hive-backup` creates encrypted hub disaster-recovery archives — everything needed to rebuild a hub. It captures: @@ -125,7 +140,7 @@ The archive has two top-level directories (`src/pkg/spokebackup/backup.go:111-11 | `beads//**` | `/data/beads//**` | `beadsSubdir`/`beadsPrefix`, `backup.go:58,114` — one subtree per agent, discovered from the archive rather than a fixed list | | `MANIFEST.json` | (not restored — it is metadata, verified by `Extract`, not spoke state) | `hubbackup` manifest format | -The mapping is a flat rename of the two archive prefixes (`spoke/` → data-dir root, `beads/` → data-dir `beads/`) — there is no repacking, renaming, or transformation needed. This is exactly the file set the entrypoint reads at boot: `HIVE_CONFIG_RUNTIME`/`HIVE_CONFIG_RUNTIME_LEGACY`/`hive.yaml.dashboard` (`src/deploy/entrypoint.sh:69-70,262-266,533-550`), the beads directory it symlinks into `/home/dev/-beads` and chowns per-agent (`entrypoint.sh:894-931`), and `gh-app-key*.pem`, read directly from `/data` (`src/pkg/dashboard/api.go:5374`, `src/pkg/hub/cluster_app_key.go:393`). +The mapping is a flat rename of the two archive prefixes (`spoke/` → data-dir root, `beads/` → data-dir `beads/`) — there is no repacking, renaming, or transformation needed. This is exactly the file set the entrypoint reads at boot: `HIVE_CONFIG_RUNTIME`/`HIVE_CONFIG_RUNTIME_LEGACY`/`hive.yaml.dashboard` (`src/deploy/entrypoint.sh:69-70,586,879`), the beads directory it symlinks into `/home/dev/-beads` and chowns per-agent (`entrypoint.sh:1239-1248,1847-1853`), and `gh-app-key*.pem`, read directly from `/data` (`src/pkg/dashboard/api_config.go:186`, `src/pkg/hub/cluster_app_key.go:393`). ### Decrypting the archive — executed @@ -221,7 +236,7 @@ Points worth knowing before you run it: 3. **Where "the target `/data`" is.** On Kubernetes it is the spoke's PVC — run `restore` from an init container or a temporary debug pod mounting the same PVC before the hive Deployment's pod starts (or into a scaled-down Deployment's pod). On Compose/Quadlet it is the `hive-data` volume — use the same container-mediated pattern as the [host-level backup pattern](#host-level-backup-pattern) / [Podman restore](#restore) sections above. The `hive-backup` binary ships inside the hive image (`src/Dockerfile`), so a debug pod or `podman exec` on the hive image already has it. 4. **A restore merges into the bead ledger rather than replacing it.** An agent present at the destination but absent from the archive keeps its beads; an agent in both gets the archive's copy. If you want a clean slate, empty `/beads` first. -5. **Modes, and what `restore` does not do.** Every restored root file — the config pair, `hive-id`, `hive-state.json` and the GitHub App keys — is written `0600`, so a hand-built archive cannot widen a credential to world-readable. Ownership is *not* set: the entrypoint re-applies `0600` to the config files (`entrypoint.sh:262-266`, `hive_harden_runtime_config`) and re-chowns `beads/` to each agent's runtime UID (`entrypoint.sh:894-931`) on every boot, so a restore run from a host shell does not need to reproduce container-internal UIDs. +5. **Modes, and what `restore` does not do.** Every restored root file — the config pair, `hive-id`, `hive-state.json` and the GitHub App keys — is written `0600`, so a hand-built archive cannot widen a credential to world-readable. Ownership is *not* set: the entrypoint re-applies `0600` to the config files (`entrypoint.sh:437-450`, `hive_harden_runtime_config`) and re-chowns `beads/` to each agent's runtime UID (`entrypoint.sh:1847-1853`) on every boot, so a restore run from a host shell does not need to reproduce container-internal UIDs. 6. **(Re)start the container**, then **verify** the same way the Podman section above does: `hive-id` and the GitHub App key SHA-256 must match the source hive; the bead ledger and dashboard config overlay must be present. `hive-state.json` round-trips, but on a live restore expect it to be overwritten at the next boot — the same caveat the Podman restore section above notes for the same file. diff --git a/docs/content/hive/contributor-relay.md b/docs/content/hive/contributor-relay.md index 40238b5..5cbc9b5 100644 --- a/docs/content/hive/contributor-relay.md +++ b/docs/content/hive/contributor-relay.md @@ -4,6 +4,8 @@ ClankeR lets a contributor lend their local AI CLI subscription to a hive. A contributor runs a small relay process on their machine; the hive assigns it real work — issues from the project's queue — and the contributor's agent executes each task locally with the CLI and model of their choice, reporting completion/PR metadata back over a WebSocket. +For labels that make work eligible or ineligible for contributors, see [Hive Labels and Control Signals](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md). For help getting to your first PR, [Join our Discord](https://hivecommons.dev/discord). + The relay turns a hive from a fixed set of resident agents into an elastic swarm: the admin curates *what* is offered (which repos, which labels, which models are acceptable), and contributors decide *how* it gets done (their CLI, their model, their compute, their tokens). The relay connects to `/api/contribute/ws`, receives one task at a time, runs the selected CLI in the contributor's environment, and reports the result back. ## How it fits together @@ -28,6 +30,11 @@ sequenceDiagram ## Basic setup +For a homelab service with accounts independent of the host, use the +[published-image Compose example](https://github.com/hivecommons/hive/blob/v5/src/examples/contributor-isolated/README.md). +It keeps GitHub, Codex, Hive configuration and work in a named volume and +requires no host GitHub or provider CLI installation. + From a checkout of this repository: ```bash @@ -46,8 +53,8 @@ Management). The dashboard shows those links on Onboarding and Operations, and the hub includes them in `auth_ok` so the relay prints them once when it connects. Use it for contributor docs, Discord/Slack invites, or a maintainer issue queue. Discord `https://discord.com/channels//` URLs only -work for people already in that server; use a `https://discord.gg/...` invite if -new contributors need to join. +work for people already in that server; use `https://hivecommons.dev/discord` if +new contributors need to join the Hive Commons Discord. `contribute-hive` starts the relay in one of two modes: @@ -116,7 +123,10 @@ Important environment variables: | --- | --- | --- | | `HIVE_HUB` | value from `contributor.env`, else public hub default | WebSocket hub(s) to subscribe to. Use comma-separated URLs for multi-hub mode. Direct Compose reads the registered value from the mounted config file. | | `HIVE_REGISTRATION_TOKEN` | value from `contributor.env` | Registration token(s), positional with `HIVE_HUB` when multiple hubs are listed. Required; run `just contribute-setup` first. | -| `AGENT_BACKEND` | `claude` | CLI/backend to run (`claude`, `copilot`, `goose`, `bob`, `codex`, `pi`, `aider`, `litellm`, `agy`, `opencode`, `kilo`, `muse`, `omp`, depending on image support and credentials). `omp` is interactive-only: Hive starts normal `omp --model ` in the prepared tmux cwd and passes no fabricated permission flags. It has no verified local confinement mechanism, so local mode refuses it without `HIVE_OMP_DANGEROUSLY_RUN_UNCONFINED=1`; container mode is the supported boundary. `agy` has the same confinement limit. `opencode`, `kilo`, and `muse` only run headless (`CONTRIBUTOR_MODE=headless`) — hive has no interactive-tmux wiring for them. | +| `HIVE_COMMONS_STRATEGY` | `ranked` | Multi-hive routing strategy for The Commons. `ranked` solicits the first subscribed hive and falls through only when it has no work; `spread` rotates with rank weights and periodic mixing; `neediest` scores each hub's `/api/contribute/status` by `actionable_items` with a small boost for idle contributor capacity (`total_registered - active_contributors`) when those fields are present. The relay evaluates this only when it is ready for another task, never while a lease is in flight. | +| `HIVE_COMMONS_SPREAD_MIX_EVERY` | `7` | In `spread`, every N completed/failed tasks the next hive is chosen randomly from the weighted rank cycle instead of by cursor. Set `0` to disable mixing. | +| `HIVE_COMMONS_NEEDIEST_REFRESH_MS` | `60000` | In `neediest`, how often to refresh each subscribed hub's `/api/contribute/status` cache. Set `0` to rely on auth-time/manual refreshes only. | +| `AGENT_BACKEND` | `claude` | CLI/backend to run (`claude`, `copilot`, `goose`, `bob`, `codex`, `pi`, `aider`, `litellm`, `agy`, `opencode`, `kilo`, `muse`, `omp`, `openhands`, depending on image support and credentials). `omp` is interactive-only: Hive starts normal `omp --model ` in the prepared tmux cwd and passes no fabricated permission flags. It has no verified local confinement mechanism, so local mode refuses it without `HIVE_OMP_DANGEROUSLY_RUN_UNCONFINED=1`; container mode is the supported boundary. `agy` has the same confinement limit. `opencode`, `kilo`, `muse`, and `openhands` only run headless (`CONTRIBUTOR_MODE=headless`) — hive has no interactive-tmux wiring for them. | | `AGENT_MODEL` | unset (backend default) | Optional model override passed to the contributor agent (e.g. `claude-sonnet-4-6`, `gpt-4o`, `gemini-2.5-pro`). Declared to the hive when the relay connects. | | `AGENT_REASONING_EFFORT` | unset | Reasoning effort override. Consumed by `codex` (`-c model_reasoning_effort`), by `agy` (`--effort low\|medium\|high`, required whenever a model is set, else agy ignores the model), by `muse` (`--reasoning-effort none\|minimal\|low\|medium\|high\|xhigh\|max\|ultra`, applied with or without a model; a value outside that set is dropped rather than passed, because muse exits 2 on it), and by `claude` (`--effort low\|medium\|high\|xhigh\|max`, applied with or without a model; a value outside that set is dropped the same way, and unset leaves Claude Code at its own default - [#8377](https://github.com/hivecommons/hive/issues/8377)). Ignored by other backends, including inference routes such as `litellm` that drive the claude binary. | | `HIVE_CONTRIBUTOR_KNOWLEDGE_EXPORT_MAX_FACTS` | `80` | Maximum entries in the `/api/knowledge/export` response when it is fetched with a contributor registration token. Owner/API-token exports remain full; contributor startup context receives a bounded summary with a truncation marker and should fetch more specific entries on demand with `hive knowledge` or the Hive MCP knowledge tool. | @@ -124,7 +134,9 @@ Important environment variables: | `CONTRIBUTOR_MODE` | `interactive` | `interactive` keeps a tmux/TTY session. `headless` is for one-shot/no-TTY task delivery. | | `HIVE_AGENT_SESSION` | `contributor` | tmux session name for interactive mode. | | `HIVE_SESSION` | backend name (`AGENT_BACKEND`) | Optional session label for running multiple relays under one GitHub account (see [Running multiple backends under one account](#running-multiple-backends-under-one-account)). Relays with distinct labels get independent session-scoped identities (`ContributorID#session`) on the hub, so their task leases, assignment cooldowns, failure streaks, and ownership fences do not collide. Auth, trust tier, model admission, and rate-limit accounting stay per-account. Sanitized on the hub: only `[A-Za-z0-9._-]` survive, capped at 32 bytes; a label that sanitizes to empty counts as unset. Set it to the **empty string** to opt out — the relay then declares no session and keeps the bare per-account identity (the historical single-session behavior). | -| `HIVE_CONTRIBUTOR_QUOTA_GUARD` | `ask` | Contributor-local subscription quota guard mode. `ask` and `pause` hold new work when a normalized quota window is at or below the configured reserve; `off` is the explicit launch-time opt-out for this relay session. Headless/no-response sessions wait safely rather than spending the last quota headroom. The guard can only act when a reading source is available — either `HIVE_CONTRIBUTOR_QUOTA_READING_FILE` / `HIVE_CONTRIBUTOR_QUOTA_READING_JSON`, or, for a **supported subscription backend** (`claude`, `pi`, `codex`, `agy`, `gemini`), the reading the Go rotation probers publish — into `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` when that is set, otherwise the per-install derived pool directory, so publishing is default-on (hivecommons/hive#6967, hivecommons/hive#6987). The publisher no longer requires provider rotation to be enabled: when `governor.rotation.enabled` is false (the default), the spoke runs a publish-only prober loop that publishes readings without rotating anything, so any host running `hive` gets a publisher. A host running **only** the relay (no local `hive` process) still has no publisher and stays on this guard's no-source behaviour. With none of those it has nothing to read, logs that it is not guarding anything, and admits work. | +| `HIVE_CONTRIBUTOR_TEAM_METADATA` | unset | Opt in to the lightweight team-structure metadata used by `/api/leaderboard/teams` and the Leaderboard tab's team leagues. When set to `1`, `true`, `yes`, or `on`, the relay sends only distro/OS family (`/etc/os-release` ID, NAME, VERSION_ID, ID_LIKE on Linux), kernel release (`uname -r`), and agent backend. It never sends hostnames, local usernames, IPs, or hardware IDs. | +| `HIVE_TEAM_OS_FAMILY` / `HIVE_TEAM_OS_ID` / `HIVE_TEAM_OS_NAME` / `HIVE_TEAM_OS_VERSION_ID` / `HIVE_TEAM_OS_ID_LIKE` / `HIVE_TEAM_KERNEL_RELEASE` / `HIVE_TEAM_AGENT_BACKEND` | auto-detected when team metadata is enabled | Optional overrides for the opt-in team declaration. Use these when a container should represent the host distro, a WSL relay should declare Windows/WSL, or an operator wants to join a cultural team without exposing exact release text. | +| `HIVE_CONTRIBUTOR_QUOTA_GUARD` | `ask` | Contributor-local subscription quota guard mode. `ask` and `pause` hold new work when a normalized quota window is at or below the configured reserve; `off` is the explicit launch-time opt-out for this relay session. Headless/no-response sessions wait safely rather than spending the last quota headroom. The guard can only act when a reading source is available — either `HIVE_CONTRIBUTOR_QUOTA_READING_FILE` / `HIVE_CONTRIBUTOR_QUOTA_READING_JSON`, or, for a **supported subscription backend** (`claude`, `pi`, `codex`, `agy`, `gemini`), the reading the Go rotation probers publish — into `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` when that is set, otherwise the per-install derived pool directory, so publishing is default-on (hivecommons/hive#6967, hivecommons/hive#6987). The publisher no longer requires provider rotation to be enabled: when `governor.rotation.enabled` is false (the default), the spoke runs a publish-only prober loop that publishes readings without rotating anything, so any host running `hive` gets a publisher. A host running **only** the relay (no local `hive` process) gets its publisher from the standalone contributor launch paths instead (`hive-quota-publisher`, hivecommons/hive#10299, below); with that opted out or unavailable it stays on this guard's no-source behaviour. With none of those it has nothing to read, logs that it is not guarding anything, and admits work. | | `HIVE_CONTRIBUTOR_QUOTA_MIN_REMAINING_PCT` | `20` | Default remaining-percentage reserve for every quota window, **including kinds this build does not recognize** — an exhausted unfamiliar window holds work rather than being skipped, because providers add window kinds over time (`weekly_scoped` was one). A *healthy* unrecognized window still admits. Valid range is `0`–`100`; invalid values fail relay startup with a message naming the bad variable. The reserve reduces the chance of consuming paid/extra usage but cannot guarantee a task will finish within a provider window. | | `HIVE_CONTRIBUTOR_QUOTA_SHORT_MIN_REMAINING_PCT` | unset | Optional short-window reserve override for normalized `session`/`five_hour` quota windows; inherits `HIVE_CONTRIBUTOR_QUOTA_MIN_REMAINING_PCT` when unset. | | `HIVE_CONTRIBUTOR_QUOTA_WEEKLY_MIN_REMAINING_PCT` | unset | Optional weekly reserve override for normalized `weekly` and `weekly_scoped` quota windows; inherits `HIVE_CONTRIBUTOR_QUOTA_MIN_REMAINING_PCT` when unset. | @@ -134,6 +146,7 @@ Important environment variables: | `HIVE_CONTRIBUTOR_QUOTA_RETRY_MS` | `60000` | How often a guarded relay re-reads the quota and re-advertises `ready` once every effective reserve is clear. Must be a positive integer of milliseconds. An explicit contributor pause outranks a recovered reading. | | `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` | derived per-install | Directory for the cross-process quota-pool store (hivecommons/hive#6953) **and** the automatically-published quota reading (hivecommons/hive#6967). When set, relays that resolve to the same pool key share one guard state through atomically-written files: a peer relay holding an in-flight task reserves the pool so a second relay on the same account cannot independently oversubscribe the reserve, and this is the directory the out-of-band `just contribute-quota …` controls write overrides into. It is also where the Go rotation probers publish a normalized reading (`.reading.json`, written temp-file-plus-rename so a reader never sees a torn file): a **supported** subscription backend then reads its guard reading from here without any hand-configuration. **Publishing is default-on (hivecommons/hive#6987):** when this is unset, the publisher and relay derive the same per-install directory (`$XDG_CONFIG_HOME` or the platform user-config dir, joined with `hive/contributor-quota`), so a default install of a supported backend gets a reading with no env var — setting this only overrides the location, or points a shared pool at a common path. With an **explicit** dir, a not-yet-landed or torn reading **holds** (the operator declared a publisher for the pool); with the **derived default** dir, a *present* reading is evaluated (a torn/`unknown` one still holds — fail-closed), a *missing* one holds when a fresh `.publisher.json` presence marker declares a live publisher for the pool (hivecommons/hive#6987 condition (b)), and only a missing reading with no fresh marker falls through to `unprovisioned`/admit, so a host with genuinely no publisher is never stranded holding forever. The cross-process store proper stays opt-in on an explicit dir. | | `HIVE_CONTRIBUTOR_QUOTA_POOL_ACCOUNT` | unset | Account component of the pool key. It is **hashed** into an opaque 16-hex key and never logged raw (privacy). Two relays with the same value share one pool; unset keys the pool off the backend name alone (same-backend relays on one host share, which can only under-subscribe the true account pool, never over-subscribe). | +| `HIVE_CONTRIBUTOR_QUOTA_PUBLISH` | on | Opt-out for the standalone contributor quota publisher (hivecommons/hive#10299). The contributor image entrypoint and `just contribute-hive local` start `hive-quota-publisher` beside the relay, so a standalone contributor feeds its own guard with no custom reading program and no local `hive` server. `0`, `false`, `off` or `no` launches without it; anything else (including unset) publishes. The publisher also stands down by itself with the guard `off`, with an explicitly configured `HIVE_CONTRIBUTOR_QUOTA_READING_FILE`/`_JSON` (it never competes with an operator-declared source), and on a backend the guard does not support — see "How the reading reaches the guard". | | `HIVE_CONTRIBUTOR_QUOTA_SESSION_ID` | `HIVE_SESSION` then `pid-` | Stable id for this relay session. A `disable-session` override targets it, and a session opt-out expires when the process exits (a new process gets a new id, so a stale opt-out never authorizes a later run). Never an account identifier. | | `HIVE_CODEX_SANDBOX_MODE` | probed (see note) | Codex `--sandbox` value. Left unset, hive resolves it at launch instead of hard-coding one: `workspace-write` everywhere it can work, and `danger-full-access` **only** inside the contributor container when that container blocks the unprivileged user namespace `workspace-write`'s bubblewrap needs (#6653). Setting this pins one value and skips the probe. | | `HIVE_CODEX_APPROVALS_REVIEWER` | `auto_review` | Codex reviewer for boundary requests. The default prevents Hive-delivered work from waiting on an interactive operator while retaining `workspace-write`; set `user` only for an intentionally attended contributor. Set it to the **empty string** to omit the `-c approvals_reviewer=` key entirely — the escape hatch if a Codex release rejects that config key at startup. Doing so keeps the sandbox posture; it is not the same as the dangerous bypass. | @@ -147,6 +160,7 @@ Important environment variables: | `HIVE_PI_DANGEROUSLY_RUN_UNCONFINED` | unset | **Required** for `just contribute-hive pi local` to launch at all. pi ships with no sandbox by default; directory confinement exists only via a third-party extension hive does not depend on. | | `HIVE_AIDER_DANGEROUSLY_RUN_UNCONFINED` | unset | **Required** for `just contribute-hive aider local` to launch at all. aider has no sandbox or OS isolation option of any kind. | | `HIVE_KILO_DANGEROUSLY_RUN_UNCONFINED` | unset | **Required** for `just contribute-hive kilo local` to launch at all. kilo's `--auto` is an unattended auto-approve flag, not a boundary; kilo has no verified sandbox, filesystem allowlist, or command deny-list hive can wire. | +| `HIVE_OPENHANDS_DANGEROUSLY_RUN_UNCONFINED` | unset | **Required** for `just contribute-hive openhands local` to launch at all. The bare `openhands` CLI runs the agent on the host with no sandbox; `--always-approve` is approval, not a boundary, and its Docker sandbox is reachable only through `openhands serve`, not the headless one-shot path. | | `HIVE_OMP_DANGEROUSLY_RUN_UNCONFINED` | unset | **Required** for `just contribute-hive omp local` to launch at all. OMP has no sandbox, filesystem allowlist, or command deny-list Hive can wire; local mode refuses to launch without this. | ### Where each backend reads its instructions @@ -168,6 +182,7 @@ mode fixed for Goose in [#2393](https://github.com/hivecommons/hive/issues/2393) | `agy` | `CLAUDE.md` | | `opencode` | `AGENTS.md`, `CLAUDE.md` | | `kilo` | `AGENTS.md`, `CLAUDE.md` | +| `openhands` | `AGENTS.md`, `CLAUDE.md` | | `muse` | `AGENTS.md`, `CLAUDE.md` | | `omp` | `AGENTS.md`, `CLAUDE.md` | | anything else | `CLAUDE.md` only — the `*` fallback | @@ -218,7 +233,7 @@ OS-enforced — Seatbelt/bubblewrap/ProcessContainer depending on platform), gated on the installed CLI actually supporting the flag; opencode gets a command-name deny-list via its own `permission.bash` config (a floor, not a filesystem boundary — opencode has no OS sandbox); goose, agy, bob, pi, aider, -kilo, and omp have no confinement mechanism this repo can wire at all, and local +kilo, omp, and openhands have no confinement mechanism this repo can wire at all, and local mode for them **refuses to launch** unless the operator sets that backend's own `HIVE__DANGEROUSLY_RUN_UNCONFINED=1`. See [sandbox-isolation.md](https://github.com/hivecommons/hive/blob/v5/src/docs/sandbox-isolation.md)'s per-backend confinement matrix @@ -274,6 +289,7 @@ The relay speaks to whatever backend you set up — pass it to `contribute-setup | `kilo` | Headless-only: `kilo run "" --auto`; set `KILO_AUTH_CONTENT` / `KILO_CONFIG_CONTENT` or `KILO_API_KEY` (optional `KILO_ORG_ID`). Hive forwards only those values and never mounts a Kilo home/config directory. `--auto` is approval, not a sandbox. | | `muse` | Muse Code (`curl -fsSL https://dev.meta.ai/install.sh | bash`). Headless-only: `muse exec ""` is its documented non-interactive sub-command. Auth is `META_API_KEY` (which muse says always takes priority) or `~/.config/muse/auth.json` written by `muse login` / `muse auth set --api-key-stdin`. Set `AGENT_MODEL` to a catalog id from `GET https://api.meta.ai/v1/models`, queried **from the machine that will run muse** — the catalog is caller-dependent (a workstation and an AWS container saw different model sets for the same key on 2026-09-08), and an id the caller cannot see fails at task time. **muse brings its own OS sandbox** (bubblewrap/seccomp on Linux, seatbelt on macOS), on by default — hive narrows it rather than refusing local launch. Installed in both images, pinned by version and per-arch SHA-256 from muse's release manifest. | | `omp` | Oh My Pi — **interactive-only**: Hive starts ordinary `omp --model ` in the prepared tmux cwd, with no fabricated permission flags; there is no headless wiring for it. Sign in once on the host (run `omp`, complete its provider setup, quit): container mode stages an allowlist of `~/.omp/agent` — only the selected provider's credential rows — into the container ([#7678](https://github.com/hivecommons/hive/issues/7678)), and the contributor image ships a pinned, checksummed `omp` ([#7661](https://github.com/hivecommons/hive/issues/7661)) so `just contribute-hive omp` runs without an `omp` on the host. No verified local confinement mechanism, so local mode refuses without `HIVE_OMP_DANGEROUSLY_RUN_UNCONFINED=1`; container mode is the supported boundary. Full setup notes: [docs/backend-setup.md](https://github.com/hivecommons/hive/blob/v5/docs/backend-setup.md) | +| `openhands` | OpenHands CLI (`uv tool install openhands --python 3.12`; PyPI `openhands`, MIT). Headless-only: `openhands --headless -t "" --always-approve --override-with-envs`. Auth is `LLM_API_KEY` (plus `LLM_BASE_URL` for an OpenAI-compatible gateway) or the `~/.openhands/settings.json` an interactive first run writes; `AGENT_MODEL=provider/model` (any litellm id) is forwarded as `LLM_MODEL` — the CLI has no `--model` flag and ignores `LLM_*` without `--override-with-envs`. No sandbox on this path, so local mode refuses without `HIVE_OPENHANDS_DANGEROUSLY_RUN_UNCONFINED=1`; not shipped in the contributor image (needs Python 3.12). Upstream has marked the CLI repo as no longer actively maintained in favour of Agent Canvas — T3 experimental. Full setup notes: [docs/backend-setup.md](https://github.com/hivecommons/hive/blob/v5/docs/backend-setup.md) | ## Running multiple backends under one account @@ -327,8 +343,24 @@ The guard evaluates a reading; something has to produce it. The Go rotation prob **Publishing is default-on for supported backends (hivecommons/hive#6987, satisfying #6967 criterion 1).** The pool directory no longer has to be hand-configured. When `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` is unset, both the Go publisher (`rotation.DefaultContributorPoolDir`) and the JS relay (`defaultContributorPoolDir`) derive the same per-install directory: `$XDG_CONFIG_HOME` (honoured explicitly first, exactly as the hivectl session cache does, because Go's `os.UserConfigDir` ignores it on darwin) or the platform user-config dir, joined with `hive/contributor-quota`. The two derivations must stay byte-for-byte identical or the publisher writes where the relay never reads — the same parity the shared pool-key vectors pin (`TestDefaultContributorPoolDir_XDGParity` on the Go side, `#6987 defaultContributorPoolDir matches Go under XDG_CONFIG_HOME` on the JS side). An explicit `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` still overrides the location. -Two safety properties are deliberate: +**A standalone contributor publishes its own readings (hivecommons/hive#10299).** The paragraphs above describe a host running `hive`. The standard standalone contributor launch paths run no `hive` process, so until #10299 they had no publisher at all: the guard reported that it was not guarding anything and admitted work unless the contributor maintained a custom reading program. They now start `hive-quota-publisher`, a small publish-only command that probes the contributor's **own** backend with the contributor's **own** credentials and writes the same pool-keyed `.reading.json` and `.publisher.json` the server-mode publisher writes — the same `pkg/rotation` provider readers, the same normalization, the same atomic write, the same bounded five-minute poll with provider-error backoff. No hub deployment, no local Hive server and no custom schema conversion are involved. + +| Launch path | What starts the publisher | +| --- | --- | +| `just contribute-hive ` (container, the default) | `bin/contributor-agent.sh`, the image entrypoint, next to the relay | +| The isolated Compose example (`src/compose-contributor.yaml`) | the same entrypoint — it runs the same image | +| `just contribute-hive local` | the recipe itself: an installed `hive-quota-publisher` if there is one, otherwise a build from this checkout when a Go toolchain is present | + +The publisher **starts and stops with the contributor**, so a stopped contributor leaves no process refreshing its pool's presence marker. It stands down, says why on stderr and exits 0 — never blocking the launch — in four cases: + +- `HIVE_CONTRIBUTOR_QUOTA_PUBLISH` is `0`/`false`/`off`/`no`: the explicit opt-out. Anything else (including unset) publishes. +- `HIVE_CONTRIBUTOR_QUOTA_GUARD=off`: nothing would read the reading, so nothing is probed. +- `HIVE_CONTRIBUTOR_QUOTA_READING_FILE` or `_JSON` is set: the operator has declared where readings come from, and a second publisher would compete with it for the same pool. +- The backend is not one of the guard-supported subscription backends (`claude`, `pi`, `codex`, `agy`, `gemini`, `kiro`): it is named as unsupported at startup rather than left waiting for a reading that can never be produced. + +In local mode a host with neither an installed binary nor a Go toolchain keeps the previous no-publisher behaviour and prints a note pointing here; nothing about the guard's own semantics changes in that case. +Two safety properties are deliberate: - **A failed probe never publishes a healthy reading.** A probe error is published as `state: unknown`, which the guard *holds* on — it is never turned into "plenty of headroom". Publishing an admit off a measurement that failed would be exactly the fail-open this guard exists to prevent. One deliberate exception in the publish-only manager (below): a probe that failed because the CLI is **not installed** publishes *nothing* — nothing is written, so the relay evaluates exactly what it would with no publisher, the unprovisioned admit. That is not a fail-open: an absent CLI cannot spend quota on that host, while a published `unknown` would hold a pool the host can never provision. Under operator-configured rotation, `not_installed` keeps publishing as `unknown` — there the provider was named explicitly, so it is real signal. - **The `unprovisioned` admit flips to a hold only on POSITIVE evidence of a publisher (hivecommons/hive#6987, condition (b)), and the remaining divergence is recorded on purpose.** #6987's title paired "make publishing default-on" with "flip `unprovisioned` to a hold". The flip is gated on the guard being able to distinguish "a publisher is expected here but has not written yet" from "nothing will ever write here" — flipping blind would strand every host with no route, the exact fleet-wide stop #6951's ruling avoided. Both previously-recorded conditions are now implemented. **Condition (a)**: the publisher no longer depends on `governor.rotation.enabled` — when rotation is disabled, `src/cmd/hive/main.go` starts a publish-only manager (`rotation.NewContributorReadingPublisher`) that probes the guard-supported providers and publishes readings without enabling failover. **Condition (b)**: the publisher writes a per-pool **presence marker** (`.publisher.json`, `rotation.ContributorPublisherMarkerPath`) when its probe loop starts and refreshes it atomically with every published reading; the relay judges it purely by mtime (a torn or unparsable marker can never strand a host) against the same TTL as reading staleness. A probe that fails because the CLI is **not installed** publishes nothing and *removes* the marker — an absent CLI cannot spend quota, and either a published `unknown` or a leftover marker would strand that pool on a hold. Under operator-configured rotation, `not_installed` keeps publishing `unknown` and keeps its marker — there the provider was named explicitly, so it is real signal. So the two dirs now behave like this: - With an **explicit** `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR`, the operator has declared a publisher feeds this pool, so "no reading yet" (or a torn file) is `unknown` and **holds** while the publisher catches up — the opt-in flip #6967 already shipped for the configured case. @@ -430,6 +462,7 @@ Positional lists have no names, and one hand-edit that drops a field transposes ```bash hivectl hives list # which hives, and which one is active hivectl hives add hive-b --hub wss://hive-b.example.com/contribute +hivectl hives reissue hive-b # rotate only hive-b's saved token hivectl hives use hive-b # make it the hub the relay starts on hivectl hives rename hive-b staging hivectl hives remove staging # asks you to type the name @@ -445,7 +478,7 @@ The same list is available in the terminal UI: `just contribute-tui` (or `hivect From a fresh checkout, `just contribute-tui` and `just contribute-hives` share `bin/hivectl-bootstrap.sh`: they first honor `HIVECTL`, then `./bin/hivectl`, user/system installs, and `PATH`; when none is present, they extract `hivectl` from the Hive image into `./bin/hivectl`. The checkout records the image digest beside the binary and refreshes when the digest changes, refusing to run an unverifiable stale copy if Podman cannot inspect or pull the image. -See [hivectl.md](/docs/hive/hivectl#hives--named-profiles-for-the-hives-you-contribute-to) for the full command reference, including adding a hive whose token you already hold (`--token-stdin`). +See [hivectl.md](/docs/hive/hivectl#hives--named-profiles-for-the-hives-you-contribute-to) for the full command reference, including adding a hive whose token you already hold (`--token-stdin`) and reissuing one saved hive token through the profile store (`hivectl hives reissue --hub `). ## Moving the relay to another machine @@ -469,7 +502,26 @@ The bundle is passphrase-encrypted and contains one profile: hub URL, contributo Use this when you want to switch back and forth, or to try the VM before committing to it. The cost is that the credential now exists in two places: remove the imported profile or the old profile once you no longer want both machines able to connect as the same contributor id. -### Option 2 — reissue the credential (`just contribute-move`) +### Option 2 — reissue the credential + +If this machine uses named profiles (`~/.config/hive/profiles.yml` exists), +reissue through `hivectl` so the profile store remains the source of truth: + +```bash +hivectl hives reissue acme --hub wss://hive.example.com/contribute +# if acme already exists locally with its hub saved: +hivectl hives reissue acme +``` + +`hivectl hives reissue` calls `POST /api/contribute/reissue-token` with your +GitHub token from `gh auth token`, replaces only that named profile's +registration token in `profiles.yml`, and then regenerates `contributor.env` +with the existing projection code. Other profile tokens are left untouched. If +the hub says the GitHub account is not registered, run +`hivectl hives add --hub ` instead. + +The legacy `just contribute-move` path is for positional `contributor.env` +setups without profiles: ```bash export HIVE_HUB=wss://hive.example.com/contribute @@ -478,6 +530,12 @@ just contribute-move claude `contribute-move` does everything `contribute-setup` does — backend preflight, `gh auth`, `gh-auth.env`, CLI config staging — except that instead of registering it calls `POST /api/contribute/reissue-token`, which authenticates with your GitHub token and therefore *can* prove you own the identity. It then writes `contributor.env` for you. +When `profiles.yml` exists, `contribute-move` aborts unless +`HIVE_FORCE_MOVE=1` is set. Forcing it is unsafe with profiles because it +rewrites `contributor.env` behind `profiles.yml`; the next `hivectl hives` +mutation regenerates `contributor.env` from `profiles.yml` and can restore the +old token. + **This rotates the credential.** Reissuing overwrites the stored hash, so a relay still running on the old machine stops authenticating the moment this succeeds. That is the point when you are moving off a machine you no longer want holding the token — but it means this is not the way to switch back and forth. Three things it does that a hand-rolled rotation makes easy to get wrong: @@ -571,10 +629,12 @@ After a contributor completes an issue, the hub keeps that issue out of the queu ### Queue hold and priority -The **Operations** tab also includes a public-safe **Most effective models** panel. It reads the same aggregate PR rework data as the Governor PRs-by-model view plus the contributor task-run log, then ranks model + CLI pairs for `7d`, `30d`, or `all` by visible measures: merged PRs, first-pass merge rate, average review rounds, average fix attempts, verified-PR run rate, failure rate, and completed-without-PR (“nothing to ship”) rate. Rows below `HIVE_CONTRIBUTE_EFFECTIVE_MODELS_MIN_PRS` merged PRs (default `5`) are shown under “Not enough data yet” instead of being ranked, and the response is aggregate-only: no contributor usernames, pane output, tokens, or per-contributor breakdowns. +The **Operations** tab also includes a public-safe, collapsible **Most effective models** panel. It reads the same aggregate PR rework data as the Governor PRs-by-model view plus the contributor task-run log, then ranks model + CLI pairs for `7d`, `30d`, or `all` by visible measures: merged PRs, first-pass merge rate, average review rounds, average fix attempts, verified-PR run rate, failure rate, and completed-without-PR (“nothing to ship”) rate. Rows below `HIVE_CONTRIBUTE_EFFECTIVE_MODELS_MIN_PRS` merged PRs (default `5`) are shown under “Not enough data yet” instead of being ranked, and the response is aggregate-only: no contributor usernames, pane output, tokens, or per-contributor breakdowns. The collapse state is remembered in browser localStorage so operators can keep the summary visible without scrolling past the controls and tables. The dashboard **Governor** card's **PRs by model** section uses that same model-effectiveness aggregation for the same `7d`, `30d`, and `all` windows. It keeps the merged/open/closed bar for PR volume, adds compact columns for merged count, first-pass merge rate, verified-PR run rate, failure rate, and “nothing to ship” rate, and badges ranked models that clear the merged-PR threshold. Operators can toggle the row order between effectiveness rank (default) and raw PR count. +The `/contribute/operations` card layout is also a per-viewer browser preference. Each card header has a grip that can be dragged with pointer/touch input or focused and moved with arrow keys; cards can move between the narrow, main, and full-width regions while the Live Activity rail stays fixed. The layout is stored in `localStorage` under `hive.ops.layout`, tolerates cards being added or removed between releases, and the **Reset layout** control restores the template order. + The **Operations** tab lets an operator reorder and park individual issues in the ready-work queue. Both controls persist on the hub configuration alongside the filters above, but they are edited only through two authenticated endpoints (owner or read-write role; a read-only or anonymous caller gets `403`): | Endpoint | Config key | Behavior | @@ -598,7 +658,7 @@ Withheld rows carry a stable reason code and, where the refusing gate had one, t | Reason | Meaning | Evidence | |---|---|---| | `open_pr_claim` | An open pull request already claims the issue. | Claiming PR URL and author | -| `issue_claim` | Someone has claimed the issue on the issue itself - a `hive-claim` marker comment posted by the hive's own App bot, or an assignee - and the claim has not expired ([#8380](https://github.com/hivecommons/hive/issues/8380)). Only while `governor.claims.enabled` is on. | `claimed_by` and `claim_expires_at` | +| `issue_claim` | Someone has claimed the issue on the issue itself - a `hive:claim` marker comment posted by the hive's own App bot, or an assignee - and the claim has not expired ([#8380](https://github.com/hivecommons/hive/issues/8380)). Only while `governor.claims.enabled` is on. | `claimed_by` and `claim_expires_at` | | `merged_claim_stale` | A merged pull request (or a verified `no_work_needed` verdict) has claimed to fix the issue for 7+ days and the issue is still open; the next step is a maintainer's — close it, or say what remains ([#8003](https://github.com/hivecommons/hive/issues/8003)). | Fixing PR URL and author; the age in days | | `issue_churn` | The issue has already absorbed several merged or abandoned pull requests without settling, so what is left is a maintainer's call ([#7995](https://github.com/hivecommons/hive/issues/7995)). | Merged / closed-unmerged counts and the PR numbers | | `workflow_blocked` | The issue carries the `blocked` workflow label. | Matched label | @@ -644,13 +704,30 @@ Contributors declare their CLI backend and model when the relay connects. The ** This is the admin's quality floor: a hive doing subtle refactors can require `claude-opus*`/`claude-sonnet*`, while a hive full of `good-first-issue` label work can accept anything, including local Ollama models. +### Minimum reasoning effort + +Relays also report their reasoning effort (`reasoning_effort` on `auth_response`). The **Effort floor** control, next to the Model Filter in **Governor → Hub**, sets a minimum: + +| Control | Config key | Behavior | +|---|---|---| +| **Effort floor** | `contribute_min_reasoning_effort` | Minimum reasoning effort on the ladder `minimal < low < medium < high < xhigh < max`. **Empty = no floor.** A relay whose effort ranks below the floor is rejected **at connect time** with an `auth_failed` message that echoes the floor (`min_reasoning_effort`). The comparison is per backend, not a raw string compare: a floor above a backend's highest level is met by that backend's highest level (e.g. a `max` floor is met by codex at `xhigh` and agy at `high`). Invalid values are rejected by `PUT /api/config/governor/hub` and ignored (with a warning) when loaded from YAML. | +| **Reject Unknown Effort** | `contribute_reject_unknown_effort` | Only applies when a floor is set. When on, a relay whose effort is empty, unrecognised, or not valid for its backend is rejected; when off (default) it is admitted. | + +A single ordered floor was chosen over a per-backend allow-list because effort vocabularies differ per backend (codex `model_reasoning_effort`, claude `--effort`, …): one floor on a shared ladder says "at least this much reasoning" once for every backend. Both keys are returned by `GET /api/config/governor` (alongside `contribute_reasoning_effort_ladder`) and on the contribute admission policy (`min_reasoning_effort`, `reject_unknown_effort`) so relays can pre-check before connecting. + +```yaml +hub: + contribute_min_reasoning_effort: high + contribute_reject_unknown_effort: true +``` + ### Trust tiers and individual controls Each trust tier can be toggled on/off and given its own rate limits (`0` = unlimited); tiers promote automatically as contributors complete tasks that open PRs. Admins can also promote, demote, or revoke individual contributors from the dashboard's contributor list (`GET /api/contributors`, with `PUT /api/contributors/{id}/trust` and `POST /api/contributors/{id}/revoke`); revoked contributors cannot reconnect. Completed-task counts and standings are public on the hive's `/leaderboard`. Tier names, promotion thresholds, and delegated roles are documented in [Contributor trust tiers and delegated agent roles](https://github.com/hivecommons/hive/blob/v5/src/docs/contributor-trust-and-roles.md). ### Filter timing -- **Queue-time vs. connect-time.** Repo, label, title, author, and assignment filters, cooldown, and the hold/priority sets apply when the queue is next built, so tightening them affects the *next* queue build. The Model Filter applies at connect time, so tightening it affects the *next* connection, not agents already mid-task. +- **Queue-time vs. connect-time.** Repo, label, title, author, and assignment filters, cooldown, and the hold/priority sets apply when the queue is next built, so tightening them affects the *next* queue build. The Model Filter and the effort floor apply at connect time, so tightening them affects the *next* connection, not agents already mid-task. - **Suspending vs. revoking.** Suspension idles everyone and is instant to undo; revocation is per-contributor and blocks reconnection. ## Kubernetes contributor workload @@ -793,24 +870,32 @@ When the agent determines that nothing is shippable until a maintainer makes a d The hub books the issue for the full with-PR cooldown immediately and records a `needs_decision` ledger/run-log marker. The relay applies the label advertised by the hub in `auth_ok` (`hub.contribute_needs_decision_label`, default `needs-decision`) using the same task credential path as the `blocked` label. Empty config disables relay labelling, but the cooldown still applies. The relay posts no extra comment: the contributor's assessment comment is the human-readable record. The label is never cleared automatically; removing it is the maintainer signal that the decision has been made. The configured label is also treated as a contributor skip label, so a custom label such as `2-discussing` keeps the issue out of the offer queue until a human removes it. -### A prompt that was typed is not a prompt that was submitted +### A prompt that was pasted is not a prompt that was submitted -The relay delivers a task prompt by typing it into the pane with `tmux send-keys -l` and then sending Enter. A task prompt is around 2 KB, so it arrives as one burst — and a TUI that implements bracketed-paste handling classifies a burst that fast as **pasted content**. codex collapses it to `[Pasted Content 1024 chars]` in its input widget and takes the Enters that follow as newlines *inside* the paste rather than as submit. The prompt sits in the widget, and the agent is never told anything. +The relay delivers a task prompt by storing it in a uniquely named tmux buffer (`tmux load-buffer -b … -`, prompt piped on stdin) and pasting it with bracketed paste (`tmux paste-buffer -p -d …`) before sending Enter. The older `tmux send-keys -l` path sent a task-sized prompt as one raw keystroke burst, and a TUI that implements bracketed-paste handling could classify a burst that fast as **pasted content**. codex collapsed it to `[Pasted Content 1024 chars]` in its input widget and took the Enters that followed as newlines *inside* the paste rather than as submit. The prompt sat in the widget, and the agent was never told anything. Observed live ([#6717](https://github.com/hivecommons/hive/issues/6717), codex-cli 0.154.0): the pane showed the launch banner, the collapsed prompt on the input line, no spinner, no tool rows and no assistant output at all, byte-identical across two consecutive five-minute checks. The relay logged `Task prompt sent to CLI` and, eight minutes later, `completed — signal=chrome_idle`. The hub booked the issue **done** with no commit, no branch and no PR, and it left `/api/contribute/queue`. `ENTER_COUNT = 3` is not the lever: the problem is not a dropped keystroke but a widget consuming newlines as content, and three are consumed exactly as one is. Three things changed instead. -1. **Settle before submitting.** The send path now waits for the widget to finish ingesting the burst before the Enter goes out, so the Enter is a keypress and not pasted text. -2. **Verify the submit.** It then re-reads the pane and re-sends Enter while the prompt is still visibly collapsed in the input widget, up to a small budget. The send loop already retried when *tmux* errored; it had never checked whether the keystrokes achieved anything. -3. **`chrome_idle` may not complete a task that never started.** The fallback infers "the agent finished" from a pane that stopped changing, and that inference had one premise it never checked: that the agent *started*. When the prompt is **still** visibly unsubmitted **and** the pane has not changed by a single byte since delivery, the task is reported `task_failed` with `failure_kind: environment` — so the hub re-offers the issue — instead of `task_complete`, which parks it as finished. +1. **Use bracketed paste for the prompt.** The prompt text is piped on stdin to `tmux load-buffer` via `execFileSync()`, never shell-interpolated and never subject to the per-argument length limit, and `paste-buffer -p -d` deletes the tmux buffer after delivery (with a best-effort cleanup fallback). +2. **Settle before submitting.** The send path still waits for the widget to finish ingesting the paste before the Enter goes out, so the Enter is a keypress and not pasted text. +3. **Verify that a turn started.** It then re-reads the pane and looks for positive evidence that the CLI accepted the prompt: working chrome, a tool row, an echoed prompt, or another non-idle change from the pre-delivery pane. A blank input widget with no activity is `submission unknown`, not success. If `[Pasted Content …]` appears after the first check, the relay re-sends Enter up to the existing small budget; it never re-pastes the prompt. +4. **`chrome_idle` may not complete a task that never started.** The fallback infers "the agent finished" from a pane that stopped changing, and that inference had one premise it never checked: that the agent *started*. When prompt submission is unconfirmed or visibly unsubmitted **and** the pane has not changed by a single byte since delivery, the task is reported `task_failed` with `failure_kind: environment` — so the hub re-offers the issue — instead of `task_complete`, which parks it as finished. -Both signals in (3) are required together, in both directions. Some CLIs echo a submitted paste back into their transcript with the same placeholder, so the placeholder alone would fail every task on such a backend; and a pane byte-identical to its pre-work state cannot belong to an agent that did anything. An agent that genuinely finished without printing `HIVE_VERDICT` still completes on the fallback exactly as it did before, which is the whole reason the fallback exists ([#5376](https://github.com/hivecommons/hive/issues/5376)). +Both signals in (4) are required together, in both directions. Some CLIs echo a submitted paste back into their transcript with the same placeholder, so the placeholder alone would fail every task on such a backend; and a pane byte-identical to its pre-work state cannot belong to an agent that did anything. An agent that genuinely finished without printing `HIVE_VERDICT` still completes on the fallback exactly as it did before, which is the whole reason the fallback exists ([#5376](https://github.com/hivecommons/hive/issues/5376)). The placeholder rendering is recorded **per backend, and only where a real capture has shown it** — codex today, in `bin/lib/pane-classifier.js` and the shared golden fixture `bin/testdata/pane-fixtures/codex_unsubmitted_paste.pane.txt`. A backend whose widget nobody has captured makes no claim either way, and gets neither the extra Enters nor the veto. This detector can fail a task, so a pattern guessed from another CLI's documentation would be a claim about a pane nobody has looked at, in the direction that costs the most. Finally, a `chrome_idle` completion carrying **neither** a verdict **nor** a PR is now logged as a warning. It is not always wrong — an agent that found nothing to do but never printed the sentinel lands there too — but it is the shape this bug takes, and nothing in the pane shows what such a task produced. +For a real tmux transport smoke test outside CI, run `bin/relay-paste-smoke.sh` on a host with tmux installed. It starts a disposable tmux session running `cat -A`, sends both a short and task-sized prompt via the same `load-buffer`/`paste-buffer -p -d` sequence, and verifies the pane received the bytes intact. + +That checks the tmux half only — the bytes land in a pane. Whether a CLI *consumes* them as one submitted turn is what [#9078](https://github.com/hivecommons/hive/issues/9078) actually broke, and a stubbed `child_process` cannot show it either way, so `bin/test_backend_smoke.sh` carries two delivery scenarios through the **real relay**: + +- **S4** (keyless, in CI's "Backend smoke (keyless subset)"): the relay, a fake hub, a real `tmux new-session -x 200 -y 50` pane, and a raw-mode codex-shaped stub that enables bracketed paste (`ESC[?2004h`), draws the real idle/working chrome, echoes the expanded prompt on submit the way codex's history cell does, reproduces the captured 0.157.1 raw-burst behaviour (nothing rendered, Enter swallowed, `[Pasted Content 4096 chars]` late), and records every submitted turn as JSONL. A 6,868-character prompt and a short control must each complete via verdict and arrive byte-for-byte as **exactly one** turn, with no `Task prompt delivery unconfirmed` in the relay log. On the pre-#9079 relay the long prompt starts zero turns, which is the incident. +- **B3** (credential-gated, on the scheduled `backend-smoke.yml` lane after the short B2 control): the same task-sized prompt against the **real** codex CLI on a 200×50 pane. Exactly-once, byte-for-byte delivery is read from the user message codex itself records in its session rollout under the throwaway `CODEX_HOME` (`event_msg/user_message` or the `response_item` user shape); when neither is found the check reports a skip naming the rollout directory rather than a false pass. B2 and B3 share `interactive_round()`. + ### Review notes that land after the verdict get one more turn omp's `--advisor` runtime reviews every turn passively and injects its notes after the turn ends, so its review of the agent's *final* turn is drawn on the pane **under** `HIVE_VERDICT`. The sentinel being final — which [#5376](https://github.com/hivecommons/hive/issues/5376), [#7662](https://github.com/hivecommons/hive/issues/7662) and [#7733](https://github.com/hivecommons/hive/issues/7733) established, and which is what makes omp bookable at all — meant the relay finalized on the line and killed the CLI with every note on the closing turn unread. Two live tasks on 2026-09-19 each ended under a stack of `⟦concern⟧`/`⟦nit⟧` notes; one was a real, cheap fix the agent would have made if it had seen it ([#7759](https://github.com/hivecommons/hive/issues/7759)). @@ -955,6 +1040,17 @@ A quota refusal now parks the loop: - **The operator is told once, clearly.** The banner previously lived only inside the agy pane while the relay log said `[environment]`; nobody reading the log could learn their quota was gone for four hours, or that switching model or backend was the remedy. - **The failure says what happened.** `[environment] … the agent CLI is not visibly working` reads as a broken contributor host. The CLI was working perfectly and the provider said no, so the reason now says so. +Codex's `■ You’ve hit your usage limit` banner also enters this hold, including +when it is followed by the optional model-switch menu ([#9247](https://github.com/hivecommons/hive/issues/9247)). +The relay preserves the displayed reset information in its failure reason and +operator log. A clock time such as `10:21 PM` has no established timezone, so +this hold does **not** expire on a guessed deadline or the generic fallback. +It survives CLI relaunches and hub reconnects, even if the new pane looks idle. +A fresh quota reading captured after the refusal, with every window above its +reserve, releases it. Without a quota reader, verify that the configured model +can run again, then restart the contributor relay explicitly. The relay never +selects the suggested model or purchases credits. + Only quota takes this path. An authorization refusal is not time-bounded, an operator has to change something, and parking the relay would hide it — a 403 still fails fast and stays available. The `failure_kind` on the wire is still `environment`: the hub's kinds are `environment` / `task` / `unspecified`, and the field is advisory — the hub records and displays it and does not route or change a work item's failure cooldown on it. A dedicated quota kind, and the cooldown exemption [#6541](https://github.com/hivecommons/hive/issues/6541) asks for, are a hub-side protocol change and are not part of this. @@ -1042,7 +1138,8 @@ Six behaviours are worth knowing: condition (b)): a fresh `.publisher.json` presence marker in the derived default dir flips a missing reading from the admit to a hold, because a live publisher has declared a reading is coming. Where no fresh marker exists — - a relay-only contributor machine runs no `hive` process and so no publisher, an + a relay-only contributor machine with the standalone publisher opted out or + unavailable (hivecommons/hive#10299) and so no publisher, an unsupported backend, an uninstalled CLI (the publisher removes the marker on a `not_installed` probe), or a dead publisher whose marker went stale — the admit stands, so no host is ever stranded holding with no route. Flipping THAT @@ -1057,6 +1154,26 @@ shape above. Two of the three read a different source than #6833's original adapter table named, and the divergence is recorded here so a later reader does not "fix" an adapter back to a source that was deliberately rejected: +- **Pi** — no built-in Pi-compatible quota reader is currently available, + including for `openai-codex`, `openrouter`, and `anthropic`. Pi's + `AGENT_MODEL=provider/model` selection is named in the startup diagnostic; + it does **not** imply Claude Code credentials or Codex CLI credentials. + Those readers use different credential stores and must not publish to Pi's + pool. Old automatic Pi pool readings are ignored, including a previous + `unknown` reading caused by missing Claude credentials. + + To retain quota protection, supply readings for the **selected Pi provider + and account** using `HIVE_CONTRIBUTOR_QUOTA_READING_FILE` (recommended for a + live, atomically refreshed source) or `HIVE_CONTRIBUTOR_QUOTA_READING_JSON`. + These explicit sources still take precedence and missing, malformed, + `unknown`, or `stale` readings still hold work; `continue-*` does not bypass + them. Setting only `HIVE_CONTRIBUTOR_QUOTA_POOL_DIR` does not create a Pi + reader. With no explicit source, Pi follows the unsupported-backend + `unprovisioned` behavior above: work is admitted, with a warning that quota + protection is unavailable and paid credits may be consumed. To explicitly + opt out instead, set `HIVE_CONTRIBUTOR_QUOTA_GUARD=off` at launch; this is + **not quota protection** and may spend paid credits. + - **Codex** — the app-server `account/rateLimits/read` method, as #6833 specified. All returned windows (`primary`, `secondary`, and any `rateLimitsByLimitId` scoped windows) fold in, worst window binds @@ -1457,7 +1574,7 @@ Trust note: hooks run with the entrypoint's full privileges inside the contributor container, and `HIVE_PRE_AGENT_HOOK` is `eval`'d verbatim — only bake hooks into images you build, and only pass `HIVE_PRE_AGENT_HOOK` values you would be willing to type into that container's shell yourself. Both knobs -are listed in the [environment variable reference](https://github.com/hivecommons/hive/blob/v5/src/docs/env-vars.md). +are listed in the [environment variable reference](/docs/hive/env-vars). ## Troubleshooting: the backend dies seconds after every task diff --git a/docs/content/hive/documentation-map.md b/docs/content/hive/documentation-map.md index e137acf..fcbd9ad 100644 --- a/docs/content/hive/documentation-map.md +++ b/docs/content/hive/documentation-map.md @@ -9,13 +9,17 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und ## Operations - [Manual provisioning](/docs/hive/manual-provisioning) — heartbeat-only cluster provisioning, hub access roles, and common gotchas. +- [DiskPressure recovery runbook](https://github.com/hivecommons/hive/blob/v5/src/docs/diskpressure-recovery-runbook.md) — spoke down with `/data` crash-looping on kubelet eviction: diagnosis, safe deletes, volume expansion/migration, and the health-check/governor-pause/self-janitor that degrade gracefully before it gets there. - [Hosted Hive Hub onboarding](https://github.com/hivecommons/hive/blob/v5/src/docs/hosted-hub.md) — signing in at `https://hive.hivecommons.dev`, requesting a hosted hive, installing the GitHub App, configuring model gateways and Copilot login, reading `/fleet`, and fixing common setup problems. - [Self-hosted hub deployment](https://github.com/hivecommons/hive/blob/v5/src/docs/hub-deployment.md) — `HIVE_MODE=hub`, hub storage, heartbeat secrets, and SaaS spoke registration. - [`CAP_NET_ADMIN` and self-hosted spokes](/docs/hive/net-admin-requirement) — the container runs with or without `NET_ADMIN`; granting it (`--cap-add NET_ADMIN` / `securityContext.capabilities.add`) enables the full forced-proxy-egress gate, and what the degraded best-effort mode means without it. - [Config layering](https://github.com/hivecommons/hive/blob/v5/src/docs/config-layering.md) — how ConfigMap seed, PVC dashboard overlay, and runtime config interact. - [Operator reference](https://github.com/hivecommons/hive/blob/v5/src/docs/operator-reference.md) — top-level config blocks, hive flags/env, GitHub token scopes, and image provenance. +- [Running a hive at ACMM Level 6](/docs/hive/running-at-level-6) — **awaiting operator review**: preparation, GitHub/coverage/knowledge/Hive/ClankeR settings, weekly routine, and incident recovery; sign-off tracked in #10515. +- [Maintainer commands](https://github.com/hivecommons/hive/blob/v5/src/docs/maintainer-commands.md) — slash-command reference for `/hive approve`, `/hive decision`, `/hive help`, `/fixed`, `/reopen`, and label/assignment helpers. - [Token mint](https://github.com/hivecommons/hive/blob/v5/src/docs/token-mint.md) — the opt-in `mint:` block (`pkg/mint`): what a minted token grants, key lifecycle, and the trust boundary an operator must get right before enabling it. Companion to [ADR-0007](/docs/hive/adr/0007-token-mint). - [Changelog](https://github.com/hivecommons/hive/blob/v5/CHANGELOG.md) — recent user-visible changes and release notes. +- [Work-source terminology](https://github.com/hivecommons/hive/blob/v5/src/docs/work-source-terminology.md) — glossary and audit notes for keeping user-facing copy neutral across GitHub, GitLab, Gitea, Linear, Jira, and future work sources. - [Release channels](/docs/hive/release-channels) — `stable`/`candidate`/`edge` moving image tags, per-line channel ownership (`v5` → `candidate`/`latest`, `v6` → `edge`, `v4` → none), switching a hive to a channel, and the version pill. - [Digest-verifiable rollback](https://github.com/hivecommons/hive/blob/v5/src/docs/release-rollback.md) - the operator runbook for pinning a hive back to a prior immutable short-SHA build for all three images (`hive`, `hive-contributor`, `hive-hub`), stopping the hub automation that would undo it, and verifying by digest rather than by tag that the pin landed on the running spoke. - [v5 GA readiness bar](https://github.com/hivecommons/hive/blob/v5/src/docs/v5-ga.md) — measurable release-train, migration, safety, and governance criteria that must be evidenced before v5 can be promoted beyond the active-development `edge` channel. @@ -28,13 +32,20 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [dibs domain cutover](https://github.com/hivecommons/hive/blob/v5/src/docs/dibs-domain-cutover.md) — staged operator sequence for moving dibs to `dibs.hivecommons.dev`, including DNS, Let's Encrypt quota hold, Certificate/Ingress manifests, redirect verification, and rollback. - [Serving spokes from the fleet wildcard certificate](https://github.com/hivecommons/hive/blob/v5/src/docs/spoke-wildcard-tls.md) — how to point a cluster's provisioned spokes at the wildcard instead of one certificate per hive, the two cluster prerequisites that must hold first, and which hosts a wildcard cannot cover. - [Tagged releases](https://github.com/hivecommons/hive/blob/v5/src/docs/releases.md) — the automated `v1.2.3` release path: what triggers a release, how the version is inferred from `CHANGELOG.md`, the commit convention that drives it, the human escape hatch, how it relates to the moving release channels above, and the per-release SPDX SBOM attached to each GitHub Release (and why it is a release artifact, not an in-image attestation — see #3760). +- [Release sentinel](https://github.com/hivecommons/hive/blob/v5/src/docs/release-sentinel.md) - the opt-in (`release_sentinel.enabled`, default off) bounded repair loop for a release whose CI fails after tagging: the per-release state machine (`awaiting_ci` -> `fixing` -> `green` | `failed` | `superseded`), stale-SHA handling, the round cap and timeout, which failures escalate to a human on the first round without a push, the separately opt-in retag after merge (`release_sentinel.retag_enabled`), and pre-tag release-workflow failures ([#9585](https://github.com/hivecommons/hive/issues/9585)). - [Dashboard-triggered standalone upgrades](https://github.com/hivecommons/hive/blob/v5/src/docs/dashboard-standalone-upgrades.md) — the dashboard upgrade button for standalone (Compose/Podman-Quadlet) hives: the `HIVE_DEPLOYMENT_RUNTIME`/`HIVE_DEPLOYMENT_PODMAN_MODE` runtime contract, the closed host-side helper (`hive-dashboard-upgrade-helper.sh`), and why an unproven runtime deliberately hides the button. +- [Work sources](/docs/hive/work-sources) — GitHub Issues, GitHub Projects, Linear, Jira Cloud, Jira Data Center / Server, run-stage, and Wavefront work enumeration. +- [Work source providers](/docs/hive/integrations/work-source-providers) — implement and contribute a provider that maps source-native work items onto `pkg/worksource.WorkSource`. +- [Hive Labels and Control Signals](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md) — authoritative operator reference for labels, holds, approvals, contributor skip labels, planning labels, and dashboard bands. The old [Hive issue labels](https://github.com/hivecommons/hive/blob/v5/src/docs/issue-labels.md) page redirects here. - [Spoke dashboard](https://github.com/hivecommons/hive/blob/v5/src/docs/dashboard.md) — the static dashboard FAQ panel contract: not ACMM-gated, no JS/fetch, grouped L1-L6/runs/contributors/claims/cost help, and guarded config-key references. +- [Dashboard feedback](https://github.com/hivecommons/hive/blob/v5/src/docs/feedback.md) — reporting bugs and requesting features from the spoke dashboard, including diagnostics, screenshot handling, hub relay, user-auth, and fallback issue paths. - [Dashboard design system](https://github.com/hivecommons/hive/blob/v5/src/docs/dashboard-design-system.md) — shared token catalogue, component variants, migration rules, ratchet plan, and #8536 theme override contract for spoke, contributor, and hub dashboard surfaces. - [The `auto-update` Compose profile](https://github.com/hivecommons/hive/blob/v5/src/docs/auto-update-profile.md) — what unattended Watchtower updates cost you, what the Docker socket proxy does and does **not** fix, and why Kubernetes should not use this profile at all. -- [Environment variable reference](https://github.com/hivecommons/hive/blob/v5/src/docs/env-vars.md) — centralized list of runtime, deployment, hub, backup, and contributor environment variables. +- [Environment variable reference](/docs/hive/env-vars) — centralized list of runtime, deployment, hub, backup, and contributor environment variables. - [Kubernetes deployment](https://github.com/hivecommons/hive/blob/v5/README.md#kubernetes-deployment) — the operator path for Kubernetes: prerequisites, namespace, secret, ConfigMap, PVC, Deployment, Service, Ingress, and published ports. Lives in the root README alongside the Compose and Podman quick starts; the manifests it applies are [`src/deploy/k8s/`](https://github.com/hivecommons/hive/blob/v5/src/deploy/k8s/). See also [dashboard route and health checks](https://github.com/hivecommons/hive/blob/v5/src/docs/health-checks.md) and the Kubernetes CronJob in [backup and restore](/docs/hive/backup-dr). - [Troubleshooting](/docs/hive/troubleshooting) — container logs, config validation, agent tmux sessions, dashboard auth, GitHub credential checks, and GitHub App workflow-permission push rejections. +- [Community and support](https://github.com/hivecommons/hive/blob/v5/src/docs/community.md) — Discord, contributor portal, issue tracker, and where to ask Hive Commons questions. +- [Integration guide](/docs/hive/integration-guide) — third-party extension surfaces for work source providers, ClankeR/Flue-style external execution, and Spektacular-backed Project Inception. - [Cross-cluster migration](https://github.com/hivecommons/hive/blob/v5/src/docs/cross-cluster-migration.md) — the manual procedure for moving a hive between clusters. - [Self-hosted Kubernetes cluster move](https://github.com/hivecommons/hive/blob/v5/src/docs/move-kubernetes.md) — moving a self-hosted (non-hub) Kubernetes hive between clusters: which PVC/Secret/ConfigMap paths hold identity and state, a generic `kubectl`-based PVC copy, target prerequisites, and verification. - [Hub-registered hive cutover](https://github.com/hivecommons/hive/blob/v5/src/docs/move-hub-registered-cutover.md) — cross-cutting rules for any hub-registered hive's move: what identity is, how the hub reads heartbeat/`dashboard_url`, why source and target must never run concurrently, the required cutover order, and rollback. @@ -48,7 +59,7 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [Fleet drift signals](https://github.com/hivecommons/hive/blob/v5/src/docs/fleet-drift-signals.md) — the per-hive deviation badges on My Hives: all signal kinds (`heartbeat-stale`, `duplicate-spoke`, `identity-split`, `version-absent`, `pinned-image`, …) with severities and who fixes each (owner vs hub operator), the derived fleet norm behind `branch-mismatch`/`version-behind`, the deliberate suppression rules (placeholders, actively-upgrading hives, status-flipping yielding to duplicate-spoke), and the in-memory first-seen semantics. - [Fleet self-reporting (`governor.fleet_report`)](https://github.com/hivecommons/hive/blob/v5/src/docs/fleet-report.md) — how a hive reports its own hive-attributable failures upstream to `hivecommons/hive`: the `file_upstream` opt-in (default off = dry-run preview on the dashboard), the two triggers (`acmm-shortfall` after two unmet weekly epochs with attributable evidence, and `hive-code-defect` on hive's own component allowlist with no shortfall required), exactly which fields leave the hive and which never do (the raw hive ID is replaced by a truncated SHA-256), fingerprint deduplication via comment + 👍 reaction instead of duplicate issues, and recovery comments that close only issues the hive itself opened. - [Agent self-healing watchdog](https://github.com/hivecommons/hive/blob/v5/src/docs/agent-watchdog.md) — liveness and readiness reconciliation for launched agents: liveness classification, restart backoff, crash-loop escalation, the auth probe that refuses to restart into dead credentials, and the `conditions` array on `/api/agents`. Ships in `mode: observe`, which audits the restarts it would have made without making them. -- [Audit log format](https://github.com/hivecommons/hive/blob/v5/src/docs/audit-log.md) — the JSONL schema of `/data/audit.jsonl`: the five fields, how to parse the flat `detail` string (and why `repo` is not first-class), the pseudo-users, and why size-triggered rotation means the effective lookback varies per hive rather than being 90 days. +- [Audit log format](https://github.com/hivecommons/hive/blob/v5/src/docs/audit-log.md) — the JSONL schema of `/data/audit.jsonl`: its fields, how to parse the flat `detail` string (and the typed `repo`/`target` fields on GitHub write entries), the pseudo-users, and why size-triggered rotation means the effective lookback varies per hive rather than being 90 days. - [Per-repo agent pause](https://github.com/hivecommons/hive/blob/v5/src/docs/repo-pause.md) — quieting ONE repository without stopping the hive: `project.paused_repos`, the dashboard toggle and `POST /api/repos/pause`, the provenance every pause carries (who/when/why), and exactly which layers enforce it — the MITM proxy, the `hive-open-pr`/`hive-merge` relays, work enumeration and the auto-merge sweeps — plus what a pause deliberately does *not* stop (reads, the hive's own control plane, and your own pushes). - [Token-access audit log](https://github.com/hivecommons/hive/blob/v5/src/docs/token-access-log.md) - the *other* audit log: `/var/run/hive-metrics/token-access.jsonl`, appended by `gh-wrapper.sh` on every agent `gh` call and by `git-credential-hive.sh` on every credential lookup, served (owner-gated, last 100 lines) by `GET /api/token-access`. Covers both line schemas, the entrypoint pre-creation/permission model, why an empty log is not "no activity", and the blind spots (contributor mode, wrapper bypass, tmpfs non-durability). - [Per-repo agents](https://github.com/hivecommons/hive/blob/v5/src/docs/per-repo-agents.md) — scoping an agent to the repositories it serves with `repos:`, so which agents exist is a per-repo answer instead of a hive-wide one: what a scope narrows (kick contents, `$HIVE_REPO`/`$HIVE_REPOS`, the AUTHORIZED REPOS block) and what enforces it deterministically (the MITM proxy plus the `hive-open-pr`/`hive-merge`/`hive-open-issue` relays), the optional `repos:` key on a BYO `AgentSpec` and why it is an optional interface rather than a sixth contract method, the `repos_owner` marker that keeps a pack apply from widening a specialist, and what a scope deliberately does *not* block. @@ -64,7 +75,9 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [State-triggered hooks](https://github.com/hivecommons/hive/blob/v5/src/docs/hooks.md) — declarative `transition → action` rules, the transition catalog, the vetted action set, and the security model (RFC #4001). - [CEL-based agent triggers](https://github.com/hivecommons/hive/blob/v5/src/docs/cel-triggers.md) — the `triggers:` config key: declarative CEL rules that kick an agent on a normalized source-control event, additive to built-in label/governor triggering, the `event.*` field reference, and the fail-closed compile/runtime contract. - [Long-running runs](https://github.com/hivecommons/hive/blob/v5/src/docs/runs.md) — how a run starts, including default-off triage admission and inception completion admitting an explicit GitHub issue into the first `spec` lease. -- [Spektacular stage runner](https://github.com/hivecommons/hive/blob/v5/src/docs/spektacular.md) — the Hive side of long-running runs (#8303, umbrella #8290): a run moves through `spec` → `plan` → `implement` stages on one task lease, Spektacular owns each artifact's state while Hive owns the workflow. Covers the default-off `runs.spektacular` config block and Governor Features toggle, the `max_stage_retries` budget, and what advances a lease. The design record for the artifacts a run leaves behind is [`design/run-artifacts.md`](https://github.com/hivecommons/hive/blob/v5/src/docs/design/run-artifacts.md). +- [Spektacular (Spek) stage runner](/docs/hive/integrations/spektacular) — the Hive side of long-running runs (#8303, umbrella #8290): a run moves through `spec` → `plan` → `implement` stages on one task lease, Spek owns each artifact's state while Hive owns the workflow. Covers the default-off `runs.spektacular` config block and Governor Features toggle, the `max_stage_retries` budget, and what advances a lease. The design record for the artifacts a run leaves behind is [`design/run-artifacts.md`](https://github.com/hivecommons/hive/blob/v5/src/docs/design/run-artifacts.md). +- [ClankeR and Flue-style external execution](/docs/hive/integrations/clanker-flue) — how the contributor relay and `pkg/extwork` let an external engine receive bounded work, report status, and return digest-verified receipts without repository credentials. +- [Spektacular and Project Inception](/docs/hive/integrations/spektacular) — the actual CLI seam Hive uses for `spec`/`plan` status, final-plan import, and Project Inception admission. - [Audit campaign](https://github.com/hivecommons/hive/blob/v5/src/docs/audit-campaign.md) — the runs three-gate model proven on a workload that never opens a pull request (#8327): owner-triggered activation, campaign/inspection/finding beads, shadow-mode execution with deterministic finding identity, per-inspection stage receipts and proof predicates, and the guard that refuses to run while `HIVE_GITHUB_TOKEN` is present. Publication remains default-off and is skipped unless `publication.enabled` is set. - [Public snapshots](https://github.com/hivecommons/hive/blob/v5/src/docs/snapshots.md) — read-only `/snapshot`, custom CSS, and frame-ancestor sharing. - [hivectl](/docs/hive/hivectl) — command-line client for the dashboard API, including @@ -97,19 +110,23 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [Formal verification](https://github.com/hivecommons/hive/blob/v5/src/docs/formal-verification.md) — the optional L5/L6 `quality.formal` capability for agent-authored Spin/Promela models, reporting-only CI, and deduplicated counterexample issues. - [Advisory digest](https://github.com/hivecommons/hive/blob/v5/src/docs/advisory.md) — what the digest shows (`max_findings`, `show_all`), how stale findings are marked unverified, and which positive signals retire findings. - [Advisory digest staleness](https://github.com/hivecommons/hive/blob/v5/src/docs/advisory-staleness.md) — when the hub raises the stale-advisory pill and alert, the gates that deliberately suppress it (undelivered App, App cannot write, all agents quiet), and the admin diagnostics that measure hidden staleness. +- [Guide flow-health diagnosis](https://github.com/hivecommons/hive/blob/v5/src/docs/guide-flow-health.md) — the `HIVE_FLOW:` kick signal and its field contract (`surge`, `mttm_sample`, `oldest_actionable`, …), the surge-coach lens that prioritizes a documentation finding under sustained congestion, and the bucketed finding/follow-up policy, without a new agent, timer, or permission grant. - [Governor mode thresholds](https://github.com/hivecommons/hive/blob/v5/src/docs/governor-thresholds.md) — how idle/quiet/busy/surge thresholds scale with repo count, the `threshold_scaling` curves, and when explicit thresholds win. - [Large-spoke scale envelope](https://github.com/hivecommons/hive/blob/v5/src/docs/scale-envelope.md) — the backlog size one spoke is known to work at, why kick-prompt caps are render-side while enumeration is uncapped, and an inventory of every cap that changes behaviour at scale (default, overflow behaviour, and whether it is operator-tunable). - [Supervisor agent](https://github.com/hivecommons/hive/blob/v5/src/docs/supervisor.md) — supervisor policy modes, bead roles, and when to enable the orchestration lane. - [Telemetry agent](https://github.com/hivecommons/hive/blob/v5/src/docs/telemetry.md) — the L5/L6-only opt-in observability agent, ACMM level gating, and the `project_observability` opt-in flow. +- [NPS feedback prompt](https://github.com/hivecommons/hive/blob/v5/src/docs/nps.md) - the dashboard's opt-in NPS question: what is collected (no user identity), how it reaches the hub (directly, or through the opt-in relay for standalone hives), why only hub admins can read it, and the `hub.nps_enabled` / `HIVE_NPS_ENABLED` switch. - [Operations agent](https://github.com/hivecommons/hive/blob/v5/src/docs/operations.md) — the L5/L6-only opt-in operational-readiness agent (health checks, SLOs, runbooks), ACMM level gating, and the `project_observability` opt-in flow. - [Custom dashboard stylesheets](https://github.com/hivecommons/hive/blob/v5/src/docs/custom-stylesheets.md) — operator-supplied CSS for the dashboard and public snapshot. - [Branding a hive](https://github.com/hivecommons/hive/blob/v5/src/docs/branding.md) — persistent per-deployment name, mark, and colours via `/branding/branding.json` and `custom.css` (`HIVE_BRANDING_JSON`/`HIVE_BRANDING_CSS`); distinct from the per-request `?style=` stylesheets above. - [Portable AgentDefinition format](https://github.com/hivecommons/hive/blob/v5/src/AGENT-DEFINITION.md) — standalone YAML schema for importing/exporting agent definitions. -- [Knowledge curator](https://github.com/hivecommons/hive/blob/v5/src/docs/knowledge-curator.md) — automatic fact extraction and promotion knobs, plus the other `knowledge:` sub-sections: `git_sources` (indexing a remote repo, layer semantics, private-repo auth (unsupported), diagnosing a failed source), `vaults` (local Obsidian vaults, git-sync), `documents` (PDF/URL import), and `bead_synthesizer` (on-by-default bead→wiki synthesis and retention policy, and how to turn it off). +- [Knowledge curator](https://github.com/hivecommons/hive/blob/v5/src/docs/knowledge-curator.md) — automatic fact extraction and promotion knobs, plus the other `knowledge:` sub-sections: `git_sources` (indexing a remote repo, layer semantics, private-repo auth (unsupported), diagnosing a failed source), `vaults` (local Obsidian vaults, git-sync), `documents` (PDF/URL import), and `bead_synthesizer` (on-by-default bead→wiki synthesis and retention policy, and how to turn it off). See also [knowledge confidence scoring](https://github.com/hivecommons/hive/blob/v5/src/docs/knowledge-confidence.md) for dashboard confidence badges and score explanations. +- [Public knowledge MCP endpoint](/docs/hive/public-knowledge-mcp) — the owner-switched, read-only `POST /mcp/knowledge` surface that lets any external agent (Goose, Claude, Copilot) read a hive's operational knowledge without a session; what is and is not exposed, and client configuration. - [Skill registry](https://github.com/hivecommons/hive/blob/v5/src/docs/skills.md) — the `/data/skills/` file format and front-matter fields. **Delivered to agents today**: skills are opt-in per agent via `skills: [name, ...]`, resolved registry-first with an `AGENTS.md` repo-local fallback, and the rendered block is prepended to the kick's `${KNOWLEDGE}`. Both sources reload on every kick. An 8 KiB whole-skill cap applies; a skill that would exceed it is dropped whole, never truncated. - [AGENTS.md repo instructions](https://github.com/hivecommons/hive/blob/v5/src/docs/agents-md.md) — the per-repo `AGENTS.md` file format Hive's parser (`pkg/agentsmd`) understands, including front-matter `skills:` and inline `## Skill:` sections. **Wired into kicks, but needs a checkout**: Hive agents keep no clones, so set `project.checkouts_dir` to a directory holding one checkout per repo. Without it there is no root to read and injection stays a no-op, which is the default. - [Agent peer-awareness logging (pluk)](https://github.com/hivecommons/hive/blob/v5/src/docs/agent-logging.md) — pluk log format, `hive-panes`, availability, and retention. - [Strategy Lab (Nous)](https://github.com/hivecommons/hive/blob/v5/src/docs/strategy-lab.md) — experiment lifecycle, dashboard/API configuration, fast-fail bounds, and the gate-decision flow. No `nous:` block in `hive.yaml`. +- [GitHub write surface](https://github.com/hivecommons/hive/blob/v5/src/docs/github-write-surface.md) - every GitHub write the hive performs for an agent (the request relays) or by itself, whether each is audited with typed `repo`/`target`, the per-lane `write_surface.allowlist`, the per-lane `write_surface.enforce` opt-in that refuses a lane's direct writes at the proxy and in the agent sandbox, and credential redaction in audit entries. - [GitHub App setup](https://github.com/hivecommons/hive/blob/v5/src/docs/github-app-setup.md) — the Forge App on GitHub and GitHub Enterprise: app creation, permissions, Setup URL, and `/gh-setup`. - [Forge setup: GitLab, Gitea, and Forgejo](https://github.com/hivecommons/hive/blob/v5/src/docs/forge-app-setup.md) — the non-GitHub forges. **Adapters exist and are tested, but are not wired into any running code path**: a hive cannot run against GitLab, Gitea, or Forgejo today, and `project.forge` only changes what the dashboard displays. Covers the `gitlab:`/`gitea:` config surface that does parse, why the `gh`-CLI agent path is GitHub-only, and how `project.forge` differs from `github.forge`. - [ACMM policy matrix](/docs/hive/acmm-policy-matrix) — capability levels and policy modes. @@ -119,9 +136,10 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [Planning intelligence](https://github.com/hivecommons/hive/blob/v5/src/docs/planning-intelligence.md) — how a large GitHub issue becomes an epic the architect lane decomposes into child beads, the human plan-review gate that withholds those children until approved, and stall-replan. - [Review swarm](https://github.com/hivecommons/hive/blob/v5/src/docs/review-swarm.md) — the five review perspectives, the verdict collector and its report contract, and the opt-in merge-gate integration and bounded auto-fix cycle for review findings. - [Duplicate PR sweep](https://github.com/hivecommons/hive/blob/v5/src/docs/duplicate-sweep.md) — the opt-in (`duplicate_sweep.enabled`) cross-PR pass that clusters open PRs by changed-file set and suggests which one to keep: the two confidence tiers, how a survivor is chosen, why a bot regeneration series gets one summary instead of a comment per PR, and why the output is always a suggestion a human acts on and never an automated close. +- [Upstream watch](https://github.com/hivecommons/hive/blob/v5/src/docs/upstream-watch.md) — the opt-in `upstream_watch` block (per-repo `upstream` or the GitHub fork parent, `sources`, `pr_labels`, `max_issues_per_run`, `label`) and the poll loop it drives: classification, fork applicability, labelled `upstream:` issues with a hidden `upstream-ref` marker, dedupe and dismissal, and the `/data/upstream-watch.json` state file. - [Hold-gated review queue triage](https://github.com/hivecommons/hive/blob/v5/src/docs/review-queue-triage.md) — the triage-class policy for the human review queue (T0 fixes, T1 behavior-adjacent, T2 additive), the shipped `review_class` presentational sort in `last-actionable.json` and the dashboard PR list, and the open mechanism options awaiting maintainer decisions (#6183). - [Retro lane](https://github.com/hivecommons/hive/blob/v5/src/docs/retro-lane.md) — the opt-in (`retro.enabled`) post-completion pass that reconstructs a record for each closed bead and flags patterns such as excessive fix attempts or kicks; deterministic by default, with LLM analysis separately opt-in. -- [Work sources](https://github.com/hivecommons/hive/blob/v5/src/docs/work-sources.md) — `governor.work_source`: the four `type` options (`github` default, `github_projects`, `linear`, `jira`), config fields, required credentials, and priority/hold-label mapping per source. +- [Work sources](/docs/hive/work-sources) — `governor.work_source`: the four `type` options (`github` default, `github_projects`, `linear`, `jira`), Jira Cloud/Data Center deployment modes, config fields, required credentials, and priority/hold-label mapping per source. - [Linear agent integration](https://github.com/hivecommons/hive/blob/v5/src/docs/linear-agent.md) — joining a Linear workspace as a first-class agent member: webhook verification, the 10-second session acknowledgement, which hive agent takes sessions, and narrating completion back as agent activities. - [Lite enrollment](https://github.com/hivecommons/hive/blob/v5/src/docs/lite-enrollment.md) — the zero-repo-secret on-ramp: `hivectl enroll OWNER/REPO` adds a repo to a spoke's `project.repos`, with prerequisites and the hosted lite-spoke path. - [ACMM policy fragments](https://github.com/hivecommons/hive/blob/v5/examples/acmm/README.md) — per-level ACMM policy references. @@ -150,7 +168,7 @@ Documentation for the current Hive line (branch `v5`; the code and docs live und - [Podman CI runner map](https://github.com/hivecommons/hive/blob/v5/src/docs/podman-ci-runner-map.md) — measured hosted-runner capabilities and which Podman lane goes where; SELinux is the only lane needing non-hosted infrastructure. - [CI runner labels](https://github.com/hivecommons/hive/blob/v5/src/docs/ci-runner-labels.md) — how `runs-on:` picks the self-hosted fleet, why a fork must degrade to a GitHub-hosted runner, and the one variable `hivecommons` sets. - [Gate-integrity invariants for agent lanes](https://github.com/hivecommons/hive/blob/v5/src/docs/design/gate-integrity-invariants.md) — proposed write-gate rules for agent lanes: no history rewrites on branches a lane did not create, no sign-off on other authors' commits, and hold-gated demotion for gate manipulation. -- [Design documents](https://github.com/hivecommons/hive/blob/v5/src/docs/design/README.md) — longer-form design records with the full reasoning behind a decision, indexed with a status each (shipped / partly shipped / design only / historical) so a proposal is not mistaken for current behaviour: master secret rotation, wrapped master delivery to pull-only spokes, PR reach telemetry, and the knowledge system. +- [Design documents](https://github.com/hivecommons/hive/blob/v5/src/docs/design/README.md) — longer-form design records with the full reasoning behind a decision, indexed with a status each (shipped / partly shipped / design only / historical) so a proposal is not mistaken for current behaviour: master secret rotation, wrapped master delivery to push-reported spokes, PR reach telemetry, and the knowledge system. - [Discord reaction-consensus issue promotion](https://github.com/hivecommons/hive/blob/v5/src/docs/design/discord-issue-promotion.md) — proposed design for turning Discord reaction consensus into an audited Hive label write that composes with `project.issue_filter.require_labels` (RFC #6239). - [GitHub @-mention triggers](https://github.com/hivecommons/hive/blob/v5/src/docs/design/github-mention-triggers.md) — v6 design, now shipped on the `v6` branch in all three phases (#7582, #7597, #7623), for the first inbound GitHub trigger: a human summons an agent by mentioning the App on an issue or PR, mirroring the Linear agent-session path, replying through `Converse` on the existing write path, with poll-first transport and guards mapped to mechanisms that already exist (RFC #7483). Not present on `v4` or `v5`. - [Podman Compose-provider selection spike](https://github.com/hivecommons/hive/blob/v5/src/docs/podman-compose-provider-spike.md) — why `podman compose` must name its provider explicitly, and which provider needs no Docker tooling. @@ -164,6 +182,7 @@ Some documents describe planned or design-only work rather than live features. T ## Security (v4) +- [Securing your hive: a first-time operator's guide](/docs/hive/securing-your-hive) — decision-oriented walkthrough of ACMM level, reporter trust, and repo scope: starting postures, the #9758/#9762 real-world walk-through, Q&A, a pre-L6 checklist, and a glossary. - [Security model — operator guide](/docs/hive/security-model) — Ed25519-only sessions/SSO, per-hive keys, master key rotation, forced proxy egress and `CAP_NET_ADMIN`, privilege model, and supply-chain posture. - [Security threat model](/docs/hive/security-threat-model) — actors, boundaries, layered defenses, known gaps, and reporting. - [Security response process](https://github.com/hivecommons/hive/blob/v5/src/docs/security-response.md) — who responds to a vulnerability report (the Maintainer Committee, rostered in `OWNERS`), the end-to-end handling flow and the 60-day fix commitment, how membership is added and rotated, the escalation path if a reporter gets no response, and the project's known limits stated plainly. diff --git a/docs/content/hive/env-vars.md b/docs/content/hive/env-vars.md index aa7cdd2..df7eb2a 100644 --- a/docs/content/hive/env-vars.md +++ b/docs/content/hive/env-vars.md @@ -50,8 +50,8 @@ This reference is compiled by hand from the Go source under `src/`, the deployme | `HIVE_METRICS_ENABLED` | No | disabled | Registers Prometheus `/metrics` when set to `1`, `true`, `yes`, or `on`. Requires `HIVE_METRICS_TOKEN` — enabled-but-tokenless returns 403 ([#3804](https://github.com/hivecommons/hive/pull/3804)). | | `HIVE_METRICS_TOKEN` | Yes when metrics enabled | none | Bearer token read by `pkg/dashboard/metrics_prometheus.go` for `/metrics` (`Authorization: Bearer ...`; Prometheus `bearer_token`). `/metrics` bypasses dashboard session auth, so this token is its only guard; enabled-but-tokenless fails closed and the cost/agent series are never served without it. | | `HIVE_METRICS_FILE` | No | `/var/run/hive-metrics/contribute.json` | Contributor metrics JSON file override. | -| `HIVE_PUBLIC_KNOWLEDGE` | No | disabled | Owner switch for the anonymous, **read-only** MCP knowledge endpoint `POST /mcp/knowledge` (`1`, `true`, `yes`, or `on` — same spelling as `HIVE_METRICS_ENABLED`). Off → 404 even for authenticated callers. Read on each request so it can be closed without a pod roll. Only operational fact types (pattern, gotcha, regression, test_scaffold, integration, coverage_rule, general) are served; ideation/governance types, sources, and usage data never leave the hive. See [public-knowledge-mcp.md](/docs/hive/public-knowledge-mcp) ([#10615](https://github.com/hivecommons/hive/issues/10615)). | -| `HIVE_PUBLIC_KNOWLEDGE_TAGS` | No | none (all public-type facts) | Comma-separated tag allow-list that further narrows what `/mcp/knowledge` serves to facts carrying at least one listed tag (case-insensitive). | +| `HIVE_PUBLIC_KNOWLEDGE` | No | disabled | Env fallback for the anonymous, **read-only** MCP knowledge endpoint `POST /mcp/knowledge` (`1`, `true`, `yes`, or `on` — same spelling as `HIVE_METRICS_ENABLED`). The owner dashboard toggle in the Knowledge header and Settings → Knowledge takes precedence once saved; with no saved setting this remains the switch. Off → 404 even for authenticated callers. Resolved on each request so it can be closed without a pod roll. Only operational fact types (pattern, gotcha, regression, test_scaffold, integration, coverage_rule, general) are served; ideation/governance types, sources, and usage data never leave the hive. See [public-knowledge-mcp.md](/docs/hive/public-knowledge-mcp) ([#10615](https://github.com/hivecommons/hive/issues/10615)). | +| `HIVE_PUBLIC_KNOWLEDGE_TAGS` | No | none (all public-type facts) | Env fallback comma-separated tag allow-list that further narrows what `/mcp/knowledge` serves to facts carrying at least one listed tag (case-insensitive). A saved dashboard tag list takes precedence. | | `HIVE_COPILOT_INTEGRATION_ID` | No | compiled Copilot integration id | Overrides the integration id used by Copilot model discovery. | | `HIVE_CONTRIBUTORS_DIR` | No | hub default | Contributor registry directory override. | | `HIVE_CONTRIBUTE_SKIP_LABELS` | No | `blocked,tracking,epic,discussion,question,needs-decision,needs-triage` | Comma-separated, case-insensitive label patterns for issues that are not contributor work and must never be offered by the relay. Patterns use `path.Match`-style `*` globs; `blocked` is always unioned into the effective set even if omitted. Same setting as `hub.contribute_skip_labels`. | diff --git a/docs/content/hive/getting-started.md b/docs/content/hive/getting-started.md index 91bece3..e34956f 100644 --- a/docs/content/hive/getting-started.md +++ b/docs/content/hive/getting-started.md @@ -2,7 +2,7 @@ # Zero to Automation: Getting Started with Hive -Hive is a team of AI agents that watch your repo and help improve it — finding bugs, adding tests, writing docs. It works in **levels (L1–L6)**: at low levels agents only *suggest* things, and at high levels they can open and even merge pull requests. You climb the levels as you build trust in what the agents produce — over **weeks per level, not days**. And here's the most important thing to know before you start: **the goal is trust, not level.** +Hive is a team of AI agents that watch your project and help improve it — finding bugs, adding tests, writing docs. It works in **levels (L1–L6)**: at low levels agents only *suggest* things, and at high levels they can open and even merge pull requests. You climb the levels as you build trust in what the agents produce — over **weeks per level, not days**. And here's the most important thing to know before you start: **the goal is trust, not level.** ## The Hive Way @@ -12,13 +12,15 @@ The biggest mistake new users make: seeing agent output and either (a) panicking ## Trust > Level. Always. -> You do not need to reach L6. Ever. L6 is full automation — agents merging code without human review. Some teams run at L4 or L5 indefinitely and that is completely fine. The number doesn't matter. What matters is whether you trust what the agents are producing. A team that runs at L3 with high confidence is in a better place than a team that jumped to L6 and is now drowning in agent PRs they don't understand. +> You do not need to reach L6. Ever. L6 is full automation — agents merging code without human review. Some teams run at L4 or L5 indefinitely and that is completely fine. The number doesn't matter. What matters is whether you trust what the agents are producing. A team that runs at L3 with high confidence is in a better place than a team that jumped to L6 and is now drowning in agent change requests they don't understand. > > **The goal is trust, not level.** ## What this guide doesn't cover (and why) -> Hive is deeply configurable. There are agent policy templates, knowledge layers, custom agents, issue label filters, multi-repo setups, and a lot more. This guide doesn't cover any of that — and that's intentional. You don't need any of it to start. The goal of your first few months is to get comfortable with one or two agents at a low level, not to explore every feature. Features will still be there when you're ready for them. +> Hive is deeply configurable. There are agent policy templates, knowledge layers, custom agents, issue label filters, multi-repo setups, and a lot more. This guide doesn't cover any of that — and that's intentional. You don't need any of it to start. The goal of your first few months is to get comfortable with one or two agents at a low level, not to explore every feature. Features will still be there when you're ready for them. When you are ready for label behavior details, use [Hive Labels and Control Signals](https://github.com/hivecommons/hive/blob/v5/src/docs/labels-and-control-signals.md). + +When you are ready for a specialist, the dashboard has **+ Add agent** in the Agents sidebar and at the top of the Agents section. It opens a create dialog with quick-start templates for common roles (scanner, reviewer, quality, CI maintainer, guide) plus import-from-YAML for shared agent definitions. ## What to expect (and what not to) @@ -27,7 +29,7 @@ The biggest mistake new users make: seeing agent output and either (a) panicking - A slow start — days or weeks before anything meaningful happens - Findings you already knew about — agents often surface obvious things first - Some findings you disagree with — that's normal, decline them and move on -- PRs with hold labels — you control every merge below L6 +- PRs with a literal `hold` level-gate label — you control every merge below L6; dashboard manual holds use `hive-pause/` - Gradual improvement in finding quality as agents learn your codebase ❌ **Don't expect:** @@ -59,6 +61,10 @@ None of the level guidance below works until your hive is connected to your git --- +## Community help + +If you get stuck, [Join our Discord](https://hivecommons.dev/discord). The invite is permanent, and the Hive Commons community can help with setup, first-run questions, contributor relay, and choosing a safe next ACMM step. + ## Common gotchas (so you don't panic) - **Dashboard full of warnings?** Normal. Most warnings clear automatically after the Forge App is installed and the first heartbeat runs. Don't panic. @@ -128,7 +134,7 @@ Read it like a weekly digest, not a to-do list. You don't have to act on everyth Every issue a hive agent opens has the agent's name in the title — for example `[scanner] Possible nil pointer dereference in handler.go:142` or `[quality] Missing test coverage for payment flow`. -When you see a new issue in your repo, check the title prefix — it tells you which agent filed it and what kind of finding it is: +When you see a new issue in your project, check the title prefix — it tells you which agent filed it and what kind of finding it is: - `[scanner]` = bugs - `[quality]` = test gaps @@ -146,12 +152,12 @@ New users often expect PRs at L2 (they don't happen) or are surprised when they | Level | GitHub activity | |-------|-----------------| | **L1, L2** | No issues, no PRs. Dashboard beads only. If you see no repo activity, that's correct. | -| **L3** | **Quality only** can open PRs. Every PR has a `hold` label — it will NOT merge until you remove the hold. No other agent opens PRs at L3. | -| **L4** | Quality **and** sec-check can open PRs (both with hold labels). Scanner and guide file issues — not PRs. | -| **L5** | All agents can open PRs, all with hold labels. Nothing auto-merges. You batch-review. | -| **L6** | PRs auto-merge when CI goes green. No hold labels. Full automation. | +| **L3** | **Quality only** can open PRs. Every PR has a literal `hold` level-gate label — it will NOT merge until you remove the hold. No other agent opens PRs at L3. | +| **L4** | Quality, ci-maintainer, **and** sec-check can open PRs (all with literal `hold` level-gate labels). Scanner and guide file issues — not PRs. | +| **L5** | All agents can open PRs, all with literal `hold` level-gate labels. Nothing auto-merges. You batch-review. | +| **L6** | Auto-merge switches on for every active repo, then PRs auto-merge when CI goes green. Owners can toggle individual repos off or back on afterward. Non-outreach PRs have no level hold; outreach PRs are still held for human review. Full automation. | -> **The hold label is your safety net.** At every level below L6, every PR an agent opens is blocked from merging until you remove the `hold` label. You are always in control. Nothing ships without your approval until you reach L6 — and you'll only reach L6 after months of trusting the system. +> **The hold label is your safety net.** Hive uses literal `hold` as the level-gate merge-blocking PR label; `hive-pause/` is the dashboard's manual hold label and `hive/` is provenance only. At every level below L6, every PR an agent opens is blocked from merging until you remove `hold`. You are always in control. Nothing ships without your approval until you reach L6 — and you'll only reach L6 after months of trusting the system. > > One exception, so it doesn't surprise you: when you **raise the level**, the hive releases the level holds *it* placed on its own open PRs that the new level no longer requires. It never removes a hold you applied yourself, and never one you re-applied after the hive removed it — those stay put until you lift them. @@ -262,15 +268,15 @@ See [sandbox-isolation.md](https://github.com/hivecommons/hive/blob/v5/src/docs/ > 💡 **Tip: scanner finds, quality fixes.** Scanner flags bugs. Quality fixes test gaps. They're a team. At L4, watch for scanner filing an issue and quality filing a PR that addresses it — that's the closed-loop feedback working. -**Shoring up security:** Sec-check's first run will probably find things. Don't panic. Read each finding, fix the critical ones yourself, and let sec-check open PRs for the medium ones — they'll have hold labels, so you approve before anything merges. +**Shoring up security:** Sec-check's first run will probably find things. Don't panic. Read each finding, fix the critical ones yourself, and let sec-check open PRs for the medium ones — they'll have literal `hold` labels, so you approve before anything merges. -**Be patient:** The first sec-check run can take a full cadence cycle to appear. And yes — you'll get more issues and PRs at this level. Still review them one by one. The hold label exists precisely so nothing merges without you. +**Be patient:** The first sec-check run can take a full cadence cycle to appear. And yes — you'll get more issues and PRs at this level. Still review them one by one. The literal `hold` label exists precisely so nothing merges without you. -**When to move up:** **4–5 weeks.** Let sec-check find and fix security issues. Watch the pattern of what agents propose. Trust is earned slowly — move up when you're approving most agent PRs without changes. +**When to move up:** **4–5 weeks.** Let sec-check find and fix security issues. Watch the pattern of what agents propose. Trust is earned slowly — move up when you're approving most agent change requests without changes. ## L5 — Propose and Review -**The level:** You're trusting *every* agent to open issues and PRs — the system proposes, you decide. Every PR still has a hold label. +**The level:** You're trusting *every* agent to open issues and PRs — the system proposes, you decide. Every agent PR still has a literal `hold` level-gate label. **What you get:** The full hive works for you. Architect produces RFCs for bigger design changes. You shift from doing the work to batch-reviewing it. @@ -283,17 +289,17 @@ See [sandbox-isolation.md](https://github.com/hivecommons/hive/blob/v5/src/docs/ **Using the findings:** Batch-review on a schedule (say, twice a week). Approve the PRs you like, decline the ones you don't, 👍 the issues that match your roadmap. -> 💡 **Tip: batch-review in one sitting.** Reviewing ten agent PRs in a single hour teaches you the agents' patterns faster than reviewing one per day. Patterns jump out when the PRs sit side by side — repeated habits, favorite files, blind spots. +> 💡 **Tip: batch-review in one sitting.** Reviewing ten agent change requests in a single hour teaches you the agents' patterns faster than reviewing one per day. Patterns jump out when the PRs sit side by side — repeated habits, favorite files, blind spots. **Building tests:** By now, quality should have already added tests for your main flows. If it hasn't, go back to L3 habits before moving on — L6 depends on it. **Be patient:** With everything un-paused, the dashboard gets busy. Give new agents a heartbeat cycle before judging their output. -**When to move up:** Only when you **genuinely trust the agents' judgment** — meaning you've reviewed enough of their PRs to know they're consistently doing the right thing, and your test suite is strong enough that green CI genuinely means "safe to ship." There's no calendar for this one. +**When to move up:** Only when you **genuinely trust the agents' judgment** — meaning you've reviewed enough of their PRs to know they're consistently doing the right thing, and your test suite is strong enough that passing checks genuinely means "safe to ship." There's no calendar for this one. ## L6 — Full Automation -**The level:** Full trust. Agents open PRs and merge them automatically when CI goes green. No hold label. +**The level:** Full trust. Auto-merge is off below L6; when you switch to L6, Hive enables auto-merge for every active repo, then agents open PRs and merge them automatically when CI goes green. Owners can toggle individual repos off or back on afterward. Non-outreach PRs have no level hold; outreach PRs remain held for human review. **What you get:** A repo that improves itself while you sleep. The tests quality built at L3 are now the guardrails that keep agents honest. @@ -302,7 +308,7 @@ See [sandbox-isolation.md](https://github.com/hivecommons/hive/blob/v5/src/docs/ ⚠️ **Cadence check:** You can shorten cadences now if your token budget allows — but 12h/1d still works fine. Faster isn't better if you're not reading the output. -Telemetry and operations don't auto-enable just because you reached L6 — they carry the same opt-in requirement here as at L5 (**Settings → Project Observability**). If you enabled them at L5, they stay on and switch to full mode (auto-merge on green CI) like the rest of your roster. +Telemetry and operations don't auto-enable just because you reached L6 — they carry the same opt-in requirement here as at L5 (**Settings → Project Observability**). If you enabled them at L5, they stay on and switch to full mode (auto-merge when checks pass) like the rest of your roster. **Using the findings:** Spot-check merged PRs weekly. 👍 issues to steer agent priorities. @@ -328,7 +334,7 @@ Telemetry and operations don't auto-enable just because you reached L6 — they ## And after that? -**Weeks 2–3** — Stay at L2. When the agents' findings match what you'd find yourself, open the **Governor config** and set the level to **3**. Now quality can open PRs (with hold labels). Review and merge the ones you like. +**Weeks 2–3** — Stay at L2. When the agents' findings match what you'd find yourself, open the **Governor config** and set the level to **3**. Now quality can open PRs (with literal `hold` labels). Review and merge the ones you like. **Weeks 4–7** — Live at L3 while quality builds your test suite. Then L4 for a month or so while sec-check hardens things. L5 and L6 come when trust is genuinely earned — *if* you ever want them at all. L4 or L5 forever is a perfectly good place to live. diff --git a/docs/content/hive/hivectl.md b/docs/content/hive/hivectl.md index e702d4c..f1d4516 100644 --- a/docs/content/hive/hivectl.md +++ b/docs/content/hive/hivectl.md @@ -127,6 +127,7 @@ hivectl agent prompt get quality --raw hivectl agent prompt set quality --file quality.md hivectl agent model-set quality claude-sonnet-4-6 hivectl agent backend-set quality claude +hivectl agent jev-mode-set scanner assist # Jev typed-decision tool: off | assist hivectl agent pipeline-set quality --file pipeline.yaml # map of step: bool ``` @@ -545,6 +546,9 @@ implementation of the profiles file here. |---|---| | `j` / `k`, `↓` / `↑` | Move the cursor | | `enter` | Make the selected hive active — the same effect as `hivectl hives use` | +| `[` / `]` | Move the selected hive up or down in The Commons rank | +| `s` | Cycle the relay routing strategy: `ranked` → `spread` → `neediest` | +| `e` | Enable/disable new work for the selected hive, retaining its credentials | | `a` | Add a hive: a two-field form (name, hub URL), then the registration POST | | `d` | Remove the selected hive — asks you to type its name, as the CLI does | | `r` | Rename the selected hive | @@ -560,12 +564,12 @@ Opening the overlay migrates a legacy positional `contributor.env` exactly as the first `hivectl hives` command would, and leaves that file byte-identical until something actually changes. Every mutation then writes `profiles.yml` first and regenerates `contributor.env` from it, so the relay keeps reading the -variables it always has. **`enter` can move a running relay without a restart** -— like the CLI, it reorders the projection and then signals the relay recorded in -`contributor-relay.pid` (or its docker/podman container) to reload it. Work -already in flight stays with the hub that assigned it; the next solicitation -goes to the newly active hive. If no relay is running, the next relay start -uses the new active hive first. +variables it always has, plus `HIVE_COMMONS_STRATEGY`. **`enter`, `[`/`]` and +`s` can move a running relay without a restart** — like the CLI, they rewrite +the projection and then signal the relay recorded in `contributor-relay.pid` +(or its docker/podman container) to reload it. Work already in flight stays with +the hub that assigned it; the next solicitation uses the new rank/strategy. If +no relay is running, the next relay start uses the saved Commons settings. Two safety properties carry over from the CLI, unchanged: @@ -676,10 +680,17 @@ repository PAT or long-lived secret. hivectl hives list # active hive marked * hivectl hives list --check -o json # also probe each hub hivectl hives add acme --hub wss://acme.hive.hivecommons.dev/contribute +hivectl hives reissue acme # rotate one saved profile +hivectl hives reissue acme --hub wss://acme.hive.hivecommons.dev/contribute hivectl hives use acme hivectl hives export acme --out acme.hive-profile hivectl hives import acme.hive-profile --name acme-laptop hivectl hives session acme --label review +hivectl hives move acme up # rank order +hivectl hives strategy spread # ranked|spread|neediest +hivectl hives disable acme # pause new work, keep credentials +hivectl hives enable acme # resume without registering again +hivectl hives web # loopback web UI for The Commons hivectl hives rename acme acme-prod hivectl hives remove acme # confirm, or --yes ``` @@ -692,6 +703,7 @@ it reads and writes your own contributor credentials, so `--server` and ```yaml version: 1 active: acme +commons_strategy: ranked profiles: - name: acme hub: wss://acme.hive.hivecommons.dev/contribute @@ -704,9 +716,11 @@ profiles: `~/.config/hive/contributor.env` becomes a **generated projection** of that file. The relay, `src/compose-contributor.yaml` and `just contribute-k8s` keep reading the same `HIVE_HUB` / `HIVE_REGISTRATION_TOKEN` / `CONTRIBUTOR_ID` -lists they always have; the active hive is written first in each, which is the -hub the relay solicits from at startup (`activeHubIndex` 0 in -`bin/contributor-relay.js`). Keys the projection does not own — +lists they always have; the active/top-ranked hive is written first in each, +which is the hub the relay solicits from at startup (`activeHubIndex` 0 in +`bin/contributor-relay.js`). `HIVE_COMMONS_STRATEGY` is written alongside the +lists so the long-running relay can choose its next hive between tasks. Keys the +projection does not own — `CONTRIBUTOR_USERNAME`, `AGENT_BACKEND`, `HIVE_LITELLM_ENDPOINT` — are carried across rather than dropped, and the previous file is kept at `contributor.env.bak`. @@ -725,14 +739,45 @@ Notes: - **The same list is in the TUI.** `just contribute-tui` (or `hivectl tui --hives`) opens the [Hives overlay](#hives-switching-the-hive-you-contribute-to): the same rows, - with `enter` to switch and `a`/`d`/`r` to add, remove and rename. It calls - these same functions, so either surface leaves the files in the same state. -- **`use` switches a running relay.** It reorders the projection, then signals - the relay advertised in `contributor-relay.pid` (or its recorded - docker/podman container) to reload `contributor.env`. The task currently in - flight finishes on the hub that assigned it; the next solicitation goes to - the newly active hive. If no relay is running, the next relay start uses the - new active hive first. + with `enter` to switch, `[`/`]` to rank, `s` to cycle the strategy, `e` to enable/disable and + `a`/`d`/`r` to add, remove and rename. It calls these same functions, so + either surface leaves the files in the same state. +- **Disable instead of removing a hive to pause it.** `hivectl hives disable ` + (or `just contribute-hives disable `) preserves its URL, registration token, + contributor ID and rank. `enable ` restores eligibility without registration + or token rotation. Existing profiles default to enabled. Disabled profiles remain + visible in the CLI/TUI but are excluded from new work under **every** strategy; + selecting one as active does not re-enable it. Disable every other hive to focus + exclusively on one, even when it has no jobs. Changes signal the running relay + to reload immediately, without interrupting work on its original hive. With all + profiles disabled, the relay reports why it is idle and waits for an enable. + The generated `HIVE_HUB_DISABLED` boolean list follows the same profile ordering + as the hub/token/contributor-ID lists; disabled entries are not removed from + any credential list. Upgrade the relay together with `hivectl`: older relays + do not understand this local routing control. +- **The Commons strategies choose only between tasks.** `ranked` (the default) + asks the highest-ranked hive first and falls through only when it reports no + work; `spread` uses rank-weighted rotation with occasional mixing so lower + ranked hives still receive some of your daily allotment; `neediest` polls + each hub's `/api/contribute/status` and prefers queued actionable work, + lightly boosted by idle contributor capacity when the hub reports + `active_contributors` and `total_registered`. The existing contributor quota + guard still runs before an offered task is accepted, so a local + daily/subscription cap can hold the relay regardless of which hive The + Commons picked. +- **The Commons web UI is local-only.** `hivectl hives web` prints a + tokenized `http://127.0.0.1:/` URL, serves only on loopback, and backs + the page with the same `profiles.yml` / generated `contributor.env` store as + the CLI and TUI. The page can subscribe (register) a hive, unsubscribe, drag + rows or use ↑/↓ to rank them, and choose `ranked`, `spread` or `neediest`. + The API and HTML never return registration tokens; the printed URL carries a + short-lived local UI token and should not be shared. +- **`use`, `move` and `strategy` switch a running relay.** They regenerate the + projection, then signal the relay advertised in `contributor-relay.pid` (or + its recorded docker/podman container) to reload `contributor.env`. The task + currently in flight finishes on the hub that assigned it; the next + solicitation goes to the newly selected hive. If no relay is running, the next + relay start uses the saved Commons settings. - **Last seen is local relay state.** `hives list` reads `~/.config/hive/hubs-seen.json`, written by the relay after `auth_ok` and successful heartbeats at most once a minute. A `-` means this machine has not @@ -742,10 +787,23 @@ Notes: `gh` login or the backend CLI preflight, so a first-time machine still wants `just contribute-setup `. No bearer credential is sent to the hub. - **Already registered elsewhere?** The register endpoint is unauthenticated, - so the hub will never hand an existing contributor's token back. Add the - credential you already hold instead: + so the hub will never hand an existing contributor's token back. On a + profile-store machine, rotate and save only that one profile with + `hivectl hives reissue acme --hub ` (or omit `--hub` when `acme` + already exists locally). If you want to keep the old machine working, use + `hivectl hives export` on the holding machine and `hivectl hives import` on + the new one. Add the credential you already hold instead only when you have + the token and contributor id in hand: `printf '%s' "$TOKEN" | hivectl hives add acme --hub --token-stdin --contributor-id `, - or move the identity with `just contribute-move`. + `just contribute-move` is unsafe with profiles because it rewrites + `contributor.env` behind `profiles.yml`; the next profile-store command + regenerates `contributor.env` from `profiles.yml`. +- **`reissue` rotates exactly one hive token.** It calls the hub's + authenticated `/api/contribute/reissue-token` endpoint using your local `gh` + token, replaces only the named profile's registration token in + `profiles.yml`, and regenerates `contributor.env` through the profile store. + If the hub says this GitHub account is not registered, run + `hivectl hives add --hub ` instead. - **Moving a profile to another machine:** `hivectl hives export --out ` writes one passphrase-encrypted bundle. On the other machine, `hivectl hives import [--name ]` decrypts, validates and diff --git a/docs/content/hive/integration-guide.md b/docs/content/hive/integration-guide.md index ef22b03..4a9bd94 100644 --- a/docs/content/hive/integration-guide.md +++ b/docs/content/hive/integration-guide.md @@ -12,16 +12,19 @@ flowchart LR Governor --> Agents[Hive agents and contributor relay] Clanker[ClankeR contributor relay\n/api/contribute/ws] --> Agents Flue[Flue external workflow\nextwork adapter] --> Clanker + Crustify[Crustify / Wavefront\nmigration graph] --> WorkSource Spek[Spektacular CLI\nspec/plan status + plan export] --> Runs[Long-running run leases] Runs --> WorkSource Runs --> Governor ``` +Two external projects shaped these surfaces and remain the reference integrations: [Flue](https://github.com/withastro/flue), the report-only external-execution pilot behind the `pkg/extwork` contract, and [Crustify](https://github.com/crustify-rs/crustify), the C/C++-to-Rust migration harness whose Wavefront migration graph is consumed as an additive work source (see [work sources](/docs/hive/work-sources) and the `wavefront-smoke.yml` canary). + ## Extension surfaces in v5 | Surface | What you can do today | Start here | | --- | --- | --- | -| Work sources | Add or configure an adapter that turns source-native items into `worksource.Issue` values. The only primary adapters linked today are GitHub Issues, GitHub Projects, Linear, and Jira; run stages and Wavefront are additive sources. | [Work source providers](/docs/hive/integrations/work-source-providers) | +| Work sources | Add or configure an adapter that turns source-native items into `worksource.Issue` values. The only primary adapters linked today are GitHub Issues, GitHub Projects, Linear, and Jira; run stages and the [Crustify](https://github.com/crustify-rs/crustify) Wavefront migration graph are additive sources. | [Work source providers](/docs/hive/integrations/work-source-providers) | | ClankeR + Flue-style external execution | Use the contributor relay as the transport and the `pkg/extwork` contract as the engine-neutral admission/observation seam. Flue is the reference HTTP adapter. | [ClankeR and Flue-style external execution](/docs/hive/integrations/clanker-flue) | | Spektacular | Let Hive poll a Spektacular-compatible CLI for `spec`/`plan` status and import final plan tasks into Hive's run flow. | [Spektacular and Project Inception](/docs/hive/integrations/spektacular) | diff --git a/docs/content/hive/integrations/work-source-providers.md b/docs/content/hive/integrations/work-source-providers.md index 2adda2f..2ffce69 100644 --- a/docs/content/hive/integrations/work-source-providers.md +++ b/docs/content/hive/integrations/work-source-providers.md @@ -94,7 +94,7 @@ The Linear adapter is the best reference for a non-GitHub work source. ## Dashboard behavior -Work-source configuration appears in Settings -> Work Source. The UI labels GitHub Issues as the default and lists GitHub Projects v2, Linear, and Jira as alternate sources (`src/pkg/dashboard/static/index.html:31330`). The Projects navigation label is intentionally neutral (`src/pkg/dashboard/static/index.html:3912`). Overview bands render source-neutral open/held item counts, while Change Throughput describes merged change requests "across tracked forges" (`src/pkg/dashboard/static/index.html:4204`). +Work-source configuration appears in Settings -> Work Source. The UI labels GitHub Issues as the default and lists GitHub Projects v2, Linear, and Jira as alternate sources (`src/pkg/dashboard/static/index.html:31330`). The Projects navigation label is intentionally neutral (`src/pkg/dashboard/static/index.html:3912`). Overview bands render source-neutral open/held item counts, while Throughput describes merged change requests "across tracked forges" (`src/pkg/dashboard/static/index.html:4204`). ## Testing diff --git a/docs/content/hive/landscape.md b/docs/content/hive/landscape.md index e376370..ade9ebe 100644 --- a/docs/content/hive/landscape.md +++ b/docs/content/hive/landscape.md @@ -12,6 +12,33 @@ after model judgment, and exposes live dashboard, cost, hub/spoke, and contributor-compute surfaces. This page positions that design against nearby agentic orchestration tools. +## Agentic AI Foundation (AAIF) and Goose + +The Linux Foundation's [Agentic AI Foundation](https://aaif.io/) hosts +[Goose](https://github.com/aaif-goose/goose), originally developed at Block. +Goose is an execution backend; Hive is a downstream orchestration layer that +dispatches to Goose at fleet scale, with operator-controlled policy and +contributor-compute relays. Hive integrates Goose as an independent CNCF +project; Hive does not join AAIF, seek AAIF membership, rely on AAIF to host or +govern Hive, or claim AAIF endorsement or compliance. + +This page tracks Goose's upstream move so Hive's backend documentation, +Dockerfiles, and pin-bump tooling follow the correct release source. It does not +request an AAIF landscape listing or describe Hive as part of the foundation's +membership, governance, or project set. + +Review material: [OpenSSF Best Practices badge](https://www.bestpractices.dev/projects/14261), +[Apache-2.0 license](https://github.com/hivecommons/hive/blob/v5/LICENSE), [security self-assessment](https://github.com/hivecommons/hive/blob/v5/src/docs/security-self-assessment.md), +and [CNCF reference architecture](https://github.com/hivecommons/hive/blob/v5/src/docs/cncf-reference-architecture.md). + +The [unattended Goose integration guide](https://github.com/hivecommons/hive/blob/v5/docs/goose-at-scale.md) is +maintained here for now. Goose's [contribution workflow](https://github.com/aaif-goose/goose/blob/main/CONTRIBUTING.md#from-issue-to-pull-request) +requires a **Ready** issue on its board before any external docs PR, and asks +humans to write new issues themselves. No open Hive issue was found there +when checked for #10627, so no upstream PR has been opened. A human sponsor +must file the documentation proposal and obtain Ready status before the +page can be proposed upstream; the local guide is not upstream approval. + ## Fullsend Public references: [fullsend.sh](https://fullsend.sh), @@ -21,7 +48,7 @@ Public references: [fullsend.sh](https://fullsend.sh), [Fullsend runtimes](https://github.com/fullsend-ai/fullsend/blob/main/docs/runtimes.md), [Fullsend intent representation](https://github.com/fullsend-ai/fullsend/blob/main/docs/problems/intent-representation.md). -Fullsend is the most directly comparable open-source project. Its public README +Fullsend is a closely comparable open-source project. Its public README positions it as autonomous agentic software development for Git-hosted organizations, including GitHub, GitLab, and Forgejo. Its docs emphasize a repo-visible coordination model: target repositories carry `.fullsend/` @@ -54,6 +81,78 @@ infrastructure. Hive is a better fit when operators need live fleet visibility, multiple runtimes, graduated autonomy, hub-managed spokes, contributor compute, or network-level enforcement independent of agent runtime hooks. +## OpenAI Symphony + +Reviewed against [OpenAI Symphony](https://github.com/openai/symphony) at +[`be10a1b`](https://github.com/openai/symphony/tree/be10a1b79df723d6d7612b5651c8522704dafb2e). +Public references: [README](https://github.com/openai/symphony/blob/be10a1b79df723d6d7612b5651c8522704dafb2e/README.md), +[Draft v1 SPEC.md](https://github.com/openai/symphony/blob/be10a1b79df723d6d7612b5651c8522704dafb2e/SPEC.md). +Symphony is Apache-2.0 licensed and describes its Elixir reference implementation +as an engineering preview for trusted environments. Its demo monitors Linear, +runs isolated coding agents, and presents proof of work (CI status, review +feedback, complexity analysis, and walkthrough videos) before accepted PRs land. +Those demo artifacts and auto-landing are **not** mandatory spec requirements: +a successful spec run can stop at a human-review handoff. + +Hive applies the Symphony operating model—manage work rather than supervise +every coding turn—to GitHub/GitLab-native projects, with deterministic policy +gating in front of the agent. This is a positioning analogy, **not** a claim +that Hive implements Symphony's `WORKFLOW.md` or Codex app-server contracts. + +| Dimension | Symphony | Hive | +| --- | --- | --- | +| Work source | Linear in the demo; the current spec defines a provider-neutral tracker adapter. | GitHub/GitLab forge workflows; primary planning adapters include GitHub Issues/Projects, Linear, and Jira. | +| Pre-agent gating | Config preflight, active/terminal states, required labels, adapter dispatchability, claims, and concurrency checks (spec §§6–8). | Deterministic shell enumeration/classification/merge eligibility before a kick, plus ACMM and network-level write enforcement. | +| Agents | Codex app-server is the specified runner protocol. | Multiple CLI backends, including Claude, Copilot, Gemini, Goose, and Codex; confinement depends on backend and deployment. | +| Scale model | Per-issue persistent workspace, bounded concurrent runs, reconciliation, continuation, and retry. | Queue-depth-driven cadence, long-lived agents, isolated execution paths, hub/spoke, and convergence audits. | +| Packaging | Language-neutral draft specification and experimental Elixir reference implementation. | Go runtime and supporting scripts, Compose/Quadlet deployment, dashboard, and contributor relay. | + +Choose Symphony when a repo-owned `WORKFLOW.md`, the Codex app-server contract, +and per-ticket workspace lifecycle are the desired integration surface. Choose +Hive when fleet operations, multiple runtimes, forge-native review/merge policy, +and graduated autonomy are central. Neither tool's workspace isolation alone +constitutes a sandbox. + +See the [section-by-section alignment review](https://github.com/hivecommons/hive/blob/v5/docs/design/symphony-spec-alignment.md) +for gaps and intentional divergences, the existing +[Linear work source](/docs/hive/work-sources), and +[work-source provider contract](/docs/hive/integrations/work-source-providers). + +## GitHub Agentic Workflows (gh-aw) + +Public references: [github/gh-aw](https://github.com/github/gh-aw), +[gh-aw documentation](https://github.github.com/gh-aw/), and the +[githubnext/agentics sample gallery](https://github.com/githubnext/agentics). + +GitHub Agentic Workflows compiles Markdown-authored agent instructions and +frontmatter into GitHub Actions workflows. It is an Actions-native on-ramp, +not a replacement for Hive's fleet scheduler. Its engines include Copilot, +Claude, Codex, and Gemini; shared engine names do not imply shared credentials, +confinement, or Hive backend-tier acceptance. + +| Dimension | gh-aw | Hive | +| --- | --- | --- | +| Execution substrate | Event, dispatch, and scheduled GitHub Actions runs | Long-running fleet with queue-depth cadence and convergence audits | +| Authoring | Markdown prompt plus workflow frontmatter compiled to Actions | Project config, deterministic pipeline, and agent policies | +| Judgment and policy | Engine judgment with declared tools, permissions, and safe outputs | Deterministic filtering/classification before judgment and gated merge authority afterward | +| Operations | Per-workflow Actions logs and artifacts | Live dashboard, budgets, hub/spoke, and contributor compute | +| Best starting point | A repo already using Actions that wants bounded agentic jobs | Operators coordinating continuous work across agents and repositories | + +Our [Hive workflow sample and installation guide](https://github.com/hivecommons/hive/blob/v5/src/deploy/gh-aw/README.md) +provides manually dispatched **report-only issue triage**. A deterministic +pre-agent admission step filters held/blocked issues, then runs Hive's existing +classifier before the configured engine produces an advisory report. It does +not give the engine merge authority or reproduce the whole production pipeline. +Keep existing gh-aw workflows when adopting Hive: add the relay/fleet for the +queues and stages requiring continuous operation, rather than rewriting those +workflows or letting both systems claim the same tasks. + +The inverse path (Hive dispatching gh-aw as an external host through +`pkg/extwork`) is **not implemented**. A future adapter is bounded to report-only +and shadow modes, with the same credential, capability, and verified-receipt +admission bar as external OMP; see +[backend support tiers](https://github.com/hivecommons/hive/blob/v5/src/docs/backend-support-tiers.md#github-agentic-workflows-on-ramp-not-a-cli-backend). + ## Single-agent and service-oriented tools ### GitHub Copilot coding agent @@ -84,9 +183,13 @@ multi-agent fleet operation. ## When to choose what +- Choose **gh-aw** when you want Markdown-authored, bounded agentic Actions + jobs in an existing repository without running a standing fleet. - Choose **GitHub Copilot coding agent** when you need the quickest hosted path for GitHub issues and do not need a separate fleet governor or custom policy plane. +- Choose **Symphony** when you want a spec-first, repo-owned `WORKFLOW.md` + and Codex app-server orchestration with per-ticket persistent workspaces. - Choose **Fullsend-style tooling** when you have a small number of repos, want GitHub Actions or native CI to be the execution substrate, prefer repo-visible `.fullsend/` configuration, and want little or no always-on infrastructure. diff --git a/docs/content/hive/manual-provisioning.md b/docs/content/hive/manual-provisioning.md index 4dca59c..a2405f5 100644 --- a/docs/content/hive/manual-provisioning.md +++ b/docs/content/hive/manual-provisioning.md @@ -110,6 +110,11 @@ model endpoint out of the box. See the overlay and its README here: [github.com/hivecommons/hive/tree/v4/src/deploy/kustomize/overlays/standalone/example-joe-spyre](https://github.com/hivecommons/hive/tree/v4/src/deploy/kustomize/overlays/standalone/example-joe-spyre). Every value there is a placeholder — copy the shape, don't apply it verbatim. +Both flows need the Gateway API Inference Extension CRDs on the cluster first. +The inference base ships an `InferencePool` (`inference.networking.k8s.io/v1`), +and without the CRDs `kubectl apply -k` creates everything else and then fails +with `no matches for kind "InferencePool"`. + There are two flows. Pick based on whether you just want to *see it run* or you're doing a *real* deployment. @@ -137,7 +142,7 @@ git clone -b v5 https://github.com/hivecommons/hive.git cd hive/src/deploy/kustomize/overlays/standalone # Edit the placeholders — see "What you must swap" below: # patch-configmap.yaml (org/repos, owner login, OAuth client id, litellm endpoint) -# patch-pvc-storageclass.yaml (your storage class — the PVC is RWO; RWX not needed) +# patch-pvc-storageclass.yaml (your storage class — RWO is fine, but network-attached, not local-path) # Review the rendered manifests, then apply: kubectl kustomize . kubectl apply -k . @@ -173,9 +178,40 @@ kubectl apply -k . > (200 = exists). Digests always work. **Storage.** The base `hive-data` PVC is `ReadWriteOnce`. That is correct for -the single-replica standalone hive — any block storage class works; you do -**not** need an RWX class. (RWX is only relevant to the hub-provisioned path -further down, whose template runs a surge rollout.) +the single-replica standalone hive — any **network-attached** block storage +class works; you do **not** need an RWX class. (RWX is only relevant to the +hub-provisioned path further down, whose template runs a surge rollout.) What +you must avoid on a multi-node cluster is a **node-local** provisioner — see +the next section. + +#### Multi-node clusters: do not put `/data` on node-local storage + +`/data` is the hive's whole state (agent homes, config overlay, audit log, +knowledge). Where it lives decides which nodes the pod can run on: + +| Class type | Examples | Effect on the spoke | +|---|---|---| +| **Node-local** (`WaitForFirstConsumer` + hostPath) | `local-path` (Talos, k3s default), `openebs-hostpath`, `hostpath` | The PV is a directory on **one node's disk**. The pod is pinned to that node forever, shares the disk with images and every other local-path PV, and the claim's size is only a label — `local-path` caps **cannot be expanded** by editing the PVC. When that node crosses kubelet's eviction threshold, every pod on it is SIGKILLed on start (exit 137) and the ReplicaSet recreates it in a loop. Observed on a 3-node Talos/AWS spoke ([#9868](https://github.com/hivecommons/hive/issues/9868)): 37 GB of local-path PVs on one node, `hive-data` at 88 % of its cap, no way to move it. | +| **Network-attached RWO** | AWS `ebs-csi` (`gp3`), AKS `managed-csi`, GKE `pd-balanced`, Ceph `rbd`, Longhorn | The volume follows the pod to any node in its zone, the data is off the node disk, and `allowVolumeExpansion: true` classes grow in place. **Use this for a production spoke.** | +| **RWX** | cephfs, NFS CSI, EFS, Azure Files | Same as above, plus multi-attach — needed only by the hub-provisioned surge rollout; optional here. | + +Check before applying: + +```bash +kubectl get storageclass +# A production spoke wants a network-attached class, ideally with allowVolumeExpansion: true. +# If the only class is local-path / openebs-hostpath, install a CSI driver for your cloud +# first — or accept that the hive is pinned to one node and size that node's disk for it. +``` + +If a spoke is already on `local-path` and must move: scale the deployment to +0, create a new PVC on the network class, copy with a one-off pod that mounts +both (`kubectl run … --overrides` or a small Job running `cp -a /old/. /new/`), +point the Deployment's `hive-data` volume at the new claim, scale back up, and +delete the old PVC once the dashboard shows the agents and config intact. The +hive does not yet split its stateless API from its stateful agent workers, so +one pod still carries both; keeping that one pod off node-local storage is +what lets the scheduler place it where there is room. **On OpenShift.** The standalone overlay does not create a Route or grant any SCC. Two more pieces supply the OpenShift-only deltas: @@ -411,7 +447,9 @@ To find the right class name: ```bash kubectl get storageclass # Look for one marked (default) or annotated storageclass.kubernetes.io/is-default-class=true -# The base PVC is ReadWriteOnce, so any block class works (AKS managed-csi, EBS, RBD, ...). +# The base PVC is ReadWriteOnce, so any NETWORK-ATTACHED block class works (AKS managed-csi, EBS, RBD, ...). +# Avoid node-local classes (local-path, openebs-hostpath): they pin the pod to one node and cannot +# be expanded — see "Multi-node clusters: do not put /data on node-local storage" above. # On Spyre/ODF clusters, ocs-storagecluster-cephfs (RWX) also works — RWX is allowed, not required. ``` @@ -516,7 +554,7 @@ kubectl kustomize overlays/spyre | less # Create the secret first (Option 1 above), then apply: kubectl apply -k overlays/spyre/ kubectl -n hive rollout status deploy/hive # wait for Ready -kubectl -n hive rollout status deploy/vllm # inference backend +kubectl -n hive-inference rollout status deploy/vllm # inference backend ``` **Changing a config value** — e.g. adding an authorized user or changing the @@ -893,9 +931,17 @@ Remove it and the hub falls back to synthesising `.`, which under the manual path for the full rationale, the YAML, and the ServiceAccount-derivation caveat. +The template also emits a read-only cluster-scoped `hive-node-health-reader-*` +ClusterRole/ClusterRoleBinding so push-reported spokes can send node health in +their outbound heartbeat. It grants `nodes get,list`, `nodes/proxy get`, +`pods list`, and `metrics.k8s.io/nodes list`. If metrics-server is absent or +that last rule is denied, the spoke still reports node count, vCPU, memory and +disk capacity from the core Node API and carries the precise `node_health_error` +reason to the hub. + The template binds `hive-sa` on `RequiresSCC` (OpenShift) clusters and `default` -elsewhere. Namespaces provisioned **before** this Role was added do not have it -and need it applied retroactively. +elsewhere. Namespaces provisioned **before** either reader was added do not have +it and need it applied retroactively. --- @@ -907,8 +953,9 @@ with `kubectl --context `. The full set, in order: 1. Namespace 2. ServiceAccount (`hive-sa`) 3. RBAC — three Roles (`hive-secrets-writer`, `hive-self-upgrade`, - `hive-route-reader`) and four RoleBindings (the three above **plus** - `hive-anyuid`) + `hive-route-reader`), one ClusterRole (`hive-node-health-reader-${NS}`), + four RoleBindings (the three above **plus** `hive-anyuid`), and one + ClusterRoleBinding (`hive-node-health-reader-${NS}`) 4. PVC (`hive-data`, RWX cephfs, 50Gi) 5. ConfigMap (`hive-config`) — the first-boot config **seed** 6. Secret (`hive-secrets`) — dashboard token, GitHub App key, LiteLLM key @@ -1004,6 +1051,30 @@ metadata: { name: hive-anyuid, namespace: ${NS} } roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: system:openshift:scc:anyuid } subjects: - { kind: ServiceAccount, name: hive-sa, namespace: ${NS} } +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRole +metadata: { name: hive-node-health-reader-${NS} } +rules: +- apiGroups: [""] + resources: ["nodes"] + verbs: ["get","list"] +- apiGroups: [""] + resources: ["nodes/proxy"] + verbs: ["get"] +- apiGroups: [""] + resources: ["pods"] + verbs: ["list"] +- apiGroups: ["metrics.k8s.io"] + resources: ["nodes"] + verbs: ["list"] +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRoleBinding +metadata: { name: hive-node-health-reader-${NS} } +roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: hive-node-health-reader-${NS} } +subjects: +- { kind: ServiceAccount, name: hive-sa, namespace: ${NS} } YAML ``` @@ -1787,6 +1858,7 @@ kubectl --context -n hive-hub exec "$HUB_POD" -- \ | Pod `1/1 Running` but hive shows **offline**; spoke logs `hub heartbeat rejected status=401` | No `HIVE_HUB_SECRET` in the deployment env | Add `HIVE_HUB_SECRET` (+ `HIVE_HUB_URL`) env from a working hive; the pod rolls and heartbeats | | Pod `1/1 Running` but hive shows **offline**; no 401, heartbeat just stopped | Heartbeat goroutine died; liveness probe on `/api/health` can't detect it | Point livenessProbe at `/api/livez`; restart to revive now | | Pod restarting repeatedly (`RESTARTS` climbing) while the app looks fine; hub unreachable/firewalled | Old liveness probe failed on stale *heartbeat success* — a connectivity condition a restart can't fix | Redeploy to pick up the attempt-based `/api/livez` (+ `startupProbe`); check `/api/health/deep` → `hub_heartbeat` for the real connectivity state | +| Pod SIGKILLed (exit 137) ~2s after every start, `CrashLoopBackOff`; node reports `DiskPressure` or the `/data` PVC is near its cap | `/data` (or the node disk behind a local-path PVC) filled up; kubelet evicts the pod on every start | Before it gets there, `/api/health/deep` → `data_disk` warns at 75%/85% and fails at 95% (the fail raises a hub alert) and the governor pauses agent kicks at the same 95% threshold; see the [DiskPressure recovery runbook](https://github.com/hivecommons/hive/blob/v5/src/docs/diskpressure-recovery-runbook.md) for what is safe to delete and how to expand | | Pod `CrashLoopBackOff` exit 255, SCC `restricted-v2` | `hive-anyuid` RoleBinding subject points at the wrong namespace | Set `subjects[].namespace` to the hive's own `$NS`, delete the pod | | Pod won't boot: `github.token or github.app_id is required` | `github.app_id` empty in the seed | Set the placeholder sentinel `app_id: 999999999` — exactly that value, see [Placeholder `app_id`](#placeholder-app_id). Any other stand-in number is treated as a real App | | Pod `CrashLoopBackOff` with `failed to init GitHub App auth` / `reading app key ...: no such file`, restarts climbing, hive offline on the hub, rollout stuck with two crashlooping pods | A non-sentinel placeholder `app_id` plus a real `installation_id`, and no private key at `key_file`. Older builds exited before the listener bound, so nothing was visible in the dashboard | Set `app_id` to the real App ID and install the PEM at `key_file` (or set `app_id` to the sentinel `999999999` to park the hive in dashboard-only mode). Patch `/data/hive.yaml.dashboard` and restart. Current builds boot degraded and show the reason in the GitHub App banner instead of crashlooping | diff --git a/docs/content/hive/net-admin-requirement.md b/docs/content/hive/net-admin-requirement.md index e82b260..b2efd02 100644 --- a/docs/content/hive/net-admin-requirement.md +++ b/docs/content/hive/net-admin-requirement.md @@ -129,9 +129,24 @@ Granting `NET_ADMIN` does **not** fix this case — the capability is already there; the kernel simply has nothing to grant access *to*. The durable fix is loading the modules on the node so they survive a node rebuild — on OpenShift/RHCOS, a MachineConfig writing an `/etc/modules-load.d/` drop-in -(for example `/etc/modules-load.d/hive-netfilter.conf` containing `xt_owner` -and `xt_REDIRECT`). Until then, taint or label the node so hive pods are not -scheduled onto it. `HIVE_PROXY_ADVISORY_OK=true` remains the explicit opt-out +(for example `/etc/modules-load.d/hive-netfilter.conf` containing `xt_mark`, +`xt_REDIRECT` and `xt_owner`). A MachineConfig rolling-reboots every node in the pool, +though, which is often unacceptable on a shared or GPU cluster. The no-reboot +alternative is the node-prep DaemonSet shipped at +[`src/deploy/k8s/node-prep/hive-netfilter-modules.yaml`](https://github.com/hivecommons/hive/blob/v5/src/deploy/k8s/node-prep/hive-netfilter-modules.yaml): + +```bash +kubectl apply -f src/deploy/k8s/node-prep/hive-netfilter-modules.yaml +kubectl -n hive-node-prep logs -l app=hive-netfilter-modules --prefix # "loaded xt_REDIRECT" per node +``` + +It runs one tiny privileged pod per worker that `modprobe`s the modules into the +host kernel (via `chroot /host`). It re-checks every five minutes, and because +it restarts with the node, it reloads the modules after a reboot. Apply it once +per cluster, not per hive. Then delete the crashlooping hive pod so it +reschedules. Used on the vllm-d cluster (2026-09-24), where one worker lacked +`xt_REDIRECT` and six lacked `xt_owner`. Until one of the two is in place, +taint or label the node so hive pods are not scheduled onto it. `HIVE_PROXY_ADVISORY_OK=true` remains the explicit opt-out here too: the spoke starts with the gate unenforced and logs a WARN saying agents can bypass the proxy on this node. diff --git a/docs/content/hive/release-channels.md b/docs/content/hive/release-channels.md index e09ab69..3c1f724 100644 --- a/docs/content/hive/release-channels.md +++ b/docs/content/hive/release-channels.md @@ -12,9 +12,16 @@ Hive publishes three **release channels** — moving GHCR image tags an operator > **Promotion policy:** the channels diverge by release line and maturity. Every green merge to **`v5`** retags **`candidate`** (and `:latest`); **`stable`** advances later by digest through the scheduled/manual stable-promotion workflow after the [stable soak and promotion policy](https://github.com/hivecommons/hive/blob/v5/src/docs/stable-soak-policy.md) passes. Merges to **`v6`** retag **`edge`**, so `edge` is an active-development v6 build, not a synonym for `stable`. **`v4`** is a maintenance line: its builds publish only `v4-latest` and short-SHA tags, no channel (#7721 Phase 1). +The hub's release-channel block also shows the stable auto-promotion state. +Hub admins see a play/pause control on the `stable` row: play (the default) +lets the hourly promotion workflow catch `stable` up to `candidate` after the +24-hour soak and maintained-hive smoke evidence; pause records who paused and +when, and the workflow skips without moving tags until resumed. Non-admins see +a read-only badge. + ## How channels are published -Channels are **retags, not rebuilds**. Each release line's `docker.yml` workflow adds fast-moving channels as extra tags in the same `docker buildx imagetools create` call that publishes the branch's `-latest` and immutable short-SHA tags, so a channel always points at an already-built, multi-arch digest. Builds of branch `v5` publish `candidate`; the separate stable-promotion workflow later retags `stable` by candidate digest after the soak gate passes. Builds of branch `v6` publish `edge`. All three images get their line's channels in both published orgs: +Channels are **retags, not rebuilds**. Each release line's `docker.yml` workflow adds fast-moving channels as extra tags in the same `docker buildx imagetools create` call that publishes the branch's `-latest` and immutable short-SHA tags, so a channel always points at an already-built, multi-arch digest. Builds of branch `v5` publish `candidate`; the separate stable-promotion workflow later retags `stable` to the digest for the newest build that crossed the 24-hour line after the soak gate passes. Builds of branch `v6` publish `edge`. All three images get their line's channels in both published orgs: - `ghcr.io/hivecommons/hive` and `ghcr.io/hivecommons/hive` - `ghcr.io/hivecommons/hive-contributor` and `ghcr.io/hivecommons/hive-contributor` @@ -71,7 +78,7 @@ Each row also carries the **commit timestamp** of the build it points at (` Distances are computed hub-side (`pkg/hub/channel_distance.go`) via GitHub's compare API and cached permanently: the distance between two fixed commits is immutable, and a moved channel is a new SHA pair, so entries become unreferenced rather than stale — there is no TTL after which a shown distance could be wrong. -Both reads, like every other hub-originated `api.github.com` call (branch tips, commit messages, workflow runs), are made anonymously unless **`HIVE_HUB_GITHUB_TOKEN`** is set on the hub. The anonymous budget is 60 requests/hour per source IP and the branch poller alone exhausts it once the hub tracks more than a couple of branches; GitHub then answers `403`/`429`, the compares fail, and the distance column and timestamps silently disappear (the hub logs `channel distance: compare failed … HTTP 403`). Set the token (a fine-grained or classic token with public-repo read; 5000 requests/hour) and the rows come back on the next 5-minute channel refresh. See [`env-vars.md`](https://github.com/hivecommons/hive/blob/v5/src/docs/env-vars.md). +Both reads, like every other hub-originated `api.github.com` call (branch tips, commit messages, workflow runs), are made anonymously unless **`HIVE_HUB_GITHUB_TOKEN`** is set on the hub. The anonymous budget is 60 requests/hour per source IP and the branch poller alone exhausts it once the hub tracks more than a couple of branches; GitHub then answers `403`/`429`, the compares fail, and the distance column and timestamps silently disappear (the hub logs `channel distance: compare failed … HTTP 403`). Set the token (a fine-grained or classic token with public-repo read; 5000 requests/hour) and the rows come back on the next 5-minute channel refresh. See [`env-vars.md`](/docs/hive/env-vars). ## Persistence: the tracked channel is durable @@ -99,7 +106,8 @@ For self-hosted Podman Quadlet spokes the selector is intentionally unavailable ## Known limitations - **Bulk actions cannot set a channel.** The bulk *Switch branch* action validates against real branches only and rejects channel names (`unknown branch`); it also never writes the tracked channel. Switching to a channel is per-hive. -- **A manual Upgrade on a channel-tracking hive transiently arms a branch-SHA target.** The manual upgrade handler still targets the tracked *branch*'s latest SHA (`getLatestSHAForBranch`, `pkg/hub/saas.go`); the heartbeat re-arm drags the hive back to the channel tag on the next non-upgrading beat. Expect a short window where the pill says `stable (v4)` while an upgrade converges on a SHA. This is now the exception — automatic targeting resolves through the channel tag (see below). +- **A manual Upgrade on a channel-tracking hive resolves through the channel tag.** If `:stable`/`:candidate`/`:edge` already points at the commit the spoke is running, the hub refuses the click with a visible explanation instead of arming a no-op heartbeat upgrade. Operators who need a newer build must wait for the channel to advance or switch the hive to a newer channel/branch tag. +- **The spoke dashboard offer uses that same target.** `/api/version` keeps reporting the branch tip for provenance, but the `behind` flag, behind count, and Upgrade button are measured against the hub-delivered upgrade policy when present. A `:stable` spoke therefore does not offer "Upgrade available → ``" while `:stable` itself still points at the running commit. ## Channel-aware upgrade targeting @@ -107,7 +115,7 @@ Automatic upgrade targeting resolves **through the tag the spoke's Deployment tr The stable soak policy made this necessary: per-merge publishes move only `candidate`, while `stable` advances later by digest. A spoke's Deployment tracks one image tag, and rolling the pod re-pulls that tag — nothing the hub instructs can make a restart land on a digest the tag does not carry. A hub that targets branch HEAD is therefore asking a `:stable` spoke to reach a digest its own tag is deliberately withholding; before the fix this looped 41 spokes into permanent `UPGRADE FAILED` (hub instructs a SHA, spoke rolls, re-pulls `:stable`, reports the SHA it started on, hub re-sends the identical instruction). -How targets are now resolved (`reachableUpgradeTarget`, used by the auto-upgrade sweep and the heartbeat spoke-managed path): +How targets are now resolved (`reachableUpgradeTarget`, used by the manual upgrade handler, the auto-upgrade sweep, and the heartbeat spoke-managed path): - **Which channel a spoke is on:** the spoke's *reported image ref* leads, because it is what the kubelet will pull; the hub-side `tracked_channel` record is intent, and the two disagree exactly while a channel switch is still on the wire. `tracked_channel` remains the fallback only for spokes too old to report an image ref. Branch tags and SHA pins resolve to branch targeting, exactly as before. - **Channel → commit:** the hub walks GHCR from the channel tag to the image index, picks the `linux/amd64` platform manifest (buildx attaches `unknown/unknown` provenance descriptors to the same index, so position is not enough), and reads the `org.opencontainers.image.revision` OCI label from the config blob — the only commit identity that survives a retag. Answers are cached for 5 minutes (`channelDigestTTL`); when a refresh fails, the last good answer is served for up to 4× that (`channelRevisionStaleGrace`, 20 minutes) since a channel moves at most hourly, then resolution is treated as failed. diff --git a/docs/content/hive/roadmap.md b/docs/content/hive/roadmap.md index 5f0fc0f..971aab0 100644 --- a/docs/content/hive/roadmap.md +++ b/docs/content/hive/roadmap.md @@ -58,7 +58,7 @@ order, not priority rank. | Cross-forge orchestration | Coordinate issues, merge requests, policy, and evidence across GitHub, GitLab, and Forgejo/Gitea-style forges. | [ADR-0005](/docs/hive/adr/0005-forge-abstraction) | | Memory and learning maturation | Turn retro findings and curated knowledge into durable, testable priming without hidden or unauditable agent memory. | [knowledge design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/knowledge-system.md), [retro lane](https://github.com/hivecommons/hive/blob/v5/src/docs/retro-lane.md) | | Kubernetes-native agent sandboxes | Graduate from tmux/container execution toward k8s-native, policy-isolated agent workloads where that complexity is justified. | [architecture](/docs/hive/architecture), [security threat model](/docs/hive/security-threat-model) | -| v6 — dashboard-optional operation (line open) | The `v6` branch opened 2026-09-18 (cut from v5); v6 implementation PRs target the `v6` branch only and v6 tracks v5 by forward-merge. Theme: every operator interaction reachable from the places humans already are — GitHub threads (@-mention triggers), Slack/Discord/Teams/Matrix/Telegram over a shared `pkg/chat` spine, email and push/on-call escalation — all behind the same role/capability/ioscan guards the dashboard uses. **Tracks merged on the `v6` branch:** the `pkg/chat` spine and Discord port ([#7572](https://github.com/hivecommons/hive/pull/7572)) with reconnect/cancellation parity ([#7586](https://github.com/hivecommons/hive/pull/7586)); Slack Socket Mode ([#7585](https://github.com/hivecommons/hive/pull/7585)); Teams, Matrix, and Telegram ([#7621](https://github.com/hivecommons/hive/pull/7621), [#7617](https://github.com/hivecommons/hive/pull/7617), [#7616](https://github.com/hivecommons/hive/pull/7616)); GitHub @-mention triggers phases 1–3 ([#7582](https://github.com/hivecommons/hive/pull/7582), [#7597](https://github.com/hivecommons/hive/pull/7597), [#7623](https://github.com/hivecommons/hive/pull/7623)); outbound email and push escalation ([#7618](https://github.com/hivecommons/hive/pull/7618)). Still unbuilt: the Slack Events API accelerator and inbound reply-to-act email. **This is branch status, not stable-release status** — none of it exists on `v4` or `v5`; since the 2026-09-21 channel re-base ([#7721](https://github.com/hivecommons/hive/issues/7721)), `edge` is built from `v6`, so operators tracking `edge` run these surfaces as active-development builds while `stable`/`candidate` (v5) do not have them. **Beyond the dashboard-optional theme, the line accepts new tracks by RFC**, each measured against the same readiness bar: task-scoped MCP — expose the hive's view of a task to contributor agents as a Model Context Protocol endpoint ([#8033](https://github.com/hivecommons/hive/issues/8033); Phase 1 merged on `v6` via [#8164](https://github.com/hivecommons/hive/pull/8164), [#8177](https://github.com/hivecommons/hive/pull/8177), [#8179](https://github.com/hivecommons/hive/pull/8179) — `pkg/taskmcp`, the `/api/contribute/mcp` endpoint, and the first four tools; Phase 2 lease-scoped auth [#8228](https://github.com/hivecommons/hive/pull/8228) and Phase 3 repo tools [#8244](https://github.com/hivecommons/hive/pull/8244) also merged on `v6`; the RFC closed 2026-09-22 with the first live token-delta measurement, and the remaining saving — eliding stuffed issue/PR lists when the MCP pointer is present — is [#8261](https://github.com/hivecommons/hive/issues/8261)) — and standby contributors — a lane paused for budget hands its queue to volunteer contributors behind a model floor ([#7629](https://github.com/hivecommons/hive/issues/7629); design [#8038](https://github.com/hivecommons/hive/pull/8038); S0–S7 merged on `v6`, S8 waits on live-hive runbook evidence). Promotion bar: [v6-readiness.md](https://github.com/hivecommons/hive/blob/v6/src/docs/v6-readiness.md), live tracker [#7683](https://github.com/hivecommons/hive/issues/7683). Full policy in [ROADMAP.md](https://github.com/hivecommons/hive/blob/v5/ROADMAP.md#v6--dashboard-optional-operation-line-open). | [#7563 epic](https://github.com/hivecommons/hive/issues/7563), [#7483](https://github.com/hivecommons/hive/issues/7483), [mention-triggers design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/github-mention-triggers.md), [Slack design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/slack-integration.md), [escalation design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/escalation-surfaces.md) | +| v6 — dashboard-optional operation (line open) | The `v6` branch opened 2026-09-18 (cut from v5); v6 implementation PRs target the `v6` branch only and v6 tracks v5 by forward-merge. Theme: every operator interaction reachable from the places humans already are — GitHub threads (@-mention triggers), Slack/Discord/Teams/Matrix/Telegram over a shared `pkg/chat` spine, email and push/on-call escalation — all behind the same role/capability/ioscan guards the dashboard uses. **Tracks merged on the `v6` branch:** the `pkg/chat` spine and Discord port ([#7572](https://github.com/hivecommons/hive/pull/7572)) with reconnect/cancellation parity ([#7586](https://github.com/hivecommons/hive/pull/7586)); Slack Socket Mode ([#7585](https://github.com/hivecommons/hive/pull/7585)); Teams, Matrix, and Telegram ([#7621](https://github.com/hivecommons/hive/pull/7621), [#7617](https://github.com/hivecommons/hive/pull/7617), [#7616](https://github.com/hivecommons/hive/pull/7616)); GitHub @-mention triggers phases 1–3 ([#7582](https://github.com/hivecommons/hive/pull/7582), [#7597](https://github.com/hivecommons/hive/pull/7597), [#7623](https://github.com/hivecommons/hive/pull/7623)); outbound email and push escalation ([#7618](https://github.com/hivecommons/hive/pull/7618)). Still unbuilt: the Slack Events API accelerator and inbound reply-to-act email. **This is branch status, not stable-release status** — none of it exists on `v4` or `v5`; since the 2026-09-21 channel re-base ([#7721](https://github.com/hivecommons/hive/issues/7721)), `edge` is built from `v6`, so operators tracking `edge` run these surfaces as active-development builds while `stable`/`candidate` (v5) do not have them. **Beyond the dashboard-optional theme, the line accepts new tracks by RFC**, each measured against the same readiness bar: task-scoped MCP — expose the hive's view of a task to contributor agents as a Model Context Protocol endpoint ([#8033](https://github.com/hivecommons/hive/issues/8033); Phase 1 merged on `v6` via [#8164](https://github.com/hivecommons/hive/pull/8164), [#8177](https://github.com/hivecommons/hive/pull/8177), [#8179](https://github.com/hivecommons/hive/pull/8179) — `pkg/taskmcp`, the `/api/contribute/mcp` endpoint, and the first four tools; Phase 2 lease-scoped auth [#8228](https://github.com/hivecommons/hive/pull/8228) and Phase 3 repo tools [#8244](https://github.com/hivecommons/hive/pull/8244) also merged on `v6`; the RFC closed 2026-09-22 with the first live token-delta measurement, and the remaining saving — eliding stuffed issue/PR lists when the MCP pointer is present — is [#8261](https://github.com/hivecommons/hive/issues/8261)) — and standby contributors — a lane paused for budget hands its queue to volunteer contributors behind a model floor ([#7629](https://github.com/hivecommons/hive/issues/7629); design [#8038](https://github.com/hivecommons/hive/pull/8038); S0–S7 merged on `v6`, S8 waits on live-hive runbook evidence). Promotion bar: [v6-readiness.md](https://github.com/hivecommons/hive/blob/v6/src/docs/v6-readiness.md), live tracker [#7563](https://github.com/hivecommons/hive/issues/7563) §v6 readiness bar. Full policy in [ROADMAP.md](https://github.com/hivecommons/hive/blob/v5/ROADMAP.md#v6--dashboard-optional-operation-line-open). **Known gap, updated 2026-10-02:** the dangerous part of the v6 release-path gap is fixed — `tagged-release.yml`/`promote-stable.yml` now derive `RELEASE_LINE`/`RELEASE_BRANCH` from `github.ref_name` and refuse to run from a branch that is not in `.github/release-lines.yml`'s `release_lines` list ([#9154](https://github.com/hivecommons/hive/issues/9154), closed, fixed by [#9307](https://github.com/hivecommons/hive/pull/9307)). What's still true: `release_lines` remains `[v4, v5]` — `v6` has not been added — and `docker.yml`'s `v6` rows still only ever publish `edge`, never `candidate`/`stable`; this is now a deliberate, guarded GA gate pending the readiness bar, not a bug. The readiness-runbook gap is resolved ([#9200](https://github.com/hivecommons/hive/issues/9200), closed). Tracked together in [#9227](https://github.com/hivecommons/hive/issues/9227), open pending a maintainer go/no-go on flipping the gate. | [#7563 epic](https://github.com/hivecommons/hive/issues/7563), [#7483](https://github.com/hivecommons/hive/issues/7483), [mention-triggers design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/github-mention-triggers.md), [Slack design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/slack-integration.md), [escalation design](https://github.com/hivecommons/hive/blob/v5/src/docs/design/escalation-surfaces.md), [#9227](https://github.com/hivecommons/hive/issues/9227) | ## Reading this roadmap diff --git a/docs/content/hive/running-at-level-6.md b/docs/content/hive/running-at-level-6.md index 9847722..a65e168 100644 --- a/docs/content/hive/running-at-level-6.md +++ b/docs/content/hive/running-at-level-6.md @@ -2,8 +2,6 @@ # Running a hive at ACMM Level 6 -> **Awaiting operator review.** This advice has not yet been reviewed by someone who operates a hive at ACMM Level 6. Maintainer sign-off is tracked in [#10515](https://github.com/hivecommons/hive/issues/10515). Until then, this page is linked only from the documentation index; it is not a signed-off operating standard. - For people who operate a hive, Level 6 means unattended merging, not just faster agents. Protect every destination branch with real required tests and coverage checks; aim for **90% or better coverage before enabling automatic merging**. Stage repositories individually, keep work creation below CI/merge capacity, and prove your emergency stops work before you need them. Human-authored PRs can merge too. A reviewer comment, a spending limit, and paused agents are not substitutes for a merge stop. This guide walks through preparation, a weekly routine, and recovery. Start with the merge-path table, configure the five areas below, then use the single pre-switch checklist. After a bad merge, stop further harm first, revert and verify, then demonstrate that a strengthened required check rejects the original change. @@ -150,6 +148,7 @@ Each row is a separate practice. Except where marked, the recommendations addres | **H13. Keep optional lanes paused until their output is wanted.** Agent controls/Cadences: retain shipped pauses for supervisor, strategist, outreach, telemetry and operations in all modes. | More issue-filers can worsen saturation; L6 does not mean every lane is active. Verify actual pause/cadence state in each mode. Hive enforces pauses; deciding which work is wanted remains yours. See [pack](https://github.com/hivecommons/hive/blob/v5/src/pkg/config/packs/level-6.yaml), [agent pause](https://github.com/hivecommons/hive/blob/v5/src/pkg/agent/manager_pause.go). | | **H14. Deliver alerts to someone who can act.** **JUDGMENT CALL:** Settings → Notifications: configure a documented ntfy/Slack/Discord destination and a responsible operator. | A dashboard-only warning may go unseen while merges continue. Send the documented test notification and verify receipt; Hive sends configured notifications, not an acknowledgement/SLA. See [notifications](https://github.com/hivecommons/hive/blob/v5/src/docs/notifications.md), [notifier](https://github.com/hivecommons/hive/blob/v5/src/pkg/notify/notify.go). | | **H15. Do not mistake review approval for a universal merge gate.** **JUDGMENT CALL:** if choosing `review.require_approval: true`, set it in Governor Config → Security → PR Review (Require review approval) and use GitHub rules or literal holds for changes that require human review. | The setting gates the scanner, not the App sweep. Verify both routes with test PRs; Hive enforces only the path-specific approval check. See [eligibility](https://github.com/hivecommons/hive/blob/v5/src/cmd/hive/merge_eligibility.go), [sweep](https://github.com/hivecommons/hive/blob/v5/src/pkg/github/automerge/automerge_sweep.go). | +| **H16. Turn issue claims on so workers do not collide.** Governor Config → Features → **Issue claims** (`governor.claims.enabled: true`); set `governor.claims.{human_ttl_s,agent_ttl_s,contributor_ttl_s}` to match how long work actually takes. Claim an item yourself with `hivectl claim o/r#N` before you start on it. | **Off by default.** With claims off, two of your own agents, a relay contributor, an outside bot and you can all start the same issue and open competing PRs; the duplicate-PR sweep only sees the collision after the PRs exist. With claims on, the hub records who is working an issue and ranks holders human > agent > contributor > external, so a lower rank backs off and a higher rank takes over with a `preempted:` label. The record is a marker comment (``) or the assignee on the GitHub issue itself, so it needs no organization: a single user's personal repos work, and any other hive or script reading the issue sees the hold. Claims lapse on their own (4h person, 2h agent kick, 30m relay task unless configured). Verify: claim a test issue, confirm `claimed` appears and that an agent kick on that issue is withheld; let the claim lapse or `hivectl unclaim`, confirm the issue is actionable again. Hive enforces the ledger and expiry, not the discipline of claiming before you start. See [claims](https://github.com/hivecommons/hive/blob/v5/src/pkg/claims/claims.go), [issue claims](https://github.com/hivecommons/hive/blob/v5/src/pkg/github/issue_claims.go), [claims config](https://github.com/hivecommons/hive/blob/v5/src/pkg/config/claims.go), [hivectl claim](/docs/hive/hivectl#claim--unclaim--claims--issue-claims). | ### ClankeR settings @@ -213,9 +212,11 @@ These are confirmations of the practices above, not additional product gates. - [ ] Real weeks at L3–L5 produced reviewed PRs matching your judgment; no calendar interval is enforced, and "not surprised" matters more than elapsed time. - [ ] GitHub installation scope/permissions and **every** merge destination's rules checked; App cannot bypass critical tests/coverage. Green CI genuinely means safe to ship, not just a passing harness. - [ ] Each proposed repo meets the 90% coverage target with demonstrated failing gates; below-target repos remain excluded. Advisor readings and their limitations recorded. +- [ ] Per-repository opt-out decided and applied: Level 6 is not all-or-nothing. Each repository card's **Auto-merge** switch is set deliberately (off for repos that are not ready or must stay human-merged; `project.repo_policies[].auto_merge` in config), and a test PR in an opted-out repo was refused while agents kept proposing work there. Remember that re-applying Level 6 re-enables auto-merge for every **active** repo, so opted-out repos are paused during the switch (R2) or re-checked afterwards (R3). - [ ] Reporter-trust posture explicitly decided before promotion. For public issues, trust is on with the intended association set; decide specifically whether `CONTRIBUTOR` belongs. Demonstrate admission and PR-side holding. - [ ] Knowledge priming enabled and actual prompts inspected; release/test/rollback constraints are accurate. Any relied-on checkout instructions demonstrably reach agents. - [ ] Hive repo scope, required contexts, proposal/plan decisions, cadence, token budget, paused lanes and alert delivery reviewed. +- [ ] Issue claims enabled (`governor.claims.enabled`) with TTLs that fit real task length; a claimed test issue was withheld from agents and released on lapse/unclaim. Humans know to `hivectl claim` before starting on an item that agents could also pick up. - [ ] ClankeR filters, model/effort admission, trust/roles, independent queueing and suspension tested. - [ ] Repository labels mapped to Hive gates; acceptance, hold and un-park canaries do not conflict with lifecycle/Prow automation. - [ ] GitHub/work-source notifications narrowed; routine chatter muted without hiding direct mentions or actionable incident alerts. diff --git a/docs/content/hive/securing-your-hive.md b/docs/content/hive/securing-your-hive.md index 15cc903..e84dfb8 100644 --- a/docs/content/hive/securing-your-hive.md +++ b/docs/content/hive/securing-your-hive.md @@ -281,9 +281,14 @@ or PR history. The association alone is enough. **What exactly does `triage/accepted` unlock, and what does it not unlock?** It unlocks **admission**: an agent may now work the issue (open a PR about it, comment, classify it) the same as any other actionable issue. It does -**not** unlock unattended merge. If the resulting PR's rationale traces back -to that untrusted-reporter issue, the reporter-trust merge-side check still -applies `hold` at every level — a human still has to remove that label. +**not** unlock unattended merge. While the issue is waiting, Hive posts one +marked explanation comment and applies `needs-triage` (or the configured +`awaiting_label`) when the issue does not already have it; adding +`triage/accepted` removes that waiting label only if Hive's marker records +that Hive added it, then admits the issue. If the resulting PR's rationale +traces back to that untrusted-reporter issue, the reporter-trust merge-side +check still applies `hold` at every level — a human still has to remove that +label. **How do I stop all auto-merges right now?** Drop below L6 — merge permission is not granted below L6 at the token/proxy diff --git a/docs/content/hive/troubleshooting.md b/docs/content/hive/troubleshooting.md index 313c1bf..8120b74 100644 --- a/docs/content/hive/troubleshooting.md +++ b/docs/content/hive/troubleshooting.md @@ -175,6 +175,7 @@ Two more card behaviors remove the old one-reason-per-click treasure hunt: - **Zero cadence is named.** An enabled, governor-kickable agent with no cadence in *any* mode that has never been kicked shows a `⏱ never scheduled — set cadences` chip; clicking it opens the agent's Cadences tab. This is the per-agent form of the fleet-level "never kicked" banner — both are driven by the same predicate, so they cannot name different agents. - **A cadence in the wrong mode is named too** ([#7474](https://github.com/hivecommons/hive/issues/7474)). An agent that *some* mode schedules but the current one does not — no entry for the current mode, and none in `idle`, which every other mode inherits — is not kicked until the mode changes, and would otherwise read as healthy: enabled, session up, a cadence configured, not on-demand. The card renders it in the same hollow-green "off" state as a governor-paused agent, the `scheduling` segment reads `no cadence in busy (only in surge)`, and the fleet banner `agent(s) reviewer (cadence only in surge) not scheduled in the current busy mode …` names the same agents. The shape to look for is a cadence **only in `surge`**: the agent runs while the backlog holds the fleet in surge, and its own success — driving the backlog below the threshold — removes its schedule. Add an entry for the mode the fleet is in, or an `idle` entry to cover every mode. An explicit `pause`/`off` entry is not this: that is an operator choice, and the card reports it as `off in mode`. +- **A crash-restarted agent sitting at an empty prompt is named too.** When an agent's CLI dies mid-turn, hive relaunches it, but the *resume kick* - the early kick that would let it pick the work back up - goes through the governor gate ([#2573](https://github.com/hivecommons/hive/issues/2573)): at most one resume kick per cadence interval, none while the mode pauses or does not schedule the agent, none for on-demand or time-of-day agents, none while the budget is exhausted. Only the first case (`interval_throttle`: a second crash inside the same interval) leaves an agent that is expected to work idle until its next scheduled slot, so only that case raises the fleet banner `agent scanner is idle at a fresh prompt: its CLI crashed and was restarted, but the resume-kick gate refused an early kick (one resume kick per cadence interval) - kick it from the agent card now, or wait for its next scheduled slot`. The other refusals get no per-agent banner ([#9612](https://github.com/hivecommons/hive/issues/9612)): a paused, unscheduled, on-demand or non-interval agent is idle by configuration and only logs `restarted agent NOT resume-kicked (cadence/budget gate); idle by configuration` (info, with `reason=`), and a budget refusal logs `...; budget exhausted, kicks suspended` and is covered by the fleet-wide `token budget exhausted` banner. The throttle log line is still `restarted agent NOT resume-kicked (cadence/budget gate); it idles until its next scheduled slot`, now with `reason=interval_throttle`. The banner is reconciled every eval cycle and clears as soon as any of these holds: a kick (scheduled or manual) reached the agent; the agent is operator-paused, disabled, or gone from the roster; its session was busy or working again; the gate would now refuse it for a non-throttle reason (for example the mode changed to one that pauses it); or the banner is older than two cadence intervals (6h when the interval is unknown). When the crash coincides with a rise in the container's cgroup `oom_kill` counter (`/sys/fs/cgroup/memory.events` or the v1 `memory/memory.oom_control`), the banner says `OOM-killed (container hit its memory limit)` instead, and the log carries `agent crash coincides with cgroup OOM kills` with the before/after counts: the kernel killed the largest process in the pod — almost always an agent CLI — without restarting the container, so Kubernetes shows no `OOMKilled` and no pod restart. The remedy is the pod memory limit (the hub provisions new spokes at 16Gi for exactly this reason) or fewer concurrent sub-agents, not another kick. - **A paused agent's primary action is always its pause toggle.** With a live session the button is `▶ resume`, and its tooltip names anything resuming will *not* clear. If the session is also down, the one button reads `▶ start & resume` and clears both flags in a single click — client-side chaining of the two existing endpoints (`POST /api/resume/{agent}` first, so the fresh session is never born paused, then `POST /api/restart/{agent}`). No new API surface; scripts can chain the same two calls. A paused agent never offers a bare Start. If the card says the agent should be running (session up, scheduling shows a cadence, next kick has an ETA) and it still misbehaves, *then* drop into the session as described below. @@ -462,6 +463,19 @@ Press **`q`** (or `Esc`) to leave copy-mode and resume following output. If the status bar shows `[live]` and output really has stopped, the agent is idle between kicks — check its next scheduled kick on the dashboard before assuming a fault. +## Copying text out of the browser terminal + +The browser terminal is a live `tmux` attach rendered by ttyd/xterm.js, which means **two** things own the mouse and neither is your browser: + +1. **tmux mouse mode owns an ordinary drag** (it is what makes the scroll wheel page back through history). Hold **⇧ Shift** while dragging to bypass it and make a normal terminal selection. +2. **xterm.js keeps that selection in its own model, not in the page**, so until [#9941](https://github.com/hivecommons/hive/issues/9941) the browser's Copy command had nothing of the pane to copy — ⌘C appeared to work and pasted whatever the clipboard already held, most visibly in Firefox. + +With the selection made, **⌘C** (macOS) or **Ctrl+Shift+C** (Linux/Windows) copies it; the terminal page now answers that gesture itself, as it does the browser's Edit ▸ Copy. A plain **Ctrl+C** is deliberately left alone — it is still SIGINT for the agent's pane. Pasting in is the browser's own: **⌘V** or **Ctrl+Shift+V**. + +**An ordinary drag inside an agent CLI selects in the CLI, not the terminal.** Agent CLIs turn mouse reporting on, so tmux hands a plain drag to the application, which copies its selection with an OSC 52 escape (and may toast a tmux paste hint). The browser terminal now forwards that copy to your clipboard ([#9941](https://github.com/hivecommons/hive/issues/9941)); if your browser refused the write, press ⌘C or Ctrl+Shift+C once on the terminal and the CLI's last copy is put on the clipboard from inside that keystroke. + +**For a login URL, prefer the dashboard's 🔑 *Copy login URL* button** on the agent card. It captures the pane server-side and rejoins wrapped lines, so the OAuth URL arrives whole; a Shift-drag selection of a wrapped URL copies the wrap's newlines with it and the link silently fails at the identity provider. + ## The dashboard says the next kick is later, but the agent is visibly working now The agent-card **last kick** / **next kick** fields describe when work is *started*, not how long it runs. A kick sends one prompt into the agent's CLI; the resulting work pass then runs as long as it needs — often hours for a deep quality or scan pass. So an agent visibly busy at 01:47 with `last kick 8:12 PM` and `next kick 2:12 AM` is not off schedule: it is still working through the pass that began at 20:12. (These fields were labelled "last run" / "next run" before [#4399](https://github.com/hivecommons/hive/issues/4399), which invited exactly this misreading.) diff --git a/docs/content/hive/work-sources.md b/docs/content/hive/work-sources.md index 0a845f7..fba432a 100644 --- a/docs/content/hive/work-sources.md +++ b/docs/content/hive/work-sources.md @@ -48,7 +48,7 @@ currently working is not offered again, and `implement` is listed only once the run's imported plan is approved. The hive binary wires that accessor during dashboard boot; if the accessor is unavailable, the additive source fails closed by listing no run stages. The Spektacular (Spek) stage runner that advances the -lease is described in [spektacular.md](https://github.com/hivecommons/hive/blob/v5/src/docs/spektacular.md). +lease is described in [spektacular.md](/docs/hive/integrations/spektacular). ## Wavefront migration graph (`wavefront.enabled: true`) @@ -80,6 +80,11 @@ is re-read on every governor cycle, so a plan Wavefront republishes is picked up without a restart. `receipts_dir` is optional; empty keeps receipts in memory for the life of the process. +Dashboard owners can configure the same `wavefront` block from **Settings → +Governor → Work Source → Wavefront (Crustify) migration graph** instead of +hand-editing `hive.yaml`; the URL source expects a plain unauthenticated JSON +document, matching the YAML-only configuration. + **Graph document.** A JSON object with a `graph` name, a `revision`, and a `nodes` list. Each node has `id`, `title`, an optional `kind`, an optional `depends_on` list of node ids, and an optional `status` (`pending` (default), @@ -233,8 +238,7 @@ that include the instance context path, such as ([Atlassian Jira Data Center REST API reference](https://docs.atlassian.com/software/jira/docs/api/REST/9.14.0/)). Atlassian's server examples use `/rest/api/2/...` endpoints for issues and searches ([Jira REST API examples](https://developer.atlassian.com/server/jira/platform/jira-rest-api-examples/)). -Data Center PATs are available in Jira Core/Software 8.14+ and are sent as -bearer tokens ([Using Personal Access Tokens](https://confluence.atlassian.com/enterprise/using-personal-access-tokens-1026032365.html)). +Data Center PATs are available in Jira Core/Software 8.14+ and are sent with Jira Data Center's bearer-token HTTP authentication ([Using Personal Access Tokens](https://confluence.atlassian.com/enterprise/using-personal-access-tokens-1026032365.html)). Jira Cloud rich text comments/descriptions use ADF JSON ([Atlassian Document Format](https://developer.atlassian.com/cloud/jira/platform/apis/document/structure/)); Data Center accepts plain text / wiki-markup string bodies. @@ -266,6 +270,10 @@ governor: # Or, for older instances without PATs: # username: hive-bot # password: ${JIRA_DATACENTER_PASSWORD} + # ca_bundle: ${JIRA_DATACENTER_CA_BUNDLE} # optional PEM, appended to system roots + # insecure_skip_verify: false # optional, unsafe; testing only + # client_cert: ${JIRA_DATACENTER_CLIENT_CERT} # optional mTLS PEM + # client_key: ${JIRA_DATACENTER_CLIENT_KEY} # optional mTLS PEM key project_keys: [ENG, OPS] repo: your-org/default-repo hold_labels: [hold, blocked] @@ -281,6 +289,10 @@ Config fields (`JiraSourceConfig`, `pkg/config/config.go`): | `username` | `Username` | DC basic auth only | Jira Data Center username when using basic auth. | | `api_token` | `APIToken` | Cloud yes; DC preferred | Cloud API token (Basic password) or Data Center Personal Access Token (Bearer). | | `password` | `Password` | DC basic auth only | Jira Data Center password, used only when `api_token` is empty. Prefer PATs where supported. | +| `ca_bundle` | `CABundle` | No | Data Center only. PEM CA certificate bundle appended to system trust roots. May be a `${ENV}` reference or pasted PEM; dashboard responses expose only `ca_bundle_set`. | +| `insecure_skip_verify` | `InsecureSkipVerify` | No | Data Center only. Disables server certificate and hostname verification. Default `false`; use only for testing. The dashboard warns loudly and the adapter logs a warning whenever this client is built. | +| `client_cert` | `ClientCert` | No | Data Center only. Optional PEM client certificate for mTLS; set with `client_key`. Dashboard responses expose only `client_cert_set`. | +| `client_key` | `ClientKey` | No | Data Center only. Optional PEM private key for mTLS; set with `client_cert`. Dashboard responses expose only `client_key_set`. | | `project_keys` | `ProjectKeys` | No¹ | Project keys to enumerate, e.g. `["ENG","OPS"]`. Used to build the default JQL when `jql` is empty. | | `jql` | `JQL` | No | Full JQL override. When empty, the adapter builds `project in () AND statusCategory != Done AND issuetype != Epic`. | | `repo` | `Repo` | Yes in practice | GitHub `owner/name` repo agents clone to work these issues; every returned `Issue.Repo` is set to this single value — Jira source config maps to exactly one repo, unlike Linear's per-team repo map. | @@ -299,6 +311,15 @@ Supply secrets via `${JIRA_API_TOKEN}` / `${JIRA_DATACENTER_PAT}` / `${JIRA_DATACENTER_PASSWORD}` environment-variable substitution in `hive.yaml` — never inline literal credentials in committed config. +**TLS for Data Center.** If Jira is signed by an internal CA, set +`ca_bundle` to a PEM bundle (or `${JIRA_DATACENTER_CA_BUNDLE}`) and Hive +appends those roots to the system trust store. mTLS is optional: provide both +`client_cert` and `client_key` as PEM values or environment references. The +dashboard's Work Source tab shows these Data Center-only fields as write-only +set/unset indicators. `insecure_skip_verify` remains off by default; when +enabled it bypasses TLS verification, should be limited to testing, and emits a +warning every time the Jira client is built. + **Data model differences.** Cloud users expose `accountId` and rich-text description/comment bodies as ADF JSON. Data Center commonly exposes users by `name` / `key`; the adapter prefers those identities when `deployment: