feat(curator): nightly local curation with safe auto-apply - #21
Merged
Merged
Conversation
New optional package `wikibricks_curator` (stdlib only, no psycopg or Databricks SDK; the core library does not import it). - `build_request` turns one curation-backlog item into a bounded request: the project's living page (first covering `topics/` page, else `topics/<project>`), covering and related pages in the shape `build_patches` expects, and only user/assistant messages from the project's sessions after the page update or cursor, newest kept first, capped at 40 events and 30,000 characters. - `chat_json` calls an OpenAI-compatible chat-completions endpoint with the remote curator's message layout and JSON parsing; `resolve_token` reads WIKIBRICKS_GATEWAY_TOKEN or `databricks auth token --profile`. Implemented by GLM 5.3 Flash; timestamp comparison across UTC offsets and per-project event loading fixed in review. Co-authored-by: Isaac <no-reply@databricks.com>
`wikibricks-curator propose --base-url URL --profile P` sends each of the top backlog projects' new session text to one model (default GLM 5.3 Flash) and turns the reply into a local curation run. - Keeps one living page per project (`topics/<project>` or the existing covering topics page) through a local prompt addendum on top of the remote curator prompt; only create_page, update_page and add_link. - Auto-applies low-risk groups with the existing `safe` policy. Updates that would drop more than half of a page's text are forced to high risk and wait for review; `wiki_index` lists pending review runs. - `store_manifest` stores a run locally; `pull_manifests` now uses it. - A cursor per project stops the same sessions from being proposed twice. Implemented by GLM 5.3 Flash. Review fixed the gateway call (positional timeout sent as the request body; `reasoning_effort: low`, without which GLM 5.3 Flash spends all 8,192 output tokens on reasoning), a dry run that advanced the cursor, and a guard that crashed on pages without a summary. Co-authored-by: Isaac <no-reply@databricks.com>
Acceptance runs against GLM 5.3 Flash on a copy of the live store showed three failure modes; each lost a whole project run: - A link from the living page created in the same reply: build_patches only knows existing pages. Such links are now dropped and counted; `curate` adds mention links once the page exists. - Omitted fields with no content value (risk_class, tags, source_ids, target_path, group; title/summary/body on links) get neutral defaults, and unknown keys are dropped. Page proposals without content still fail. - An undecodable or invalid reply is retried once. After the fixes: 8 of 8 projects applied, 0 errors, 100 s for 6 projects. Co-authored-by: Isaac <no-reply@databricks.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The store records about 240 sessions but has 19 curated pages. Nothing turns sessions into pages unless an agent is prompted during a session. The remote curator job needs Lakebase and never ran in production.
Change
A new optional package and console script,
wikibricks-curator. The core library still never calls a model.build_requestcollects the living page, covering and related pages, and the new sessions' user and assistant messages only. Tool output and reasoning stay out, capped at 40 messages and 30,000 characters.system.ai.glm-5-3-flash, withreasoning_effort: low. It uses the standard library only, with no Lakebase, SDK or job. The token comes fromWIKIBRICKS_GATEWAY_TOKENordatabricks auth token --profile.topics/<project>or its existing coveringtopics/page, with sections Current state, Decisions, Open issues and Key facts. Onlycreate_page,update_pageandadd_linkare allowed.build_patches(moved in refactor(remote): move build_patches into a psycopg-free module #20) and is stored with the newstore_manifest.apply_run(policy="safe")then applies it automatically.wiki_indexshows_meta/curation-reviewwhile patches have no receipt.curatelinks it later).Verification
uv.lockis unchanged. The curator tests cover evidence selection across UTC offsets, gateway calls (keyword timeout,reasoning_effort, parsing, HTTP errors), create, update, shrink guard, review entry, errors, retries, dry runs, the CLI, packaging and idempotentstore_manifest.reasoning_effort: low, all 8,192 output tokens went to reasoning. A positional timeout was sent as the request body. There were same-run links and an omittedrisk_class.topics/pages for 7 projects, plus an update totopics/agent-atlas(4,684 to 5,651 characters, citing 8 messages). All have the four sections and dated facts, and a secret-pattern scan found nothing.~/code, the emails folder) get project pages too.created_by = remote-curator, which the apply path hard-codes for all curation runs.GLM 5.3 Flash wrote the implementation in two steps. Review fixed the gateway call, timestamp comparison, per-project event loading, dry-run cursor writes and the guard crash. The acceptance fixes came from real model runs.
This pull request and its description were written by Isaac.