Use case
Databricks Discover Pages define governed business concepts. Published Pages become human-modeled context in the Genie Ontology, and Genie One cites them in answers.
WikiBricks already syncs immutable page versions from local SQLite to Lakebase. The weekly remote maintenance job reads that archive at a fixed watermark. Discover publication belongs in this remote workflow because the data and Databricks credentials are already available there.
The feature needs a Lakehouse staging layer for review, audit, retries, and publication state. This layer is optional. The default remote maintenance path continues to use Lakebase without a second store.
Proposed API
Add two optional tasks after remote maintenance:
local SQLite
-> immutable archive sync
Lakebase
-> remote maintenance at a fixed watermark
Unity Catalog Delta staging table
-> selection, field mapping, and policy checks
Databricks Discover Page drafts
-> owner or curator review and publication
Genie Ontology
Enable the tasks through bundle variables and a remote policy block:
databricks bundle deploy -t personal \
--var="enable_discover_pages=true" \
--var="discover_catalog=main" \
--var="discover_schema=wikibricks"
discover_pages:
enabled: true
candidates_table: main.wikibricks.discover_page_candidates
publish_state: draft
include_page_types: [entity, concept, synthesis, comparison]
include_prefixes: [topics/, guides/, comparisons/, synthesis/]
domain_map:
topics/: "<domain-id>"
The remote job projects eligible page versions from Lakebase into the Delta staging table. Use MERGE with the replica ID, path, version ID, and content hash as the source key.
| WikiBricks source |
Lakehouse staging |
Discover Page |
title |
page_name |
Page name |
summary |
description |
Description |
body |
body |
Page body |
| explicit synonyms |
synonyms |
Synonyms |
| typed links |
related_assets |
Related assets or parent and child Pages |
| source provenance |
sources |
Sources |
| replica, path, version, hash |
source identity columns |
Idempotency and conflict checks |
| configured domain map |
domain_id |
Domain or subdomain |
Deterministic logic handles selection, field mapping, and source identity. An optional model can propose synonyms, domains, or related assets. Store its evidence and confidence in the staging row. Low-confidence rows remain pending.
The publisher creates Discover Pages as drafts through a supported Databricks API or SDK. It never calls an undocumented endpoint. If no programmatic Pages interface is available, rows remain ready_for_publish for Genie Code bulk import.
Acceptance criteria
- Remote maintenance reads eligible pages from a fixed Lakebase watermark.
- A Delta table registered in Unity Catalog stores candidates and publication status.
- The job excludes raw sessions, archives,
_meta/, and pages without a domain mapping.
- Configuration selects pages by path prefix, page type, or tag for each replica.
MERGE makes staging idempotent across retries and unchanged source versions.
- The publisher creates drafts and records the Discover Page ID, source hash, state, and error details.
- A manual edit or source mismatch creates a conflict row instead of an overwrite.
- Published Pages never change without an explicit policy and supported revision contract.
- A failed staging or publication task does not advance its watermark.
- The feature is disabled by default, and its Databricks resources can stop when idle.
- Tests cover cross-replica isolation, retries, conflicts, domain mapping, and sensitive-data filters.
- Documentation warns that Discover Page data lacks customer-managed key encryption and can be replicated globally.
Alternatives considered
A local export command puts Databricks credentials and publication state on each laptop. Direct publication from Lakebase lacks an auditable review queue. MCP bulk import remains a fallback when Databricks does not expose a supported Pages API.
Additional context
Use case
Databricks Discover Pages define governed business concepts. Published Pages become human-modeled context in the Genie Ontology, and Genie One cites them in answers.
WikiBricks already syncs immutable page versions from local SQLite to Lakebase. The weekly remote maintenance job reads that archive at a fixed watermark. Discover publication belongs in this remote workflow because the data and Databricks credentials are already available there.
The feature needs a Lakehouse staging layer for review, audit, retries, and publication state. This layer is optional. The default remote maintenance path continues to use Lakebase without a second store.
Proposed API
Add two optional tasks after remote maintenance:
Enable the tasks through bundle variables and a remote policy block:
The remote job projects eligible page versions from Lakebase into the Delta staging table. Use
MERGEwith the replica ID, path, version ID, and content hash as the source key.titlepage_namesummarydescriptionbodybodysynonymsrelated_assetssourcesdomain_idDeterministic logic handles selection, field mapping, and source identity. An optional model can propose synonyms, domains, or related assets. Store its evidence and confidence in the staging row. Low-confidence rows remain pending.
The publisher creates Discover Pages as drafts through a supported Databricks API or SDK. It never calls an undocumented endpoint. If no programmatic Pages interface is available, rows remain
ready_for_publishfor Genie Code bulk import.Acceptance criteria
_meta/, and pages without a domain mapping.MERGEmakes staging idempotent across retries and unchanged source versions.Alternatives considered
A local export command puts Databricks credentials and publication state on each laptop. Direct publication from Lakebase lacks an auditable review queue. MCP bulk import remains a fallback when Databricks does not expose a supported Pages API.
Additional context