diff --git a/.gitignore b/.gitignore index 7e201db2..09659c10 100644 --- a/.gitignore +++ b/.gitignore @@ -70,7 +70,6 @@ instance/ # Sphinx documentation docs/_build/ -docs/ # PyBuilder .pybuilder/ @@ -223,7 +222,7 @@ __marimo__/ **/.DS_Store # Local design notes and plans -/docs/ +# /docs/ # impeccable design skill (local design context and live-mode config) /PRODUCT.md diff --git a/docs/STIT-766-design.md b/docs/STIT-766-design.md new file mode 100644 index 00000000..902dec16 --- /dev/null +++ b/docs/STIT-766-design.md @@ -0,0 +1,317 @@ +# STIT-766 — Current-state read model: design proposal + +**Status:** draft +**Author:** Michael Barlow +**Reviewer:** John McGrath, Alex Axthelm +**Jira:** [STIT-766](https://rmi1.atlassian.net/browse/STIT-766) (epic [STIT-689](https://rmi1.atlassian.net/browse/STIT-689)) +**Created**: 2026-09-17 + +## 1. Objective + +Design a database architecture and application logic to serve the key Stitch endpoints with minimal latency. + +## 2. Background + +### Problem + +The `list`, `filter-options`, and `detail` endpoints all rebuild the same thing on every request: one coalesced, flat row per resource that represents the current top-level "state". `queries.py` does this by joining membership → resource → default priority → source → value rows, left-joining per-field overrides, ranking candidates with a four-key `ROW_NUMBER()` window, and then pivoting the winners into columns with `max(case(...))` per field. + +Combined with serializing to/from Pydantic models, this has led to poor response times that +negatively impact user experience. Furthermore, given that our entire db is less than 100MB +in size, this raises questions about our overall db design and access patterns. + +### Observations + +With the newly added attribute-level priority feature ([STIT-494](https://rmi1.atlassian.net/browse/STIT-494)), we now have a more deterministic and bounded way to establish the different variations we display for a flat resource. Our permissions model only has 2 restricted options: `wm` and `cc`. All the remaining sources are public, and ALL users have read permissions for them. + +> [!NOTE] +> It was more straightforward to treat these individually when implemented, but reviewing our permissions model may help simplify in the future. However, it's not strictly necessary for this work. + +If we think of the set of `llm`, `rmi`, `bc`, `alb`, and `gem` as a single `public` visibility. +Then there are some fairly significant implications in how our coalescing logic actually +plays out. For any given attribute on a coalesced resource, users will see at most 3 +possible options: `wm`, `cc`, and the highest priority `public` source. To illustrate, +see the following scenarios for a resource's `latitude` data: + +- priorities: `wm`, `cc`, `gem`, `bc` + - user with `wm + cc` sees `wm` + - user with `wm` sees `wm` + - user with `cc` sees `cc` + - remaining users see `gem` +- priorities: `wm`, `cc`, `gem`, `rmi`, `llm`, `alb` + - user with `wm + cc` sees `wm` + - user with `wm` sees `wm` + - user with `cc` sees `cc` + - remaining users see `gem` +- priorities: `cc`, `gem`, `wm`, `bc` + - user with `wm + cc` sees `cc` + - user with `cc` sees `cc` + - all remaining users (including `wm` viewers) see `gem` +- priorities: `bc`, `cc`, `wm`, `gem` + - all users see `bc` +- priorities: `bc`, `gem`, `cc`, `wm` + - all users see `bc` + +For any view or operation that interacts only with the flattened representation of a resource: + +- we only need to store/cache up to the first public source in the priority list +- a resource attribute has at most 3 variations: `wm`, `cc`, and `public` +- an entire coalesced resource has **at most 4 possible variations** depending on user permissions and attribute priorities: + - `public`: all attributes have a public source as the highest priority + - `public + wm`: 1 or more attributes has `wm` as highest priority and all next-highest priorities are public + - said another way, all top 2 priorities are either `wm` or `public` and at least 1 is `wm` + - `public + cc`: 1 or more attributes has `cc` as highest priority and all next-highest priorities are public + - all top 2 priorities are either `cc` or `public` and at least 1 is `cc` + - `public + wm + cc`: at least 1 attribute where `wm` is highest **AND** 1 attribute where `cc` is highest + - Note: `wm + cc` must see different data from both `wm` and `cc` individually + +Thus, whatever structure we use here is bounded by at most 4 x the number of top-level (unmerged) resources: 1 variant for each possible licensed representation. + +> [!NOTE] +> One caveat here is that the number of possible variants grows exponentially with the number of distinct licenses: +> +> - 2 licensed sources = 4 variants +> - 3 licensed sources = 8 variants +> - 4 licensed sources = 16 variants +> +> so even if we add 2 proprietary sources, this structure grows to a max of 16 x the number of top-level (unmerged) resources. I don't see this as a real issue even in the long term. How many licensed sources are even realistic? 6? 10? Even at 10 proprietary sources, the view/table would be 2^10 = 1024 x resources. The upper bound of **all** named fields is probably somewhere in the ~65k range, so ~65M rows is still manageable...if we'd ever even get close to that. + +## 3. Goals and non-goals + +**Goals** + +1. Serve `list` and `filter-options` with <100ms latency +2. Give the database one unambiguous stored answer to "what is the value for each attribute of a resource for a user with X permissions" +3. As a follow-up to 2, provide one unambiguous stored representation for each resource for each effective user permission set. +4. Express the invariants declaratively in schema, so the sync/refresh mechanism has lower/zero risk regarding correctness +5. Preserve the public REST surface and today's behavior + +**Non-goals** + +- A general permission-based solution that accounts for all future/unknown licensed and public permissions +- The details of the sync mechanism. See §6. +- Redesigning the permission model. + +## 4. Schema Changes + +Part of the goal is to trade application complexity for schema complexity. Making schemas more complex but in a way that enables tighter constraints and simpler queries (i.e. through mostly joins) frees our application code from being the enforcer of invariants. + +### Resource State View/Table + +The purpose is to provide a durable store that houses the flattened/coalesced resources. We're effectively precomputing the coalescing logic and saving it to a single table. The actual schema is less important than its function and the constraints we place on it. + +It must: + +- only store unmerged resources (i.e. where `repointed_id == NULL`), merging triggers deletion from the table +- provide highly performant, permission-scoped querying of resource data +- not expose proprietary information to unlicensed users +- provide resource representations that are consistent with existing coalescing logic +- be able to be rebuilt from scratch at any time, we should be able to derive the data for the table easily & quickly + +The top contender for a schema is: + +```mermaid +erDiagram + og_field_resource_view { + bignt resource_id + smallint permission_mask + jsonb record + } +``` + +**permission mask** +The `permission_mask` is a bitmask where `public` = 0, `cc` = 1, `wm` = 2, and `wm + cc` = 3. + +| wm | cc | perm | +| --- | --- | ---- | +| 0 | 0 | 0 | +| 0 | 1 | 1 | +| 1 | 0 | 2 | +| 1 | 1 | 3 | + +This allows for 2 filtering options: + +- strict `permission_mask = ` (incurs minor cost of possibly duplicating data across rows) +- bit comparison to filter where `( | permission_mask) = ` + - if we sort by `permission_mask` (desc), this would allow us to only store the minimum number of resource variants + - for example, if a resource had ALL public sources as the highest priorities, we'd only need 1 row in the table with `permission_mask = 0` because ALL users would see the same version + - or if a resource had only `wm` and `public` variants, a `wm + cc` permission would get the `wm` version: 3 (`cc + wm`) | 2 (`wm`) = 3 + - but a `cc` permission would skip the `wm` row and see the `public` variant: + 1 (`cc`) | 2 (`wm`) = 3 => exclude + 1 (`cc`) | 0 (`pub`) = 1 => include + +> [!NOTE] Permission Alternative +> We can also use permission columns for the minor cost of duplicating data across columns. Benefits from being a simpler more understandable approach. + +**record column as jsonb** +We'd effectively house the entire flat Pydantic model in json. The main reasoning is that we can likely expand to near-instant full-text search without much difficulty. It's also partially as an experiment to assess the difficulty of working with Postgres JSON syntax and investigate whether there are performance trade-offs. Should it prove easy to use while still being performant, it opens the door to using it in other places across our application where we might want greater flexibility in our data handling. + +**sync overview** +Very roughly speaking, when a user or process updates relevant data (merge resources, reprioritize, new sources from ETL), we compute the updated view state(s), and write them to the table, deleting where necessary. + +```mermaid +flowchart LR + A[db updates] -->|Trigger| B[compute coalesced state] + B --> C[Write coalesced state to table] + D(merge) --> A + E(reprioritize) --> A + F(llm enrich) --> A + G(ETL) --> A +``` + +### Attribute metadata: `og_field_attributes` + +```mermaid +erDiagram + og_field_attributes { + serial id + text name + } + +``` + +Priority rows and EAV rows reference `og_field_attribute.id` + +### Single priority store: `og_field_resource_attribute_priority` + +Replace the two-table priority split with a single table at the +`(resource, attribute, source record)` grain, and use defaults when +writing new data rather than as a SQL fallback rule. + +```mermaid +erDiagram + og_field_resource_attribute_priority { + bigint resource_id + bigint attribute_id + bigint source_id + int priority + boolean is_curated + } +``` + +- Drop `og_field_source_priority` and `og_field_resource_source_priority`. +- On resource or source creation, write explicit priority rows for every attribute + the source has a value for, seeded from the existing `SOURCE_PRIORITY` tuple in + `stitch.ogsi.model`. +- Add constraints: + - unique priority across resource and attribute: no 2 rows for a resource & attribute can have the same priority + `UNIQUE (resource_id, attribute_id, priority)` + - priority > 0 + - fk constraint to memberships on resource_id, source_id: ensure source is attached to resource + - fk constraint to `og_field_source_values` on `source_id, attribute_id` +- `oil_gas_field_source_values` currently has an `id` primary key, so we could replace `(attribute_id, source_id)` with `source_value_id` +- we set `is_curated` to `True` when a user makes an update + +### Optional Additional Tables + +These are primarily for some convenience in constructing simpler SQL statements and permissions handling. + +```mermaid +erDiagram + og_source_keys { + bigint id + text key_name + } +``` + +References to `source` or `src_key` become foreign keys with `source_key_id`. + +```mermaid +erDiagram + user_source_key_permissions { + uuid user_id + bigint source_key_id + source_key_permission perm + } +``` + +Where the `source_key_permission` is a new ENUM type with `read`, `edit`. + +What this would allow is comparatively smaller and more straightforward SQL statements. + +**single resource for a user** + +```sql +WITH resolved AS ( + SELECT DISTINCT ON ( + p.resource_id, + p.attribute_id + ) + p.resource_id, + p.attribute_id, + p.source_id, + s.source_key_id, + p.priority, + v.value_text, + v.value_num, + v.value_json + FROM og_field_resource_attribute_priority AS p + JOIN oil_gas_field_source_values AS v + ON v.source_id = p.source_id + AND v.attribute_id = p.attribute_id + JOIN oil_gas_field_sources AS s + ON s.source_id = p.source_id + JOIN user_source_key_permission AS permission + ON permission.source_key_id = s.source_key_id + AND permission.user_id = $1 -- pass in specific user id + AND permission.action = 'read' + WHERE p.resource_id = $2 -- drop this predicate to fetch many resources + ORDER BY + p.resource_id, + p.attribute_id, + p.priority + -- p.source_id <-- unnecessary with the priority uniqueness constraint +) +SELECT + r.resource_id, + jsonb_object_agg( + a.name, + COALESCE( + to_jsonb(r.value_text), + to_jsonb(r.value_num), + r.value_json + ) + ) AS attributes +FROM resolved AS r +JOIN og_field_attributes AS a + ON a.attribute_id = r.attribute_id +GROUP BY r.resource_id; + +``` + +Note: The above can be paged and totaled as well with minimal alteration. + +**single resource detailed provenance** + +```sql +SELECT + p.resource_id, + a.name AS attribute, + p.source_id, + sk.key_name AS source_key, + p.priority, + v.value_text, + v.value_num, + v.value_json +FROM og_field_resource_attribute_priority AS p +JOIN oil_gas_field_source_values AS v + ON v.source_id = p.source_id + AND v.attribute_id = p.attribute_id +JOIN og_field_attributes AS a + ON a.attribute_id = p.attribute_id +JOIN og_field_sources AS s + ON s.source_id = p.source_id +JOIN og_source_keys AS sk + ON sk.source_key_id = s.source_key_id +JOIN user_source_key_permission AS permission + ON permission.source_key_id = s.source_key_id + AND permission.user_id = $1 + AND permission.action = 'view' +WHERE p.resource_id = $2 +ORDER BY p.priority, p.source_id; +``` + +## 5. Open Issues + +1. Priority row count at current and projected resource volume. +2. How to handle priorities upon merge? diff --git a/docs/draft-schema-operations-api.md b/docs/draft-schema-operations-api.md new file mode 100644 index 00000000..cc0efc79 --- /dev/null +++ b/docs/draft-schema-operations-api.md @@ -0,0 +1,675 @@ +# Source-Prioritized Resource Schema: Operational API + +This document describes how application code should operate on the schema in +`01_schema.sql`. It treats the database model as an API: the supported commands, +read semantics, invariants, transaction boundaries, and expected failure modes. + +## 1. Core semantics + +A **resource** is a minimal logical record assembled from one or more **sources**. +Each source belongs to a **source key**, such as `A`, `B`, or `licensed_vendor`. +A user may view only the source keys for which they have `view` permission. + +Source data is represented as populated attribute rows: + +```text +(source_id, attribute_id) -> exactly one of value_text, value_num, value_json +``` + +For each resource and attribute, `resource_attribute_priority` defines an ordered +list of source candidates. Lower numeric priority wins. + +The resolved value is therefore: + +> the highest-priority populated value whose source key is visible to the user + +Permission filtering must happen **before** selecting the winning priority. If a +restricted source at priority 0 contains a value and the user cannot view its +key, the user receives the next visible candidate. + +## 2. Important invariants + +The schema enforces these relationships: + +1. A source can participate in a resource only through `resource_source`. +2. A priority entry can reference only a source attached to the resource. +3. A priority entry can reference only a populated source attribute. +4. A source attribute contains exactly one typed value. +5. Two sources cannot occupy the same priority position for the same resource + and attribute. +6. A source cannot appear twice in the priority list for the same resource and + attribute. +7. Visibility is determined by the source's key, not by the resource or + attribute independently. + +The schema does **not** require priorities to be contiguous. Positions `0, 10, +20` are valid and often easier to maintain than `0, 1, 2`. + +## 3. Suggested application service surface + +A service layer can expose operations similar to the following: + +```text +create_resource() -> resource_id +create_source(source_key, external_id?) -> source_id +attach_source(resource_id, source_id) +define_attribute(name, description?, unit?) -> attribute_id +put_source_value(source_id, attribute, typed_value) +delete_source_value(source_id, attribute) +set_attribute_priorities(resource_id, attribute, ordered_source_ids) +grant_source_key_permission(user_id, source_key, action = "view") +revoke_source_key_permission(user_id, source_key, action = "view") +resolve_resource(user_id, resource_id) -> flat object +search_resources(user_id, predicates) -> resource ids or flat objects +explain_resolved_value(user_id, resource_id, attribute) -> provenance +``` + +Database functions are optional. These can be implemented in application code +using parameterized SQL while preserving the same semantics. + +## 4. Creating resources and sources + +### Create a resource + +```sql +INSERT INTO resource DEFAULT VALUES +RETURNING resource_id; +``` + +### Resolve or create a source key + +```sql +INSERT INTO source_key (key_name) +VALUES ($1) +ON CONFLICT (key_name) +DO UPDATE SET key_name = EXCLUDED.key_name +RETURNING source_key_id; +``` + +### Create a source + +```sql +INSERT INTO source (source_key_id, external_id) +VALUES ($1, $2) +RETURNING source_id; +``` + +### Attach a source to a resource + +```sql +INSERT INTO resource_source (resource_id, source_id) +VALUES ($1, $2) +ON CONFLICT DO NOTHING; +``` + +Attaching a source does not automatically prioritize every value from that +source. Priority remains explicit per attribute. + +## 5. Defining attributes + +Use a stable attribute name at the API boundary, but use `attribute_id` +internally. + +```sql +INSERT INTO attribute_definition (name, description, unit) +VALUES ($1, $2, $3) +ON CONFLICT (name) +DO UPDATE SET + description = COALESCE(EXCLUDED.description, + attribute_definition.description), + unit = COALESCE(EXCLUDED.unit, attribute_definition.unit) +RETURNING attribute_id; +``` + +Renaming an attribute should be treated as a controlled schema operation. The +priority and value tables remain stable because they reference `attribute_id`, +not the textual name. + +## 6. Writing source values + +A source value exists only when it is populated. SQL `NULL` means "no candidate +value" and should normally be represented by deleting the row. + +### Text value + +```sql +INSERT INTO source_attribute_value ( + source_id, + attribute_id, + value_text +) +VALUES ($1, $2, $3) +ON CONFLICT (source_id, attribute_id) +DO UPDATE SET + value_text = EXCLUDED.value_text, + value_num = NULL, + value_json = NULL, + updated_at = now(); +``` + +### Numeric value + +```sql +INSERT INTO source_attribute_value ( + source_id, + attribute_id, + value_num +) +VALUES ($1, $2, $3) +ON CONFLICT (source_id, attribute_id) +DO UPDATE SET + value_text = NULL, + value_num = EXCLUDED.value_num, + value_json = NULL, + updated_at = now(); +``` + +### JSON value + +```sql +INSERT INTO source_attribute_value ( + source_id, + attribute_id, + value_json +) +VALUES ($1, $2, $3::jsonb) +ON CONFLICT (source_id, attribute_id) +DO UPDATE SET + value_text = NULL, + value_num = NULL, + value_json = EXCLUDED.value_json, + updated_at = now(); +``` + +The JSON literal `null` is rejected. Use row deletion to represent absence. + +### Delete a value + +```sql +DELETE FROM source_attribute_value +WHERE source_id = $1 + AND attribute_id = $2; +``` + +This deletion fails while a priority row references the value. The application +must first remove or replace that source in affected priority lists. This is an +intentional consistency guard. + +A safe transaction is: + +```sql +BEGIN; + +DELETE FROM resource_attribute_priority +WHERE source_id = $1 + AND attribute_id = $2; + +DELETE FROM source_attribute_value +WHERE source_id = $1 + AND attribute_id = $2; + +COMMIT; +``` + +## 7. Managing priority lists + +Treat a priority list as one ordered aggregate rather than many unrelated rows. +Replacing the complete list in one transaction is usually safer than issuing +individual moves. + +Given an ordered input such as: + +```json +[17, 42, 91] +``` + +assign priorities using ordinality: + +```sql +BEGIN; + +DELETE FROM resource_attribute_priority +WHERE resource_id = $1 + AND attribute_id = $2; + +INSERT INTO resource_attribute_priority ( + resource_id, + attribute_id, + source_id, + priority +) +SELECT + $1, + $2, + input.source_id, + input.ordinality - 1 +FROM unnest($3::bigint[]) WITH ORDINALITY AS input(source_id, ordinality); + +COMMIT; +``` + +The insert fails when: + +- a source is not attached to the resource; +- a source lacks a value for the attribute; +- the source appears more than once; +- the input creates duplicate priority positions. + +Those failures are useful API validation errors and should normally map to a +client error rather than an internal server error. + +### Incremental insertion with sparse ranks + +To avoid renumbering, use ranks such as `1000, 2000, 3000`. Insert a new source +between the first and second at `1500`. Renormalize only when no gap remains. + +## 8. Granting and revoking visibility + +### Grant permission + +```sql +INSERT INTO user_source_key_permission ( + user_id, + source_key_id, + action +) +VALUES ($1, $2, 'view') +ON CONFLICT DO NOTHING; +``` + +### Revoke permission + +```sql +DELETE FROM user_source_key_permission +WHERE user_id = $1 + AND source_key_id = $2 + AND action = 'view'; +``` + +Permission changes take effect immediately for dynamically resolved reads. +Caches must include the user's effective visibility profile in the cache key or +be invalidated when permissions change. + +## 9. Resolving one resource for one user + +The permission join removes inaccessible candidates before `DISTINCT ON` +chooses the winner. + +```sql +WITH resolved AS ( + SELECT DISTINCT ON ( + p.resource_id, + p.attribute_id + ) + p.resource_id, + p.attribute_id, + p.source_id, + s.source_key_id, + p.priority, + v.value_text, + v.value_num, + v.value_json + FROM resource_attribute_priority AS p + JOIN source_attribute_value AS v + ON v.source_id = p.source_id + AND v.attribute_id = p.attribute_id + JOIN source AS s + ON s.source_id = p.source_id + JOIN user_source_key_permission AS permission + ON permission.source_key_id = s.source_key_id + AND permission.user_id = $1 + AND permission.action = 'view' + WHERE p.resource_id = $2 + ORDER BY + p.resource_id, + p.attribute_id, + p.priority, + p.source_id +) +SELECT + r.resource_id, + jsonb_object_agg( + a.name, + COALESCE( + to_jsonb(r.value_text), + to_jsonb(r.value_num), + r.value_json + ) + ) AS attributes +FROM resolved AS r +JOIN attribute_definition AS a + ON a.attribute_id = r.attribute_id +GROUP BY r.resource_id; +``` + +The result is a user-specific flat JSON object. + +### Empty resources + +The query above returns no row when the user can see no resolved attributes. +To return the resource with an empty JSON object, start from `resource` and use +a lateral or left join. + +## 10. Resolving many resources + +The same resolution CTE can omit the `resource_id` predicate. For large result +sets, page by `resource_id` and resolve only the requested page. + +Avoid flattening every attribute for every resource before applying selective +filters. Resolve only the attributes needed for filtering or output. + +## 11. Provenance and explanation + +A useful API should expose why a user received a particular value. + +```sql +SELECT + p.resource_id, + a.name AS attribute, + p.source_id, + sk.key_name AS source_key, + p.priority, + v.value_text, + v.value_num, + v.value_json +FROM resource_attribute_priority AS p +JOIN source_attribute_value AS v + ON v.source_id = p.source_id + AND v.attribute_id = p.attribute_id +JOIN attribute_definition AS a + ON a.attribute_id = p.attribute_id +JOIN source AS s + ON s.source_id = p.source_id +JOIN source_key AS sk + ON sk.source_key_id = s.source_key_id +JOIN user_source_key_permission AS permission + ON permission.source_key_id = s.source_key_id + AND permission.user_id = $1 + AND permission.action = 'view' +WHERE p.resource_id = $2 + AND a.name = $3 +ORDER BY p.priority, p.source_id; +``` + +The first row is the visible winner. Remaining rows are visible fallbacks. +Do not reveal hidden candidates or their keys unless the user has separate +permission to inspect restricted provenance. + +## 12. Dynamic search by unknown attributes + +The search API may accept an arbitrary list of predicates. An extensible input +shape is: + +```json +[ + { + "attribute": "country", + "operator": "eq", + "value_text": "US" + }, + { + "attribute": "production", + "operator": "gte", + "value_num": 1000 + } +] +``` + +The correct order of operations is: + +```text +parse predicates + -> resolve attribute names to IDs + -> restrict candidates to visible source keys + -> choose the highest-priority visible value + -> apply predicates + -> require all predicates to match +``` + +Filtering raw source values before priority resolution answers a different +question: "does any permitted source contain this value?" The usual API should +instead filter the flattened representation the user actually sees. + +### Equality-only search + +This query accepts typed equality predicates as JSON and requires every +predicate to match. + +```sql +WITH requested_filters AS ( + SELECT + definition.attribute_id, + filter.attribute, + filter.value_text, + filter.value_num, + filter.value_json + FROM jsonb_to_recordset($2::jsonb) AS filter( + attribute text, + value_text text, + value_num numeric, + value_json jsonb + ) + JOIN attribute_definition AS definition + ON definition.name = filter.attribute +), +resolved AS ( + SELECT DISTINCT ON ( + p.resource_id, + p.attribute_id + ) + p.resource_id, + p.attribute_id, + v.value_text, + v.value_num, + v.value_json + FROM resource_attribute_priority AS p + JOIN requested_filters AS filter + ON filter.attribute_id = p.attribute_id + JOIN source_attribute_value AS v + ON v.source_id = p.source_id + AND v.attribute_id = p.attribute_id + JOIN source AS s + ON s.source_id = p.source_id + JOIN user_source_key_permission AS permission + ON permission.source_key_id = s.source_key_id + AND permission.user_id = $1 + AND permission.action = 'view' + ORDER BY + p.resource_id, + p.attribute_id, + p.priority, + p.source_id +), +matched AS ( + SELECT + resolved.resource_id, + resolved.attribute_id + FROM resolved + JOIN requested_filters AS filter + ON filter.attribute_id = resolved.attribute_id + AND ( + ( + filter.value_text IS NOT NULL + AND resolved.value_text = filter.value_text + ) + OR ( + filter.value_num IS NOT NULL + AND resolved.value_num = filter.value_num + ) + OR ( + filter.value_json IS NOT NULL + AND resolved.value_json = filter.value_json + ) + ) +) +SELECT matched.resource_id +FROM matched +GROUP BY matched.resource_id +HAVING count(*) = (SELECT count(*) FROM requested_filters) +ORDER BY matched.resource_id; +``` + +Input validation should require exactly one filter value column per predicate, +just as the value table requires exactly one storage column. + +### Supporting comparison operators + +Add an `operator` field and dispatch by value type: + +```sql +AND CASE filter.operator + WHEN 'eq' THEN resolved.value_num = filter.value_num + WHEN 'neq' THEN resolved.value_num <> filter.value_num + WHEN 'lt' THEN resolved.value_num < filter.value_num + WHEN 'lte' THEN resolved.value_num <= filter.value_num + WHEN 'gt' THEN resolved.value_num > filter.value_num + WHEN 'gte' THEN resolved.value_num >= filter.value_num + ELSE false +END +``` + +Do not accept arbitrary SQL operators or column names from clients. Parse a +small allowlisted operator vocabulary and bind all values as parameters. + +### Duplicate attributes + +Decide explicitly whether repeated predicates for one attribute mean AND, OR, +or invalid input. A simple first version should reject duplicates. + +### Unknown attributes + +The join to `attribute_definition` silently drops unknown names. For an API, +validate before executing the search and return a clear client error listing +unknown attributes. + +## 13. Pagination and result shape + +A search endpoint may return only resource IDs first: + +```text +search_resources(user, predicates, cursor, limit) -> [resource_id] +``` + +Then fetch flattened representations for those IDs in a second query. This is +often easier to optimize than resolving and aggregating every matching resource +inside one large statement. + +Prefer keyset pagination: + +```sql +WHERE resource_id > $cursor +ORDER BY resource_id +LIMIT $limit +``` + +## 14. Concurrency + +Priority-list replacement should occur in one transaction. Concurrent writers +for the same `(resource_id, attribute_id)` can serialize using a row lock or an +advisory transaction lock. + +One option is to lock the resource row: + +```sql +SELECT 1 +FROM resource +WHERE resource_id = $1 +FOR UPDATE; +``` + +A narrower advisory lock can be derived from the resource and attribute IDs: + +```sql +SELECT pg_advisory_xact_lock($1, $2); +``` + +Choose one strategy and use it consistently in all priority mutation paths. + +## 15. Row-level security + +Application-level permission joins make semantics explicit. PostgreSQL row-level +security can add defense in depth so restricted source values cannot be read by +accident. + +A common request transaction is: + +```sql +BEGIN; +SET LOCAL app.user_id = ''; +-- Execute permission-aware queries. +COMMIT; +``` + +The authenticated identity must come from trusted server-side authentication, +not a client-controlled query parameter. + +RLS does not eliminate the need to reason about provenance leakage. Protect +`source`, `resource_attribute_priority`, and diagnostic views as needed if even +the existence of restricted sources is confidential. + +## 16. Caching and materialization + +The flattened representation is user-dependent. A cache keyed only by +`resource_id` is incorrect when users have different source-key permissions. + +Valid cache keys include: + +```text +(resource_id, user_id, permission_version) +(resource_id, role_id, permission_version) +(resource_id, visibility_profile_hash) +``` + +Role- or profile-based caching is preferable when many users share identical +grants. + +A globally materialized resolved table works only for a canonical visibility +profile. For multiple profiles, materialize per profile or resolve dynamically. + +## 17. Deletion behavior + +Deleting a resource cascades through its attachments and priority rows. + +Deleting a source cascades through its values, attachments, and priority rows. +This is convenient but may be too destructive for audited systems. Alternatives +include soft deletion or `ON DELETE RESTRICT` plus explicit archival workflows. + +Deleting an attribute definition is restricted while values or priorities refer +to it. Treat attribute deletion as a migration rather than a routine API call. + +## 18. Recommended validation errors + +Map database constraint failures to stable API errors: + +| Condition | Suggested API error | +|---|---| +| Source not attached to resource | `SOURCE_NOT_ATTACHED` | +| Source has no value for attribute | `SOURCE_ATTRIBUTE_MISSING` | +| Duplicate priority position | `PRIORITY_CONFLICT` | +| Duplicate source in list | `DUPLICATE_PRIORITY_SOURCE` | +| No typed value or multiple typed values | `INVALID_TYPED_VALUE` | +| JSON literal null supplied | `INVALID_JSON_NULL` | +| Unknown attribute name | `UNKNOWN_ATTRIBUTE` | +| Unsupported operator | `UNSUPPORTED_OPERATOR` | +| User lacks permission | `FORBIDDEN_SOURCE_KEY` or omit existence details | + +For security-sensitive operations, avoid distinguishing "does not exist" from +"exists but forbidden" unless the caller is allowed to know that distinction. + +## 19. Recommended initial implementation + +A practical first version is: + +1. Use the canonical typed EAV tables. +2. Store direct `view` grants by source key. +3. Replace full priority lists transactionally. +4. Resolve with a permission join followed by `DISTINCT ON`. +5. Accept equality-only dynamic filters initially. +6. Resolve only requested attributes during search. +7. Add RLS after application queries and transaction identity handling are + stable. +8. Add role-based permission profiles if caching or permission-row volume + becomes significant. + +This preserves strong relational constraints while keeping the application API +predictable and extensible. diff --git a/docs/draft-schema.sql b/docs/draft-schema.sql new file mode 100644 index 00000000..4df990ca --- /dev/null +++ b/docs/draft-schema.sql @@ -0,0 +1,223 @@ +-- Source-prioritized, permission-aware EAV schema +-- Target database: PostgreSQL 15+ +-- +-- Core assumptions: +-- * Lower integer values represent higher priority. +-- * A source attribute row exists only when the source has a value. +-- * Exactly one typed value column is populated per source attribute. +-- * Source visibility is granted by source key and action. +-- * Permission filtering occurs before priority resolution. + +BEGIN; + +CREATE TYPE source_key_action AS ENUM ( + 'view', + 'edit' +); + +CREATE TABLE users ( + user_id uuid PRIMARY KEY +); + +CREATE TABLE og_field_resources ( + resource_id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, + created_at timestamptz NOT NULL DEFAULT now() +); + +CREATE TABLE source_keys ( + source_key_id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, + key_name text NOT NULL UNIQUE +); + +CREATE TABLE oil_gas_field_sources ( + source_id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, + source_key_id bigint NOT NULL, + external_id text, + created_at timestamptz NOT NULL DEFAULT now(), + + CONSTRAINT source_source_key_fk + FOREIGN KEY (source_key_id) + REFERENCES source_keys(source_key_id) +); + +CREATE TABLE attribute_definition ( + attribute_id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY, + name text NOT NULL UNIQUE, + description text, + unit text, + created_at timestamptz NOT NULL DEFAULT now() +); + +CREATE TABLE memberships ( + resource_id bigint NOT NULL, + source_id bigint NOT NULL, + attached_at timestamptz NOT NULL DEFAULT now(), + + PRIMARY KEY (resource_id, source_id), + + CONSTRAINT resource_source_resource_fk + FOREIGN KEY (resource_id) + REFERENCES og_field_resources(resource_id) + ON DELETE CASCADE, + + CONSTRAINT resource_source_source_fk + FOREIGN KEY (source_id) + REFERENCES oil_gas_field_sources(source_id) + ON DELETE CASCADE +); + +CREATE TABLE oil_gas_field_source_values ( + source_id bigint NOT NULL, + attribute_id bigint NOT NULL, + + value_text text, + value_num numeric, + value_json jsonb, + + created_at timestamptz NOT NULL DEFAULT now(), + updated_at timestamptz NOT NULL DEFAULT now(), + + PRIMARY KEY (source_id, attribute_id), + + CONSTRAINT oil_gas_field_source_values_source_fk + FOREIGN KEY (source_id) + REFERENCES oil_gas_field_sources(source_id) + ON DELETE CASCADE, + + CONSTRAINT oil_gas_field_source_values_attribute_fk + FOREIGN KEY (attribute_id) + REFERENCES attribute_definition(attribute_id), + + CONSTRAINT oil_gas_field_source_values_exactly_one_value_ck + CHECK (num_nonnulls(value_text, value_num, value_json) = 1), + + CONSTRAINT oil_gas_field_source_values_nonnull_json_ck + CHECK (value_json IS NULL OR value_json <> 'null'::jsonb) +); + +CREATE TABLE og_field_resource_attribute_priority ( + resource_id bigint NOT NULL, + attribute_id bigint NOT NULL, + source_id bigint NOT NULL, + priority integer NOT NULL, + created_at timestamptz NOT NULL DEFAULT now(), + + PRIMARY KEY (resource_id, attribute_id, source_id), + + CONSTRAINT resource_attribute_priority_position_uq + UNIQUE (resource_id, attribute_id, priority), + + CONSTRAINT resource_attribute_priority_nonnegative_ck + CHECK (priority >= 0), + + -- The source must be attached to the resource. + CONSTRAINT resource_attribute_priority_resource_source_fk + FOREIGN KEY (resource_id, source_id) + REFERENCES memberships(resource_id, source_id) + ON DELETE CASCADE, + + -- The source must have a populated value for the selected attribute. + CONSTRAINT resource_attribute_priority_source_value_fk + FOREIGN KEY (source_id, attribute_id) + REFERENCES oil_gas_field_source_values(source_id, attribute_id) +); + +-- handles the permissions for a user associated with a given source key +-- user_id | src_key_id | action +-- 1 | 1 (wm) | view -> only wm + public +-- 1 | 2 (rmi) | view +-- 1 | 3 (gem) | view +-- 1 | 4 (alb) | view +-- 1 | 5 (bc) | view +-- 1 | 6 (llm) | view + +-- 2 | 7 (cc) | view -> only cc + public +-- 2 | 2 (rmi) | view +-- 2 | 3 (gem) | view +-- 2 | 4 (alb) | view +-- 2 | 5 (bc) | view +-- 2 | 6 (llm) | view + +-- 3 | 2 (rmi) | view -> only public +-- 3 | 3 (gem) | view +-- 3 | 4 (alb) | view +-- 3 | 5 (bc) | view +-- 3 | 6 (llm) | view + +-- 4 | 1 (wm) | view -> all: wm + cc + public +-- 4 | 7 (cc) | view +-- 4 | 2 (rmi) | view +-- 4 | 3 (gem) | view +-- 4 | 4 (alb) | view +-- 4 | 5 (bc) | view +-- 4 | 6 (llm) | view +CREATE TABLE user_source_key_permission ( + user_id uuid NOT NULL, + source_key_id bigint NOT NULL, + action source_key_action NOT NULL, + granted_at timestamptz NOT NULL DEFAULT now(), + + PRIMARY KEY (user_id, source_key_id, action), + + CONSTRAINT user_source_key_permission_user_fk + FOREIGN KEY (user_id) + REFERENCES users(user_id) + ON DELETE CASCADE, + + CONSTRAINT user_source_key_permission_key_fk + FOREIGN KEY (source_key_id) + REFERENCES source_keys(source_key_id) + ON DELETE CASCADE +); + +-- Resolution path: resource/attribute -> priority -> source -> source key. +CREATE INDEX resource_attribute_priority_resolution_idx + ON og_field_resource_attribute_priority ( + resource_id, + attribute_id, + priority, + source_id + ); + +CREATE INDEX resource_attribute_priority_source_idx + ON og_field_resource_attribute_priority ( + source_id, + attribute_id + ); + +CREATE INDEX source_key_lookup_idx + ON oil_gas_field_sources ( + source_key_id, + source_id + ); + +CREATE INDEX user_source_key_view_idx + ON user_source_key_permission ( + user_id, + source_key_id + ) + WHERE action = 'view'; + +-- Optional lookup indexes for filtering before or after resolution. +CREATE INDEX oil_gas_field_source_values_text_lookup_idx + ON oil_gas_field_source_values ( + attribute_id, + value_text, + source_id + ) + WHERE value_text IS NOT NULL; + +CREATE INDEX oil_gas_field_source_values_num_lookup_idx + ON oil_gas_field_source_values ( + attribute_id, + value_num, + source_id + ) + WHERE value_num IS NOT NULL; + +CREATE INDEX oil_gas_field_source_values_json_gin_idx + ON oil_gas_field_source_values + USING gin (value_json) + WHERE value_json IS NOT NULL; + +COMMIT;