Skip to content

fix(athena): S3 Tables macros, bucket groups, schema LF-Tags, catalog-qualified columns - #1

Merged
aoelvp94 merged 4 commits into
aoelvp94:athena/part-4-macro-packagefrom
fpiped:fix/athena-part4-s3-tables-lf
Sep 28, 2026
Merged

aoelvp94 merged 4 commits into
aoelvp94:athena/part-4-macro-packagefrom
fpiped:fix/athena-part4-s3-tables-lf

Conversation

@fpiped

@fpiped fpiped commented Sep 25, 2026

Copy link
Copy Markdown

Fixes for dbt-labs#16369 found running the whole Athena stack against a real work group; one commit each, so any can be dropped:

Tested: cargo nextest run -p dbt-loader (228 passed) on this branch; end to end with the other parts (Iceberg bucket partitions, schema LF-Tags, catalogs.yml v2 s3_tables table and incremental models built twice). The companion fix for list_relations is on athena/part-5-metadata-adapter.

The vendored package predates dbt-adapters 62a5495cf (S3 Tables catalogs).
From that commit, applied as is: drop_relation drops an S3 Tables relation
through Glue; the table materialization drops and recreates it, since S3
Tables allows no ALTER TABLE RENAME; create_table_as sets no S3 location and
skips the direct S3 cleanup for it; incremental skips the Glue table version
expiry, which the federated catalog does not serve.
dbt v2's regex match objects do not support indexing (dbt-labs#16451), so
bucket_match[1] rendered empty; .group(n) works in dbt v1 and v2.
As dbt-athena does after the DDL; adapter.add_lf_tags_to_database is
implemented in dbt-labs#16376.
Athena's information_schema covers only the catalog it is read from, so the
columns of a relation in another catalog (S3 Tables, s3tablescatalog/<bucket>)
came back empty.
aoelvp94 added a commit that referenced this pull request Sep 27, 2026
…mation_schema

Relation and column lookups went to Glue unconditionally. Glue serves only the
Data Catalog, so a relation in a federated connector or in S3 Tables
(`s3tablescatalog/<bucket>`) came back empty and the next build re-ran its CTAS
(TABLE_ALREADY_EXISTS). Reading everything from `information_schema` instead
would fix that at the cost of an Athena query per relation, each billed at the
10 MB minimum.

Dispatch on the catalog: Glue when it is the Data Catalog, otherwise that
catalog's own `information_schema` — Athena's covers only the catalog it is
read from, so the catalog belongs in the table reference, not just in a filter.
`list_relations` also falls back when the Glue call fails at all, since Lake
Formation can deny GetTables on tables that stay queryable through Athena.

The `information_schema` queries and their tests are @fpiped's, from
#1 and #2; this keeps the Glue fast path in front of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@aoelvp94
aoelvp94 merged commit 5c7c852 into aoelvp94:athena/part-4-macro-package Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants