Skip to content

feat(athena): Part 4 — vendor the dbt-athena 1.11.0 macro package - #16369

Open
aoelvp94 wants to merge 11 commits into
dbt-labs:mainfrom
aoelvp94:athena/part-4-macro-package
Open

aoelvp94 wants to merge 11 commits into
dbt-labs:mainfrom
aoelvp94:athena/part-4-macro-package

Conversation

@aoelvp94

@aoelvp94 aoelvp94 commented Sep 18, 2026 •

Copy link
Copy Markdown

Part of #16252. Related: #13822. Independent PR; the sequence is at the bottom.

Problem

Every supported adapter has its Jinja package under crates/dbt-loader/src/dbt_macro_assets/. Athena has none, so parse/ls stop at "missing dbt-athena package" once the profile loads.

Solution

Vendors dbt/include/athena/ from the dbt-athena 1.11.0 wheel as dbt-athena/ (41 files, 2,212 SQL lines; dbt-redshift is 38 files / 1,754 for comparison). Sourcing note in the README: the macros ship in the dbt-athena distribution, not dbt-athena-community, which is an empty shim depending on it.

No loader code change: rust-embed picks the folder up and internal_package_names derives dbt-{adapter_type}. Athena inherits no other adapter's macros, so it needs no entry in the match that gives Redshift dbt-postgres.

Six macros differ from upstream dbt-athena, and one file is left out. Each is marked with a comment at the site:

  • athena__list_schemas and athena__get_columns_in_relation delegated to adapter.list_schemas / adapter.get_columns_in_relation. In dbt-core those are Python methods; in Fusion AdapterImpl::list_schemas runs the Jinja macro of the same name, so the delegation recursed until the Jinja stack overflowed. Both are now SQL over information_schema, which is what the Python did.

  • athena__create_schema called render_hive() on the relation. Athena CREATE SCHEMA is Hive DDL, so it now renders the backticked bare name directly.

  • athena__check_schema_exists added: the default implementation queries a column Athena's information_schema.schemata does not have.

  • athena__get_catalog and athena__get_catalog_relations delegated to adapter.get_catalog / adapter.get_catalog_by_relations, Python methods that read the Glue catalog through boto3. Under Fusion compile --write-catalog (and so docs generate) reported "unknown method: get_catalog_by_relations" per schema and produced an empty catalog.json. Both now query information_schema.tables joined to information_schema.columns, returning the Postgres-shaped rows the metadata adapter already parses. Known limit: a schema-wide information_schema.columns scan fails with ICEBERG_MISSING_METADATA if any table in the schema has unreadable Iceberg metadata, which loses that schema's catalog for the run; the Glue GetTables API does not have that problem and a Glue-backed catalog is the faithful follow-up.

  • resolve_table_type copied catalog_relation.table_format verbatim. Fusion's TableFormat renders default where dbt-athena expects hive, so an unconfigured model resolved to 'default' and table.sql took the else (Iceberg) branch of {% if table_type == 'hive' %}. It now maps the format instead of copying it.

sample_profiles.yml is not vendored: is_metadata_file in load_packages.rs accepts only dbt_project.yml, packages.yml, profile_template.yml and __init__.py in an internal package, and construct_internal_packages asserts everything else lives under macros/ or tests/, so a debug build tripped during dbt parse while the file was present. The profile shape it documented is what the Part 9 wizard generates.

Everything else is byte-identical to 1.11.0, including python_submissions.sql, which Fusion cannot use but which keeps the diff against upstream dbt-athena trivial to re-vendor.

Permission requested: .agents/adapters.md says vendored macros are not to be touched without explicit approval. The six edits above are the minimum for the package to load, for docs to work, and for an unconfigured model to land as Hive under Fusion. If you would rather keep the package byte-identical, the same six can be implemented as Rust arms in AdapterImpl and the macros left as shipped. Say which.

Verification

  • With Parts 1, 2, 3: dbt parse on a 750-model project using dbt_utils, elementary and a root-project override of iceberg_merge: 0 errors, 1 deprecation warning unrelated to Athena. Every macro in the package is parsed; no Jinja construct is rejected. Root-project macro overrides still shadow package macros as in dbt-core.
  • With the later execution parts as well: view, table and incremental materializations run unmodified end to end, and compile --write-catalog produces a populated catalog.json.

Feedback wanted: permission for the six macro edits, or the Rust-arm alternative.

Sequence (independent PRs, each compiles alone against main)

Checklist

  • I have read the contributing guide and understand what's expected of me.
  • I have run this code in development, and it appears to resolve the stated issue.
  • This PR includes tests, or tests are not required or relevant for this PR. (loader macro tests for the modified macros are a pending follow-up on this branch)
  • This PR has no interface changes. (adds the dbt-athena macro namespace)

aoelvp94 and others added 4 commits September 18, 2026 12:47
Adds dbt/include/athena/ from the dbt-athena 1.11.0 wheel as
crates/dbt-loader/src/dbt_macro_assets/dbt-athena — 42 files, 2212 SQL
LOC, in line with the other 14 vendored adapter packages (dbt-redshift
is 38 files / 1754 LOC for comparison).

No loader code change is required. The package is embedded by the
rust-embed `#[folder = "src/dbt_macro_assets/"]` attribute, and
internal_package_names() derives the directory name as
`dbt-{adapter_type}`, which is `dbt-athena` for AdapterType::Athena.
Athena inherits no other adapter's macros, so it needs no entry in the
match that gives Redshift dbt-postgres and Databricks dbt-spark.

Sourcing note recorded in the README: these macros ship in the
`dbt-athena` distribution, not `dbt-athena-community`, which is an empty
shim depending on it.

This is the mechanical half of macro support. Whether Fusion's Jinja
accepts every construct dbt-athena uses is unverified — that is what
`dbt parse` against a real project will answer, and the ~38 adapter.*
methods these macros call remain unimplemented in Rust.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hema Fusion-safe

Three vendored dbt-athena macros delegated to Python-side adapter methods.
In Fusion the Rust methods `AdapterImpl::list_schemas` and
`get_columns_in_relation` execute these very macros, so `dbt run` recursed
until the blocking worker overflowed its stack (minijinja eval_impl <->
execute_macro_with_package <-> list_schemas, seen under lldb).

- `athena__list_schemas`: `information_schema.schemata`, filtered on the
  catalog with the rendered quotes stripped.
- `athena__get_columns_in_relation`: `information_schema.columns` ordered by
  `ordinal_position`, with null length/precision/scale (Trino has none),
  through `sql_convert_columns_in_relation`.
- `athena__create_schema`: Hive DDL takes a bare, backtick-quoted schema
  name; the catalog-qualified double-quoted rendering was rejected by Athena.
  `render_hive()` has no Rust counterpart, nor does
  `add_lf_tags_to_database`, which is skipped.

Verified as root-project overrides against a scratch schema: schema created,
`create or replace view` executed, view readable back. The view
materialization itself still calls `adapter.expire_glue_table_versions`,
which needs a Rust shim; the table materialization stops at
`adapter.build_catalog_relation`.
Fusion's `default__check_schema_exists` calls
`information_schema.replace(information_schema_view='SCHEMATA')`, a method
Fusion relations do not have, so it fails with "map has no method named
replace" for any adapter that relies on the default. Every other adapter
package in `dbt_macro_assets` overrides it; Athena now does too, with a
count over `information_schema.schemata` filtered on the lowercased schema
and catalog (Glue stores both lowercased).

Surfaced by elementary's on-run-start hook, which probes for the legacy
`<schema>__tests` schema on every `dbt test`; with this macro both elementary
hooks complete on Athena.
@aoelvp94
aoelvp94 requested a review from a team as a code owner September 18, 2026 19:54
@cla-bot cla-bot Bot added the cla:yes label Sep 18, 2026
…write-catalog and docs

`athena__get_catalog` and `athena__get_catalog_relations` delegated to
Python-side adapter methods (`adapter.get_catalog`,
`adapter.get_catalog_by_relations`) that read the Glue catalog through
boto3. Fusion's `compile --write-catalog` (and therefore `docs generate`)
runs `get_catalog_relations` per schema and reported "[Non-critical]
unknown method: map has no method named get_catalog_by_relations" for every
schema, so `catalog.json` came out empty and the docs carried no actual
column types.

Both macros now query Trino's `information_schema.tables` joined to
`information_schema.columns`, returning the Postgres-shaped result the
metadata adapter's `build_schemas_from_stats_sql` /
`build_columns_from_get_columns` already parse: one row per column with the
table columns repeated, `column_index` from `ordinal_position` (bigint),
`table_comment` and `table_owner` null since Glue exposes neither there.
Relation filters compare lowercase names, matching Athena's identifier
folding.

Known limit of this approach: Athena fails a schema-wide
`information_schema.columns` scan with ICEBERG_MISSING_METADATA when any
table in that schema has unreadable Iceberg metadata, which loses that
schema's catalog for the run (reported as non-critical, as before). The
Glue GetTables API does not have that problem; a Glue-backed catalog would
be the faithful port of dbt-athena's behaviour and is the follow-up.

@fpiped fpiped left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested with the rest of the stack (and Part 6) against a real workgroup. Two things:

  1. crates/dbt-loader/src/dbt_macro_assets/dbt-athena/sample_profiles.yml: the loader accepts only dbt_project.yml, packages.yml, profile_template.yml and __init__.py in an internal package (load_packages.rs, is_metadata_file), so dbt parse panics in debug builds while this file is present. The sample profile is covered by the Part 9 wizard.
  2. resolve_table_type copies catalog_relation.table_format verbatim. Fusion's TableFormat renders default, not hive, so an unconfigured model resolves to 'default' and table.sql takes the else (Iceberg) branch of {% if table_type == 'hive' %}. {%- set table_type = 'iceberg' if catalog_relation.table_format == 'iceberg' else 'hive' -%} fixes it.

…table_type

Two problems found running the vendored package against a real workgroup.

`sample_profiles.yml` is not an accepted file in an internal package:
`is_metadata_file` (`load_packages.rs`) allows only `dbt_project.yml`,
`packages.yml`, `profile_template.yml` and `__init__.py`, and
`construct_internal_packages` assumes everything else lives under `macros/`
or `tests/`, so a debug build trips its assertion during `dbt parse`. The
file only documents the profile shape, which `dbt init` now generates.

`resolve_table_type` copied `catalog_relation.table_format` verbatim.
Fusion's `TableFormat` renders `default` where dbt-athena expects `hive`,
so an unconfigured model resolved to 'default' and `table.sql` took the
`else` branch of `{% if table_type == 'hive' %}`, creating an Iceberg table
where Hive was meant. Map the format instead of copying it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment on lines +9 to +10
{%- set ns.bucket_column = bucket_match[1] -%}
{%- set bucket_num = adapter.murmur3_hash(col, bucket_match[2] | int) -%}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Match objects can't be indexed in dbt v2 (#16451), so these render empty and bucketed Iceberg tables fail. .group() works today:

Suggested change
{%- set ns.bucket_column = bucket_match[1] -%}
{%- set bucket_num = adapter.murmur3_hash(col, bucket_match[2] | int) -%}
{%- set ns.bucket_column = bucket_match.group(1) -%}
{%- set bucket_num = adapter.murmur3_hash(col, bucket_match.group(2) | int) -%}

@fpiped

fpiped commented Sep 25, 2026

Copy link
Copy Markdown

athena__create_schema skips adapter.add_lf_tags_to_database(relation), which dbt-athena calls after the DDL. #16376 now implements it, so the call can come back:

  {%- call statement('create_schema') -%}
    create schema if not exists `{{ relation.schema }}`
  {% endcall %}
  {{ adapter.add_lf_tags_to_database(relation) }}

Checked on a real work group: the profile's lf_tags_database lands on the schema the run creates.

The vendored package predates dbt-adapters 62a5495cf (S3 Tables catalogs).
From that commit, applied as is: drop_relation drops an S3 Tables relation
through Glue; the table materialization drops and recreates it, since S3
Tables allows no ALTER TABLE RENAME; create_table_as sets no S3 location and
skips the direct S3 cleanup for it; incremental skips the Glue table version
expiry, which the federated catalog does not serve.
dbt v2's regex match objects do not support indexing (dbt-labs#16451), so
bucket_match[1] rendered empty; .group(n) works in dbt v1 and v2.
As dbt-athena does after the DDL; adapter.add_lf_tags_to_database is
implemented in dbt-labs#16376.
Athena's information_schema covers only the catalog it is read from, so the
columns of a relation in another catalog (S3 Tables, s3tablescatalog/<bucket>)
came back empty.
@fpiped

fpiped commented Sep 25, 2026

Copy link
Copy Markdown

Fixes from the end-to-end runs, as a PR to this branch: aoelvp94#1 (S3 Tables macros from dbt-adapters 62a5495cf, bucket_match.group(), the add_lf_tags_to_database call, catalog-qualified information_schema.columns).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants