Conversation
Adds dbt/include/athena/ from the dbt-athena 1.11.0 wheel as
crates/dbt-loader/src/dbt_macro_assets/dbt-athena — 42 files, 2212 SQL
LOC, in line with the other 14 vendored adapter packages (dbt-redshift
is 38 files / 1754 LOC for comparison).
No loader code change is required. The package is embedded by the
rust-embed `#[folder = "src/dbt_macro_assets/"]` attribute, and
internal_package_names() derives the directory name as
`dbt-{adapter_type}`, which is `dbt-athena` for AdapterType::Athena.
Athena inherits no other adapter's macros, so it needs no entry in the
match that gives Redshift dbt-postgres and Databricks dbt-spark.
Sourcing note recorded in the README: these macros ship in the
`dbt-athena` distribution, not `dbt-athena-community`, which is an empty
shim depending on it.
This is the mechanical half of macro support. Whether Fusion's Jinja
accepts every construct dbt-athena uses is unverified — that is what
`dbt parse` against a real project will answer, and the ~38 adapter.*
methods these macros call remain unimplemented in Rust.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hema Fusion-safe Three vendored dbt-athena macros delegated to Python-side adapter methods. In Fusion the Rust methods `AdapterImpl::list_schemas` and `get_columns_in_relation` execute these very macros, so `dbt run` recursed until the blocking worker overflowed its stack (minijinja eval_impl <-> execute_macro_with_package <-> list_schemas, seen under lldb). - `athena__list_schemas`: `information_schema.schemata`, filtered on the catalog with the rendered quotes stripped. - `athena__get_columns_in_relation`: `information_schema.columns` ordered by `ordinal_position`, with null length/precision/scale (Trino has none), through `sql_convert_columns_in_relation`. - `athena__create_schema`: Hive DDL takes a bare, backtick-quoted schema name; the catalog-qualified double-quoted rendering was rejected by Athena. `render_hive()` has no Rust counterpart, nor does `add_lf_tags_to_database`, which is skipped. Verified as root-project overrides against a scratch schema: schema created, `create or replace view` executed, view readable back. The view materialization itself still calls `adapter.expire_glue_table_versions`, which needs a Rust shim; the table materialization stops at `adapter.build_catalog_relation`.
Fusion's `default__check_schema_exists` calls `information_schema.replace(information_schema_view='SCHEMATA')`, a method Fusion relations do not have, so it fails with "map has no method named replace" for any adapter that relies on the default. Every other adapter package in `dbt_macro_assets` overrides it; Athena now does too, with a count over `information_schema.schemata` filtered on the lowercased schema and catalog (Glue stores both lowercased). Surfaced by elementary's on-run-start hook, which probes for the legacy `<schema>__tests` schema on every `dbt test`; with this macro both elementary hooks complete on Athena.
…write-catalog and docs `athena__get_catalog` and `athena__get_catalog_relations` delegated to Python-side adapter methods (`adapter.get_catalog`, `adapter.get_catalog_by_relations`) that read the Glue catalog through boto3. Fusion's `compile --write-catalog` (and therefore `docs generate`) runs `get_catalog_relations` per schema and reported "[Non-critical] unknown method: map has no method named get_catalog_by_relations" for every schema, so `catalog.json` came out empty and the docs carried no actual column types. Both macros now query Trino's `information_schema.tables` joined to `information_schema.columns`, returning the Postgres-shaped result the metadata adapter's `build_schemas_from_stats_sql` / `build_columns_from_get_columns` already parse: one row per column with the table columns repeated, `column_index` from `ordinal_position` (bigint), `table_comment` and `table_owner` null since Glue exposes neither there. Relation filters compare lowercase names, matching Athena's identifier folding. Known limit of this approach: Athena fails a schema-wide `information_schema.columns` scan with ICEBERG_MISSING_METADATA when any table in that schema has unreadable Iceberg metadata, which loses that schema's catalog for the run (reported as non-critical, as before). The Glue GetTables API does not have that problem; a Glue-backed catalog would be the faithful port of dbt-athena's behaviour and is the follow-up.
fpiped
left a comment
There was a problem hiding this comment.
Tested with the rest of the stack (and Part 6) against a real workgroup. Two things:
crates/dbt-loader/src/dbt_macro_assets/dbt-athena/sample_profiles.yml: the loader accepts onlydbt_project.yml,packages.yml,profile_template.ymland__init__.pyin an internal package (load_packages.rs,is_metadata_file), sodbt parsepanics in debug builds while this file is present. The sample profile is covered by the Part 9 wizard.resolve_table_typecopiescatalog_relation.table_formatverbatim. Fusion'sTableFormatrendersdefault, nothive, so an unconfigured model resolves to'default'andtable.sqltakes theelse(Iceberg) branch of{% if table_type == 'hive' %}.{%- set table_type = 'iceberg' if catalog_relation.table_format == 'iceberg' else 'hive' -%}fixes it.
…table_type
Two problems found running the vendored package against a real workgroup.
`sample_profiles.yml` is not an accepted file in an internal package:
`is_metadata_file` (`load_packages.rs`) allows only `dbt_project.yml`,
`packages.yml`, `profile_template.yml` and `__init__.py`, and
`construct_internal_packages` assumes everything else lives under `macros/`
or `tests/`, so a debug build trips its assertion during `dbt parse`. The
file only documents the profile shape, which `dbt init` now generates.
`resolve_table_type` copied `catalog_relation.table_format` verbatim.
Fusion's `TableFormat` renders `default` where dbt-athena expects `hive`,
so an unconfigured model resolved to 'default' and `table.sql` took the
`else` branch of `{% if table_type == 'hive' %}`, creating an Iceberg table
where Hive was meant. Map the format instead of copying it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| {%- set ns.bucket_column = bucket_match[1] -%} | ||
| {%- set bucket_num = adapter.murmur3_hash(col, bucket_match[2] | int) -%} |
There was a problem hiding this comment.
Match objects can't be indexed in dbt v2 (#16451), so these render empty and bucketed Iceberg tables fail. .group() works today:
| {%- set ns.bucket_column = bucket_match[1] -%} | |
| {%- set bucket_num = adapter.murmur3_hash(col, bucket_match[2] | int) -%} | |
| {%- set ns.bucket_column = bucket_match.group(1) -%} | |
| {%- set bucket_num = adapter.murmur3_hash(col, bucket_match.group(2) | int) -%} |
|
{%- call statement('create_schema') -%}
create schema if not exists `{{ relation.schema }}`
{% endcall %}
{{ adapter.add_lf_tags_to_database(relation) }}Checked on a real work group: the profile's |
The vendored package predates dbt-adapters 62a5495cf (S3 Tables catalogs). From that commit, applied as is: drop_relation drops an S3 Tables relation through Glue; the table materialization drops and recreates it, since S3 Tables allows no ALTER TABLE RENAME; create_table_as sets no S3 location and skips the direct S3 cleanup for it; incremental skips the Glue table version expiry, which the federated catalog does not serve.
dbt v2's regex match objects do not support indexing (dbt-labs#16451), so bucket_match[1] rendered empty; .group(n) works in dbt v1 and v2.
As dbt-athena does after the DDL; adapter.add_lf_tags_to_database is implemented in dbt-labs#16376.
Athena's information_schema covers only the catalog it is read from, so the columns of a relation in another catalog (S3 Tables, s3tablescatalog/<bucket>) came back empty.
|
Fixes from the end-to-end runs, as a PR to this branch: aoelvp94#1 (S3 Tables macros from dbt-adapters |
Part of #16252. Related: #13822. Independent PR; the sequence is at the bottom.
Problem
Every supported adapter has its Jinja package under
crates/dbt-loader/src/dbt_macro_assets/. Athena has none, soparse/lsstop at "missingdbt-athenapackage" once the profile loads.Solution
Vendors
dbt/include/athena/from the dbt-athena 1.11.0 wheel asdbt-athena/(41 files, 2,212 SQL lines; dbt-redshift is 38 files / 1,754 for comparison). Sourcing note in the README: the macros ship in thedbt-athenadistribution, notdbt-athena-community, which is an empty shim depending on it.No loader code change:
rust-embedpicks the folder up andinternal_package_namesderivesdbt-{adapter_type}. Athena inherits no other adapter's macros, so it needs no entry in the match that gives Redshiftdbt-postgres.Six macros differ from upstream dbt-athena, and one file is left out. Each is marked with a comment at the site:
athena__list_schemasandathena__get_columns_in_relationdelegated toadapter.list_schemas/adapter.get_columns_in_relation. In dbt-core those are Python methods; in FusionAdapterImpl::list_schemasruns the Jinja macro of the same name, so the delegation recursed until the Jinja stack overflowed. Both are now SQL overinformation_schema, which is what the Python did.athena__create_schemacalledrender_hive()on the relation. AthenaCREATE SCHEMAis Hive DDL, so it now renders the backticked bare name directly.athena__check_schema_existsadded: the default implementation queries a column Athena'sinformation_schema.schematadoes not have.athena__get_catalogandathena__get_catalog_relationsdelegated toadapter.get_catalog/adapter.get_catalog_by_relations, Python methods that read the Glue catalog through boto3. Under Fusioncompile --write-catalog(and sodocs generate) reported "unknown method: get_catalog_by_relations" per schema and produced an emptycatalog.json. Both now queryinformation_schema.tablesjoined toinformation_schema.columns, returning the Postgres-shaped rows the metadata adapter already parses. Known limit: a schema-wideinformation_schema.columnsscan fails withICEBERG_MISSING_METADATAif any table in the schema has unreadable Iceberg metadata, which loses that schema's catalog for the run; the GlueGetTablesAPI does not have that problem and a Glue-backed catalog is the faithful follow-up.resolve_table_typecopiedcatalog_relation.table_formatverbatim. Fusion'sTableFormatrendersdefaultwhere dbt-athena expectshive, so an unconfigured model resolved to'default'andtable.sqltook theelse(Iceberg) branch of{% if table_type == 'hive' %}. It now maps the format instead of copying it.sample_profiles.ymlis not vendored:is_metadata_fileinload_packages.rsaccepts onlydbt_project.yml,packages.yml,profile_template.ymland__init__.pyin an internal package, andconstruct_internal_packagesasserts everything else lives undermacros/ortests/, so a debug build tripped duringdbt parsewhile the file was present. The profile shape it documented is what the Part 9 wizard generates.Everything else is byte-identical to 1.11.0, including
python_submissions.sql, which Fusion cannot use but which keeps the diff against upstream dbt-athena trivial to re-vendor.Permission requested:
.agents/adapters.mdsays vendored macros are not to be touched without explicit approval. The six edits above are the minimum for the package to load, for docs to work, and for an unconfigured model to land as Hive under Fusion. If you would rather keep the package byte-identical, the same six can be implemented as Rust arms inAdapterImpland the macros left as shipped. Say which.Verification
dbt parseon a 750-model project usingdbt_utils,elementaryand a root-project override oficeberg_merge: 0 errors, 1 deprecation warning unrelated to Athena. Every macro in the package is parsed; no Jinja construct is rejected. Root-project macro overrides still shadow package macros as in dbt-core.compile --write-catalogproduces a populatedcatalog.json.Feedback wanted: permission for the six macro edits, or the Rust-arm alternative.
Sequence (independent PRs, each compiles alone against main)
table/incrementalhelpers (athena/parts-6-7-execution); see also feat(athena): Part 6 — AthenaAdapter methods behind the dbt-athena macros #16376Checklist
dbt-athenamacro namespace)