Skip to content

Refactor: split ASTDatabaseParser into per-dialect strategy classes #81

Description

@FabianClemenz

Context

src/erdify/parser.py (~1037 LOC) holds a single ASTDatabaseParser class with ~45 methods that mixes the parsing logic for all supported ORM dialects:

  • Django — _inherits_django_model, _extract_django_fields, _parse_django_fk_column, _parse_django_relationship, _django_*
  • SQLAlchemy Core — _parse_core_association_tables, _build_core_table_entity, _parse_core_column
  • SQLAlchemy ORM / Mapped — _has_mapped_field, _is_mapped_annotation, _unwrap_mapped
  • Pydantic / dataclass — _inherits_basemodel, _is_dataclass, _parse_field

_classify_source already routes a class to its dialect — a natural seam for a Strategy/Dialect pattern.

Proposal

Extract one DialectParser per ORM behind a small interface, selected by _classify_source. The top-level ASTDatabaseParser keeps file discovery, exclude handling, and orchestration.

Benefits:

  • Clear boundaries + isolated, per-dialect tests
  • Easier to add a new dialect without touching the others
  • Smaller, focused units

Priority

Low / deferred. The current class is large but cohesive and well-namespaced (_-prefixed internals). No new dialect is planned right now. This is worth doing when the next ORM dialect lands or a larger parser change is needed — not on spec. Documenting now so the seam isn't lost.

The rest of the codebase (generator.py, cli.py, config.py) is well structured and out of scope.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    deferredParked by design; revisit when a real need shows upenhancementNew feature or requestrefactorCode refactoring / tech debt

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions