Context
src/erdify/parser.py (~1037 LOC) holds a single ASTDatabaseParser class with ~45 methods that mixes the parsing logic for all supported ORM dialects:
- Django —
_inherits_django_model, _extract_django_fields, _parse_django_fk_column, _parse_django_relationship, _django_*
- SQLAlchemy Core —
_parse_core_association_tables, _build_core_table_entity, _parse_core_column
- SQLAlchemy ORM /
Mapped — _has_mapped_field, _is_mapped_annotation, _unwrap_mapped
- Pydantic / dataclass —
_inherits_basemodel, _is_dataclass, _parse_field
_classify_source already routes a class to its dialect — a natural seam for a Strategy/Dialect pattern.
Proposal
Extract one DialectParser per ORM behind a small interface, selected by _classify_source. The top-level ASTDatabaseParser keeps file discovery, exclude handling, and orchestration.
Benefits:
- Clear boundaries + isolated, per-dialect tests
- Easier to add a new dialect without touching the others
- Smaller, focused units
Priority
Low / deferred. The current class is large but cohesive and well-namespaced (_-prefixed internals). No new dialect is planned right now. This is worth doing when the next ORM dialect lands or a larger parser change is needed — not on spec. Documenting now so the seam isn't lost.
The rest of the codebase (generator.py, cli.py, config.py) is well structured and out of scope.
Context
src/erdify/parser.py(~1037 LOC) holds a singleASTDatabaseParserclass with ~45 methods that mixes the parsing logic for all supported ORM dialects:_inherits_django_model,_extract_django_fields,_parse_django_fk_column,_parse_django_relationship,_django_*_parse_core_association_tables,_build_core_table_entity,_parse_core_columnMapped—_has_mapped_field,_is_mapped_annotation,_unwrap_mapped_inherits_basemodel,_is_dataclass,_parse_field_classify_sourcealready routes a class to its dialect — a natural seam for a Strategy/Dialect pattern.Proposal
Extract one
DialectParserper ORM behind a small interface, selected by_classify_source. The top-levelASTDatabaseParserkeeps file discovery, exclude handling, and orchestration.Benefits:
Priority
Low / deferred. The current class is large but cohesive and well-namespaced (
_-prefixed internals). No new dialect is planned right now. This is worth doing when the next ORM dialect lands or a larger parser change is needed — not on spec. Documenting now so the seam isn't lost.The rest of the codebase (
generator.py,cli.py,config.py) is well structured and out of scope.