Skip to content

fix: keep hyperlink addresses in DOCXToDocument table cells - #12988

Merged
sjrl merged 1 commit into
deepset-ai:mainfrom
Jayanth-reflex:claude/brave-clarke-wuu6la
Sep 28, 2026
Merged

sjrl merged 1 commit into
deepset-ai:mainfrom
Jayanth-reflex:claude/brave-clarke-wuu6la

Conversation

@Jayanth-reflex

Copy link
Copy Markdown
Contributor

Related Issues

Proposed Changes:

DOCXToDocument formats links in body paragraphs through _process_links_in_paragraph(), but tables were built from cell.text, which python-docx fills with each paragraph's display text only. So with link_format="markdown" or "plain", a link in a table cell lost its address.

A new _cell_text() helper passes each cell paragraph through _process_links_in_paragraph() and joins them with newlines, the same way cell.text does. _table_to_markdown() and _table_to_csv() now use it. With the default link_format="none" the output is unchanged, because that path returns paragraph.text.

How did you test it?

  • New test_link_extraction_in_table, covering both table formats (markdown, csv) and both link formats that output addresses (markdown, plain). All 4 cases fail without the fix and pass with it.
  • test_docx_file_to_document.py: 36 passed
  • All converter unit tests: 305 passed
  • hatch run fmt and mypy on the changed files: clean

Notes for the reviewer

  • Markdown pipe and line-break escaping still runs on the formatted cell text, so a link can't break a table row.
  • Nested tables inside a cell are still skipped, as before.

This PR was fully generated with an AI assistant. I have reviewed the changes and run the relevant tests.

Checklist

Table cells were serialized from cell.text, which holds only the display
text of a hyperlink, so link_format="markdown" or "plain" kept addresses
in body paragraphs but dropped them in tables. Cell paragraphs now go
through the same link handling as body paragraphs, in both the markdown
and csv table formats. link_format="none" output is unchanged.

Fixes deepset-ai#12977

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PQv3cLuiTKPtLNj9i1cma
Signed-off-by: Jayanth <jayanthreddy268.jr@gmail.com>
@Jayanth-reflex
Jayanth-reflex requested a review from a team as a code owner September 27, 2026 15:51
@Jayanth-reflex
Jayanth-reflex requested review from sjrl and removed request for a team September 27, 2026 15:51
@vercel

vercel Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

@Jayanth-reflex is attempting to deploy a commit to the deepset Team on Vercel.

A member of the Team first needs to authorize it.

@CLAassistant

CLAassistant commented Sep 27, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@Jayanth-reflex

Copy link
Copy Markdown
Contributor Author

@sjrl , could you please review and auth the workflows

@github-actions github-actions Bot added the type:documentation Improvements on the docs label Sep 28, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Coverage report

Click to see where and how coverage changed

FileStatementsMissingCoverageCoverage
(new stmts)
Lines missing
  haystack/components/converters
  docx.py
  haystack/dataclasses
  chat_message.py
  haystack/document_stores/in_memory
  document_store.py
Project Total  

This report was generated by python-coverage-comment-action

@sjrl sjrl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@sjrl
sjrl merged commit 3153560 into deepset-ai:main Sep 28, 2026
26 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

topic:tests type:documentation Improvements on the docs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DOCXToDocument drops hyperlink addresses inside tables

3 participants