Skip to content

fix(CSVDocumentSplitter): report column positions when read_csv_kwargs sets a header - #12968

Closed
sclfcz wants to merge 2 commits into
deepset-ai:mainfrom
sclfcz:fix/csv-splitter-header-kwargs-col-idx
Closed

sclfcz wants to merge 2 commits into
deepset-ai:mainfrom
sclfcz:fix/csv-splitter-header-kwargs-col-idx

Conversation

@sclfcz

@sclfcz sclfcz commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Fixes #12965

What

CSVDocumentSplitter forces header=None so that columns are integer positions, and then reads col_idx_start straight off the column label:

resolved_read_csv_kwargs = {"header": None, "skip_blank_lines": False, "dtype": object, **self.read_csv_kwargs}
...
"col_idx_start": int(split_df.columns[0]),

A caller who passes read_csv_kwargs={"header": 0} (documented as supported, and "infer" behaves the same) gets the first row as string labels, so int("name") raised ValueError for every sub-table and the document could not be split at all:

splitter = CSVDocumentSplitter(row_split_threshold=1, column_split_threshold=None, read_csv_kwargs={"header": 0})
splitter.run([Document(content="name,score\nAda,9\n\nBob,8")])
# ValueError: invalid literal for int() with base 10: 'name'

Change

Map the labels back to their position in the original frame, so col_idx_start keeps meaning "starting column index of the sub-table in the original table" regardless of how the CSV was read:

column_positions = {label: position for position, label in enumerate(df.columns)}
...
"col_idx_start": column_positions[split_df.columns[0]],

Testing

read_csv_kwargs before after
default (header=None) [(0, 0), (3, 0)] [(0, 0), (3, 0)] (unchanged)
{"header": 0} ValueError: invalid literal for int() with base 10: 'name' [(0, 0), (2, 0)]
{"header": "infer"} same ValueError [(0, 0), (2, 0)]

The row indices differ because with header=0 the header row is consumed, so the sub-tables really do start at rows 0 and 2; col_idx_start stays 0.

  • pytest test/components/preprocessors/test_csv_document_splitter_header_kwargs.py → 2 passed
  • reverting only the col_idx_start line → the header test fails and the default case still passes

AI assistance

Written with an AI coding assistant (Claude-based agent): it located the cast, ran the before/after table above and wrote the test. I reviewed the diff, the reproduction output and the reasoning, and I take responsibility for the change.

When every document lacks the ranking field, run() returned the input
documents unchanged, bypassing the configured missing_meta policy: a batch with
one rated document dropped the unrated ones, while an entirely unrated batch
passed through intact even with missing_meta="drop".

Apply the policy in that branch as well (drop -> empty list, top/bottom keep
the documents, as before) and keep the warning that explains the situation.
…s sets a header

The splitter overrides header to None so columns are integer positions and
col_idx_start is read straight from columns[0]. A caller passing
read_csv_kwargs={"header": 0} (or "infer") gets the first row as string
labels, and int(columns[0]) then raised ValueError for every sub-table, so the
document could not be split at all.

Map the labels back to their position in the original frame.
@sclfcz
sclfcz requested a review from a team as a code owner September 26, 2026 01:51
@sclfcz
sclfcz requested review from bogdankostic and removed request for a team September 26, 2026 01:51
@vercel

vercel Bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

@sclfcz is attempting to deploy a commit to the deepset Team on Vercel.

A member of the Team first needs to authorize it.

@github-actions

Copy link
Copy Markdown
Contributor

Hi @sclfcz, thanks for your interest in contributing to Haystack! 🙏

⛔ First-time contributors can have at most 1 open pull request in this repository until it has been approved, so this PR was closed automatically. Your open pull request #12963 is unaffected. Once it has been approved by a maintainer, you are welcome to open more PRs. Feel free to reopen this one at that point.

See the contributing guidelines for details.

This is an automated message to help us keep the review queue healthy.

@sclfcz

sclfcz commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Added a second commit to the same branch, because it is the same label-vs-position confusion in the same method:

The sort key was dataframe.columns[0] — an integer position only while header=None. With a caller-supplied header the sub-tables were ordered alphabetically, so header=0 on z,_,a\n1,,2\n3,,4 (the empty middle column splits it) returned the sub-table at position 2 with split_id=0 and the one at position 0 with split_id=1, i.e. the table's columns came back swapped. It now sorts on the mapped position, exactly like col_idx_start.

read_csv_kwargs before after
default header=None [(0, 0, 0)] unchanged
{"header": 0} [(0, 2, 0), (0, 0, 1)] [(0, 0, 0), (0, 2, 1)]

pytest test/components/preprocessors/test_csv_document_splitter_header_kwargs.py → 3 passed; reverting only the sort key fails the ordering test and leaves the other two passing.

@sclfcz

sclfcz commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Understood - staying within the one-open-PR rule for first-time contributors.

This PR carries a real fix (two commits, both the same label-vs-position confusion) and its tests pass locally:

  • header=0 no longer raises ValueError: int('name') when computing col_idx_start;
  • sub-tables are ordered by column position instead of header label, so z,_,a keeps column 0 before column 2;
  • pytest test/components/preprocessors/test_csv_document_splitter_header_kwargs.py → 3 passed, and reverting either change fails exactly one test.

I will reopen it as soon as #12963 has been reviewed, or let me know if you would rather I fold it into that PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CSVDocumentSplitter crashes when read_csv_kwargs sets header=0

1 participant