Conversation
When every document lacks the ranking field, run() returned the input documents unchanged, bypassing the configured missing_meta policy: a batch with one rated document dropped the unrated ones, while an entirely unrated batch passed through intact even with missing_meta="drop". Apply the policy in that branch as well (drop -> empty list, top/bottom keep the documents, as before) and keep the warning that explains the situation.
…s sets a header
The splitter overrides header to None so columns are integer positions and
col_idx_start is read straight from columns[0]. A caller passing
read_csv_kwargs={"header": 0} (or "infer") gets the first row as string
labels, and int(columns[0]) then raised ValueError for every sub-table, so the
document could not be split at all.
Map the labels back to their position in the original frame.
|
@sclfcz is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
|
Hi @sclfcz, thanks for your interest in contributing to Haystack! 🙏 ⛔ First-time contributors can have at most 1 open pull request in this repository until it has been approved, so this PR was closed automatically. Your open pull request #12963 is unaffected. Once it has been approved by a maintainer, you are welcome to open more PRs. Feel free to reopen this one at that point. See the contributing guidelines for details. This is an automated message to help us keep the review queue healthy. |
|
Added a second commit to the same branch, because it is the same label-vs-position confusion in the same method: The sort key was
|
|
Understood - staying within the one-open-PR rule for first-time contributors. This PR carries a real fix (two commits, both the same label-vs-position confusion) and its tests pass locally:
I will reopen it as soon as #12963 has been reviewed, or let me know if you would rather I fold it into that PR. |
Fixes #12965
What
CSVDocumentSplitterforcesheader=Noneso that columns are integer positions, and then readscol_idx_startstraight off the column label:A caller who passes
read_csv_kwargs={"header": 0}(documented as supported, and"infer"behaves the same) gets the first row as string labels, soint("name")raisedValueErrorfor every sub-table and the document could not be split at all:Change
Map the labels back to their position in the original frame, so
col_idx_startkeeps meaning "starting column index of the sub-table in the original table" regardless of how the CSV was read:Testing
read_csv_kwargsheader=None)[(0, 0), (3, 0)][(0, 0), (3, 0)](unchanged){"header": 0}ValueError: invalid literal for int() with base 10: 'name'[(0, 0), (2, 0)]{"header": "infer"}ValueError[(0, 0), (2, 0)]The row indices differ because with
header=0the header row is consumed, so the sub-tables really do start at rows 0 and 2;col_idx_startstays 0.pytest test/components/preprocessors/test_csv_document_splitter_header_kwargs.py→ 2 passedcol_idx_startline → the header test fails and the default case still passesAI assistance
Written with an AI coding assistant (Claude-based agent): it located the cast, ran the before/after table above and wrote the test. I reviewed the diff, the reproduction output and the reasoning, and I take responsibility for the change.