fix: preserve extra_columns field on ragged CSV rows - #13001
Closed
ansh-rohilla wants to merge 1 commit into
Closed
ansh-rohilla wants to merge 1 commit into
ansh-rohilla wants to merge 1 commit into
Conversation
CSVToDocument in row mode previously passed restkey="extra_columns" to csv.DictReader. When a CSV file had an actual header column named "extra_columns", a ragged row with surplus columns wrote overflow values into row["extra_columns"], silently discarding the real cell value. This change: - Uses a private sentinel (_RAGGED_ROW_RESTKEY) as csv.DictReader restkey so it cannot collide with any CSV header. - Maps the sentinel to "extra_columns" during metadata merging, routing through existing collision resolution (prefixed as csv_extra_columns if extra_columns already exists in row_meta). - Adds regression test in test_csv_to_document.py. - Adds reno release note. Closes deepset-ai#12994
Contributor
|
@ansh-rohilla is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
Member
|
Thanks for the contribution but I am closing this PR. See #12994 (comment) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related Issues
Proposed Changes:
In row mode,
CSVToDocumentpreviously configuredcsv.DictReaderwithrestkey="extra_columns". If the CSV file contained an actual header column namedextra_columns, ragged rows with surplus fields overwrote the cell's actual value with the overflow list.This PR:
_RAGGED_ROW_RESTKEY = object()) forDictReader.restkeyso it never collides with any CSV header."extra_columns"during metadata construction, ensuring that if a realextra_columnsfield exists, the surplus values are placed undercsv_extra_columnsusing existing collision resolution.test_row_mode_ragged_row_with_real_extra_columns_column.How did I test it?
test/components/converters/test_csv_to_document.py.18 passed),reno lint, andruff check/ruff format(all passed).Checklist
fix:.