Skip to content

Fix: Decode null characters in old MAT formats - #58

Merged
foreverallama merged 1 commit into
mainfrom
char-fixes-5
Mar 16, 2026
Merged

foreverallama merged 1 commit into
mainfrom
char-fixes-5

Conversation

@foreverallama

@foreverallama foreverallama commented Mar 16, 2026 •

Copy link
Copy Markdown
Owner

Previously, we were filtering out 0x00 bytes from the uint16 utf-8 encoded char array in legacy MAT-file formats. This resulted in null characters being stripped from the original char array, if any.

This PR fixes the bug by constructing a new byte array from (msb, lsb) of each uint16 integer. The MSB is included only if it's non-zero, but the LSB is always included. This ensures an accurate UTF-8 encoded buffer that can be appropriately decoded.

Note: Trailing null characters are still removed by NumPy, however the array itemsize or dtype can be used to infer the number of trailing null characters.

@foreverallama
foreverallama merged commit cd51391 into main Mar 16, 2026
3 checks passed
@foreverallama
foreverallama deleted the char-fixes-5 branch March 16, 2026 07:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant