Roadmap: The future of csv-parser, CSVStat deprecation, and the csvzall project #304
vincentlaucsb
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
AI Disclosure: I wrote all of this myself. :)
Hello everybody. I'd like to thank everybody for their detailed PRs and their continued support, whether by emailing me kind comments (and asking how they may cite the project in their PhD dissertation) or adding celebratory emojis on new releases.
I am humbled to see a project I wrote in college to parse CSVs has taken off.
I have revived this project after a long hiatus because CSVs continue to be the format that won't die, and they are still relevant in 2026. I apologize to everybody who left detailed Issues or Pull Requests that I left hanging.
GitHub Actions & the Single Header Format
One of the reasons (not an excuse) for leaving quality PRs hanging was that I was worried they would introduce a random segfault without me knowing. As a result, I have spent (and will continue to spend) considerable effort revamping the automated GitHub Workflows pipeline and test suite to ensure new PRs are adequate tested, so I can merge them quickly and with confidence.
The extensive sanitization suite identified latent memory and thread bugs which I have patched as of 3.0.0.
Single Header
I decided to make the single header a release asset + published on my GitHub Pages to simplify PRs. Many times, PR commiters would forget to update the single header file.
Personally, I understand, because I forgot to update the single header file on many occasions myself. I think this is a reasonable move to reduce:
With that being said, if this move has seriously inconvenienced anybody, please let me know.
AI Guidelines
Obviously, like many SWEs in the year 2026, I have grown accustomed to using LLM-based coding agents to help write improvements. Like many, I have seen they are both very powerful, yet at the same time, capable of writing lots of slop. For example, Copilot has been very helpful in fixing many compiler-specific bugs, but I've also had to manually refactor a lot of AI code since it tends to like writing macro soup, or adding locks to every hot path for thread "safety".
The test suites I have added should fish out any logic or safety bugs introduced by new code, whether human-written, AI-generated, or a combination of both.
At the end of the day, I will judge all PRs based on whether or not they add to the library, and whether or not they pass the tests. That's it.
csvzall: The High Performance Data Science CSV Command Line Toolkit
csvzall is a new project I am working on that will be both a general purpose C++ library and command line program for managing CSVs.
csvzall= CSV + SawZall. The SawZall is a versatile tool, not because it can do anything and everything, but because the concept of a blade with an orbital action accomplishes a lot of common construction/demolition tasks. My philosophy forcsvzallis similar. It will be a versatile tool, not because it implements 67 different commands (like other CSV command line tools), but because it's a efficient SQLite/Postgres dumping tool.It is not a replacement for csv-parser, but rather a library built on top of it. My main motivation for building it was so I could graph my gym logs and Garmin data in Obsidian. It's main goal is to be a high-level "batteries-included" CSV library for common ETL and data science workflows. (My 2nd motivation is to be able to play around with Census data w/o waiting for Pandas to load).
Currently, I have implemented a few basic transforms as well as a CSV -> SQLite loader. In the future, I plan to add:
Many frictions I experienced developing csvzall have already motivated
csv-parserimprovements, such as yesterday'sparse_unsafe()function which allows handling piped CSVs fromstdinwithout unnecessary copying.CSVStat: Likely to be deprecated soon
CSVStatwas a class I used to calculate add column-parallel streaming statistics. This allowed you to calculate mean, variance, and schema information in a streaming manner.The concept is valuable, but I will be moving it to
csvzallin the future because:csv-parserusers to include code to calculate statistics that they may or may not needcsv-parser, will only add more code that most users won't want or needCSVReaderthat accepts arbitrary functions which better threading behavior.DataFrame: Here to stay
I debated adding the
DataFrameclass because I didn't want to add scope creep or create another half-baked Pandas clone.However, the
DataFramesolves a lot of high-level problems:CSVReader's streaming nature was good for memory inefficiency, but was annoying for many medium-sized CSVs that could easily fit in memoryMy goal for the
DataFrameis:Conclusion: why I keep working on this
Besides wanting to build out
csv-parserso I have a strong foundation for thecsvzallapp, my motivation for building this library is the same as when I started. Most C++ libraries force you to choose between performance, correctness, or usability.You either have to between:
malloc()pressure due to pushing everything into astd::vector<std::string>My goal for this parser is to not force you to choose. Recently, the 3.3.0 version was able to achieve 550MB/s on my 2022 Intel Core i5 desktop processor, all while delivering a nice user-friendly API that C++ programmers deserve.
Addendum: Near-term items (just a brain dump)
std::metaallows for doing this in a zero-cost manner. The main issue right now is lack of compiler supportcsv-parserwork with common ZIP and GZIP librariesCSVWriteroptimizations: Currently CSVRow and DataFrame objects can be serialized to CSV viaCSVWriter << vector<string>. This works fine, but incurs unnecessary memory allocations as mentioned earlier. I will be adding more direct pipelines in the near future.csv-parserworks on most modern compilers without spamming warnings and errorsFeel free to reply with your own ideas or pain points.
Thanks again to everybody for your support.
All reactions