Senior Data and AI Engineer in Porto Alegre, Brazil (UTC-3).
I build data platforms and LLM systems, and most of my time goes into the part that decides whether they keep working: pipelines that rerun cleanly, retrieval you can actually evaluate, and gates that stop a worse model from shipping.
LinkedIn · email.rodrigopalma@gmail.com
I fix bugs in the tools I use at work. More than 50 pull requests merged in 20 projects, with repeat merges in great-tables, fsspec, Apache Iceberg, Delta Lake, tox, kedro and feast. All of them, merged only: great-tables, fsspec, iceberg-python, delta-rs, tox, kedro, feast, Pillow, redis-py, py-shiny, pint, pyjanitor, apprise, cachetools, nox, onnx, pdm, plotnine, polars and zarr-python. Everything, merged, open and closed: search.
Three that show the kind of bug I go after:
- delta-rs #4747. The doc comment said a
directory counts as a partition when it is named
partitionCol=value, but the code only compared the prefix. With a partition column called_date, an unrelated_dates_backup/directory matched, lost the protection a leading underscore is supposed to give it, and became a deletion candidate forvacuum. - iceberg-python #3995. Rewriting a
predicate to DNF distributed every
ANDover theORs below it with nothing bounding the result. Twenty groups of two branches, about 40 predicates in all and a plausible size for a filter built from user input, expanded to 1,048,576 terms in 16.9s and about 1 GiB. The same input now fails in 0.22s against an explicit limit. - great-tables #869.
date_style="iso"used the CLDR patterny, which has no minimum width, so the year 999 rendered as999-01-05and the package's own ISO parser rejected the package's own output.
Most of them come from two habits: round-tripping a value through the library's own parser, and reading what a function's name promises against what the body actually does.
Data and AI engineer at the public defender's office of Rio Grande do Sul. The system I am responsible for measures excess caseload and allocates more than R$85 million a year in compensation to public defenders. A wrong number there is not a bad chart; it is somebody's pay, so most of the engineering went into making the comparison between units hold up when a unit disputes its own result. I am a co-author on it: the criterion for comparing units across defensorias came from a public defender, not from engineering. Second place in digital innovation at the 2nd CNTI.Def / 5th Enastic, the Brazilian public-defender technology conference, 2026. Around it: ETL for data nobody could query before, and analytics with NLP, all internal.
edgar-rag answers questions over SEC 10-K filings, citing the passage it used or declining. The evaluation was pre-registered: hypotheses, margins and arms fixed before the run, on 300 unanswerable questions across 20 companies. With no gate the model answered 13 of them (4.3%, Wilson 95% [2.5%, 7.3%]), already under the 5-point margin the main hypothesis needed, so it was not attainable on this set and the report says so. A period check plus a cosine gate cuts that to 3 of 300 (1.0%) at half the generation time, 6.59s against 13.21s per question. The service still ships with no gate, because the hypothesis that would have justified one was not met. anchora is the earlier, smaller version of the same idea over Brazilian public law, with a LoRA fine-tune and a 28-question holdout.
quantlens is a quant analyst over B3 (Brazilian exchange) data running entirely on local models. Retrieval with guardrails, an offline eval suite that CI enforces as a gate, and a benchmark that fails the build on regression (the offline path end to end, no LLM call, p50 441 µs). Local models were a deliberate trade: no per-query cost and no data leaving the machine, paid for with weaker generation, which is why the guardrails and the evals exist at all. The decisions behind it are written down as ADRs.
market-elt is market-data ELT on dbt and DuckDB, where data-quality tests break the build instead of letting bad rows travel downstream. Deliberately boring.
About five years as the only developer of my own PC-gaming e-commerce (backend, frontend, payments and the 2am operations), then public-sector systems and data. I started out in psychology. BSc in Computer Engineering, MSc in Big Data and Business Intelligence, Microsoft certified in DP-100 (Azure Data Scientist) and DP-600 (Fabric Analytics Engineer).
Proven in the repositories above: Python, SQL, PyTorch, FastAPI, dbt, DuckDB, Polars, Docker, GitHub Actions, Ollama, LoRA fine-tuning.
Used at work, where the code is not public: Spark, Airflow, Kubernetes, Terraform, AWS, Azure.


