Personal portfolio showcasing my work across data science, machine learning, statistical modeling, algorithms, and deployed data products.
Live site: https://pedrohgp02.github.io/
An end-to-end public-health data product that ingests official Brazilian surveillance data, evaluates forecasting models using time-aware backtesting, and serves interactive forecasts through Streamlit.
A personalized machine-learning study using 92,445 listening events to compare audio models, Audio Spectrogram Transformers, and behavioral context for skip prediction.
Hierarchical Bayesian modeling of Argentine football attendance using PyMC, Negative Binomial models, PSIS-LOO, posterior predictive checks, and missing-data estimation.
- Dynamic programming for reconstructing genealogical relationships from DNA sequences
- Constraint-aware task scheduling with heaps and dependency handling
- Cellular automata and Monte Carlo simulation for wildfire intervention analysis
- AI memory-support prototype built at UC Berkeley Hack for Impact
Python R SQL PyTorch scikit-learn PyMC pandas NumPy Streamlit
- Portfolio: https://pedrohgp02.github.io/
- GitHub: https://github.com/pedrohgp02
- LinkedIn: https://www.linkedin.com/in/pedrohgpaiva/
- Writing: https://oodnotes.substack.com/