Taking a raw travel-industry extract and making it fit to model — cleaning, imputation, feature engineering and reshaping.
Written in 2020 during my MSc Data Analytics at London Metropolitan University. The dataset is fabricated; the preparation steps are the point.
999 customer enquiries about holiday packages, across 23 variables. The modelling target is Booked.Status — whether an enquiry converted into a booking.
The script reads Data for cleaning.csv and writes ReadyforModelling.csv, which is the input the companion EDA repository picks up.
- Profiling the raw extract to find what is actually wrong with it
- Handling missing values, including time-series-aware imputation rather than blanket mean-filling
- Parsing and deriving date features
- Type correction and consistent categorical encoding
- Feature engineering aimed at the conversion target
- Producing a clean, reproducible modelling table
The imputation choice is the interesting part: filling a gap in a time-ordered field with the column mean destroys exactly the signal you were hoping to model.
| File | Contents |
|---|---|
DataPreparation.pdf |
The full written walkthrough |
DataPreparation.Rmd |
R Markdown source |
DataPreparation.R |
Script form |
R · lubridate · zoo · imputeTS · DataExplorer · data.table
Rscript install.R # DataExplorer, data.table, imputeTS, lubridate, zoodata for cleaning.csv is not committed — it was fabricated for coursework rather than drawn from a public source. The compiled report carries the full output, and the companion Exploratory-Data-Analysis repository holds ReadyforModelling.csv, the prepared result this script produces.
To run it against your own data, the script expects 999 rows of holiday-package enquiries across 23 variables, including date fields, city and contact-channel columns, and a Booked.Status target.
For current work, see rag-eval-harness and medallion-duckdb.