Skip to content

Repository files navigation

Data preparation in R

Taking a raw travel-industry extract and making it fit to model — cleaning, imputation, feature engineering and reshaping.

Written in 2020 during my MSc Data Analytics at London Metropolitan University. The dataset is fabricated; the preparation steps are the point.

The dataset

999 customer enquiries about holiday packages, across 23 variables. The modelling target is Booked.Status — whether an enquiry converted into a booking.

The script reads Data for cleaning.csv and writes ReadyforModelling.csv, which is the input the companion EDA repository picks up.

What it covers

  • Profiling the raw extract to find what is actually wrong with it
  • Handling missing values, including time-series-aware imputation rather than blanket mean-filling
  • Parsing and deriving date features
  • Type correction and consistent categorical encoding
  • Feature engineering aimed at the conversion target
  • Producing a clean, reproducible modelling table

The imputation choice is the interesting part: filling a gap in a time-ordered field with the column mean destroys exactly the signal you were hoping to model.

Files

File Contents
DataPreparation.pdf The full written walkthrough
DataPreparation.Rmd R Markdown source
DataPreparation.R Script form

Stack

R · lubridate · zoo · imputeTS · DataExplorer · data.table

Running it

Rscript install.R      # DataExplorer, data.table, imputeTS, lubridate, zoo

Data

data for cleaning.csv is not committed — it was fabricated for coursework rather than drawn from a public source. The compiled report carries the full output, and the companion Exploratory-Data-Analysis repository holds ReadyforModelling.csv, the prepared result this script produces.

To run it against your own data, the script expects 999 rows of holiday-package enquiries across 23 variables, including date fields, city and contact-channel columns, and a Booked.Status target.


For current work, see rag-eval-harness and medallion-duckdb.

About

Cleaning, imputation and feature engineering in R, turning a raw travel-industry extract into a modelling table.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages