Production-quality baseline models for HWIN-Bench evaluation.
This repository contains only classical machine learning baselines for evaluating HWIN-Bench. No deep learning, no HWIN-Net. Pure scikit-learn, XGBoost, LightGBM, CatBoost.
| Model | File |
|---|---|
| Linear Regression | models/train_linear.py |
| Ridge Regression | models/train_ridge.py |
| Lasso Regression | models/train_lasso.py |
| Elastic Net | models/train_elasticnet.py |
| Random Forest | models/train_random_forest.py |
| Extra Trees | models/train_extra_trees.py |
| XGBoost | models/train_xgboost.py |
| LightGBM | models/train_lightgbm.py |
| CatBoost | models/train_catboost.py |
- Python 3.10+
- See
requirements.txt
git clone https://github.com/Zakariya-Q/hwin-baselines.git
cd hwin-baselines
pip install -r requirements.txtExpects HWIN-Bench data in data/:
data/
train.parquet
val.parquet
test.parquet
- Target column:
y - No preprocessing, no feature engineering, no normalization
- Load with Polars, convert to numpy for training
Run all baselines sequentially:
bash scripts/run_all.shOr run individual models:
# Linear Regression
python models/train_linear.py --config configs/default.yaml
# Ridge Regression
python models/train_ridge.py --config configs/default.yaml
# Lasso Regression
python models/train_lasso.py --config configs/default.yaml
# Elastic Net
python models/train_elasticnet.py --config configs/default.yaml
# Random Forest
python models/train_random_forest.py --config configs/default.yaml
# Extra Trees
python models/train_extra_trees.py --config configs/default.yaml
# XGBoost
python models/train_xgboost.py --config configs/default.yaml
# LightGBM
python models/train_lightgbm.py --config configs/default.yaml
# CatBoost
python models/train_catboost.py --config configs/default.yamlEdit configs/default.yaml:
data:
data_dir: data
train_file: train.parquet
val_file: val.parquet
test_file: test.parquet
target_column: y
output:
dir: outputs
seed: 42
models:
linear:
fit_intercept: true
ridge:
alpha: 1.0
lasso:
alpha: 1.0
elasticnet:
alpha: 1.0
l1_ratio: 0.5
random_forest:
n_estimators: 200
max_depth: null
min_samples_split: 2
min_samples_leaf: 1
n_jobs: -1
random_state: 42
extra_trees:
n_estimators: 200
max_depth: null
min_samples_split: 2
min_samples_leaf: 1
n_jobs: -1
random_state: 42
xgboost:
n_estimators: 500
max_depth: 6
learning_rate: 0.1
subsample: 0.8
colsample_bytree: 0.8
n_jobs: -1
random_state: 42
tree_method: hist
lightgbm:
n_estimators: 500
max_depth: 6
learning_rate: 0.1
subsample: 0.8
colsample_bytree: 0.8
n_jobs: -1
random_state: 42
verbose: -1
catboost:
iterations: 500
depth: 6
learning_rate: 0.1
l2_leaf_reg: 3
thread_count: -1
verbose: false
random_state: 42Each model creates:
outputs/
linear/
model.pkl # joblib dump
metrics.json # RMSE, MAE, R2, MAPE
predictions.csv # index, y_true, y_pred
ridge/
lasso/
elasticnet/
random_forest/
extra_trees/
xgboost/
lightgbm/
catboost/
Plus benchmark summary:
outputs/benchmark_results.csv
Format: Model,RMSE,MAE,R2,MAPE
| Model | RMSE | MAE | R² | MAPE |
|---|---|---|---|---|
| CatBoost | 1.029 | 0.884 | 0.713 | 114.35 |
| Extra Trees | 1.035 | 0.833 | 0.710 | 110.53 |
| Ridge | 1.045 | 0.851 | 0.704 | 133.05 |
| Linear | 1.046 | 0.847 | 0.704 | 134.53 |
| Random Forest | 1.075 | 0.848 | 0.687 | 113.63 |
| XGBoost | 1.108 | 0.934 | 0.668 | 126.72 |
| Elastic Net | 1.425 | 1.143 | 0.450 | 84.71 |
| Lasso | 1.518 | 1.224 | 0.376 | 91.60 |
| LightGBM | 1.377 | 1.147 | 0.487 | 182.91 |
On Lightning AI Studio (FREE CPU):
git clone https://github.com/Zakariya-Q/hwin-baselines.git
cd hwin-baselines
pip install -r requirements.txt
bash scripts/run_all.shNo GPU required. All models run on CPU.
- Fixed
seed=42everywhere - Deterministic training
random_state=42passed to all estimators- Results logged with progress bars
pytest tests/ -v# Format code
black .
# Lint
ruff check .
# Pre-commit hooks
pre-commit installMIT License - see LICENSE
@software{hwin_baselines,
title = {hwin-baselines: Classical ML Baselines for HWIN-Bench},
author = {Q, Zakariya},
year = {2025},
url = {https://github.com/Zakariya-Q/hwin-baselines}
}Q: Why no deep learning baselines? A: This repo is specifically for classical ML baselines. Deep learning baselines (FT-Transformer, TabTransformer, TabPFN) are maintained separately.
Q: Can I add my own model?
A: Yes! See examples/custom_model.py for a template. Add your model to configs/default.yaml, create a training script in models/, and update scripts/run_all.sh and scripts/benchmark_summary.py.
Q: What if my data has different columns?
A: The code automatically uses all columns except y as features. Just ensure your target column is named y.
Q: How do I run on real HWIN-Bench data?
A: Download HWIN-Bench from the official repository, place parquet files in data/, and run bash scripts/run_all.sh.