This project is a complete healthcare analytics and time-series forecasting system for estimating care load, discharge demand, intake-discharge imbalance, and capacity stress indicators for HHS-style daily operational datasets.
It is designed for:
- VS Code execution on Windows
- GitHub upload
- Streamlit dashboard demos
- Resume and portfolio use
- College final-year project submission
- Research or stakeholder presentation
Source context: HHS
Live deployed app: Predictive Forecasting Dashboard
- Dynamic CSV loading with safe column-name detection.
- Missing-value handling, duplicate removal, and date conversion.
- Daily continuity checks with date interpolation.
- Outlier detection using the IQR method.
- Premium dark UI using the approved government-grade palette:
#2563EB,#0F172A,#F59E0B,#EF4444,#F8FAFC, and#CBD5E1. - Glassmorphism KPI cards, risk alert panels, responsive layout, hover effects, Plotly dark charts, and Streamlit 1.56-compatible APIs.
- Feature engineering for lags, rolling averages, rolling standard deviation, calendar fields, weekend flag, and net pressure.
- Baseline models: Naive Forecast and Moving Average.
- Statistical models: ARIMA, SARIMA, and Exponential Smoothing.
- Machine learning models: Random Forest Regressor and Gradient Boosting Regressor.
- Model evaluation with MAE, RMSE, MAPE, R2 Score, and accuracy percentage.
- Automatic best-model selection by lowest RMSE.
- 7-day, 14-day, and 30-day forecasting.
- Confidence intervals for future predictions.
- Surge detection with capacity-risk banners.
- Professional Streamlit dashboard with upload, model selection, forecast horizon, KPI cards, charts, and CSV download.
- Deployment files for Streamlit Cloud, Render, and Hugging Face Spaces.
- Python 3.11+
- Pandas
- NumPy
- Matplotlib
- Seaborn
- Plotly
- Scikit-learn
- Statsmodels
- Streamlit
- Joblib
Predictive_Forecasting_Project/
|
|-- app/
| |-- __init__.py
| |-- main.py # Streamlit dashboard
| |-- forecast.py # Future forecasting and surge detection
| |-- preprocessing.py # Data cleaning and validation
| |-- utils.py # Shared constants and helpers
| |-- visualizations.py # EDA and Plotly charts
| |-- model_training.py # Model training and evaluation
|
|-- data/
| |-- dataset.csv # Place your dataset here
| |-- cleaned_dataset.csv # Generated after running the pipeline
|
|-- models/
| |-- saved model files # Generated after training
|
|-- notebooks/
| |-- EDA.ipynb
|
|-- outputs/
| |-- charts/ # Generated chart images
| |-- reports/ # Generated CSV and JSON reports
|
|-- .streamlit/
| |-- config.toml
|
|-- app.py # Deployment wrapper
|-- main.py # Command-line pipeline
|-- requirements.txt
|-- README.md
|-- .gitignore
|-- Procfile
|-- runtime.txt
|-- packages.txt
|-- run_project.bat
Expected columns are similar to:
- Date
- Children apprehended and placed in CBP custody
- Children in CBP custody
- Children transferred out of CBP custody
- Children in HHS Care
- Children discharged from HHS Care
The code dynamically cleans and maps column names. Exact capitalization and punctuation are not required.
- Open python.org/downloads.
- Download Python 3.11 or newer.
- During installation, select Add python.exe to PATH.
- Verify in PowerShell:
python --version- Open VS Code.
- Click File > Open Folder.
- Select:
Predictive_Forecasting_Project
- Open the terminal:
Ctrl + `
python -m venv .venvPowerShell:
.venv\Scripts\Activate.ps1Command Prompt:
.venv\Scripts\activate.batIf PowerShell blocks activation, run:
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUserThen activate again.
python -m pip install --upgrade pip
pip install -r requirements.txtPlace your CSV file at:
data/dataset.csv
This project already includes the mentioned HHS CSV as data/dataset.csv. You can replace it with a new dataset later.
python main.pyRun with a specific horizon:
python main.py --horizon 7
python main.py --horizon 14
python main.py --horizon 30Run with a specific target:
python main.py --target children_discharged_from_hhs_care --horizon 30streamlit run app/main.pyAlternative:
python -m streamlit run app/main.py.\run_project.batThe batch file creates a virtual environment, installs packages, runs the pipeline, and starts Streamlit.
- Run:
python main.py --horizon 30- Confirm these files were created:
outputs/reports/model_leaderboard.csv
outputs/reports/forecast_children_in_hhs_care_30_days.csv
outputs/reports/surge_warning_report.json
models/best_model.pkl
- Start the dashboard:
streamlit run app/main.py- In the sidebar:
- Upload
data/dataset.csv, or use the default dataset. - Select forecast target.
- Select model or choose Auto Best.
- Select 7, 14, or 30 days.
- Click Run Forecast.
- Review:
- KPI cards
- Forecast chart
- Confidence interval band
- Capacity risk indicator
- Model comparison table
- Downloadable forecast CSV
The project calculates:
- MAE: Mean Absolute Error
- RMSE: Root Mean Squared Error
- MAPE: Mean Absolute Percentage Error
- R2 Score
- Forecast Accuracy %
The best model is automatically selected using the lowest RMSE.
The system calculates a recent 30-day rolling average for the selected target. If forecast values or upper confidence interval values exceed that rolling average by 10% or more, the system flags capacity risk.
Risk levels:
LOW CAPACITY RISKMODERATE CAPACITY RISKHIGH CAPACITY RISK
Run these commands in Windows PowerShell from the project root:
git init
git add .
git commit -m "Initial commit: predictive forecasting project"
git branch -M main
git remote add origin https://github.com/YOUR_USERNAME/YOUR_REPOSITORY_NAME.git
git push -u origin mainSteps:
- Create a new repository on GitHub.
- Do not initialize it with a README because this project already has one.
- Copy the repository URL.
- Replace
YOUR_USERNAMEandYOUR_REPOSITORY_NAMEin the command above. - Run the commands.
Required files:
requirements.txtapp/main.py.streamlit/config.toml
The dashboard uses Streamlit 1.56+ features such as st.iframe and width="stretch", so keep streamlit>=1.56.0 in requirements.txt.
Steps:
- Push the project to GitHub.
- Go to Streamlit Cloud.
- Click New app.
- Select your GitHub repository.
- Set the main file path:
app/main.py
- Click Deploy.
Common fixes:
- If packages fail, confirm
requirements.txtis in the repository root. - If the dataset is missing, upload from the dashboard sidebar or commit
data/dataset.csv. - If the app sleeps, wake it from Streamlit Cloud dashboard.
Required files:
requirements.txtProcfileruntime.txt
Steps:
- Push the project to GitHub.
- Go to Render.
- Click New Web Service.
- Connect your repository.
- Select Python environment.
- Build command:
pip install -r requirements.txt- Start command:
streamlit run app/main.py --server.port=$PORT --server.address=0.0.0.0- Deploy.
Common fixes:
- If
$PORTfails locally, ignore it locally. Render provides it in deployment. - If the app cannot find data, upload the CSV in the dashboard or commit
data/dataset.csv.
Required files:
requirements.txtapp.pyapp/data/
Steps:
- Create a new Space at Hugging Face Spaces.
- Choose Streamlit as the SDK.
- Upload all project files.
- Use the root
app.pywrapper as the entry point. - Confirm
requirements.txtis at the root.
Common fixes:
- If imports fail, confirm the
app/folder contains__init__.py. - If packages fail, check Python package versions in
requirements.txt. - If data is not included, use the dashboard uploader.
Reinstall Python and select Add python.exe to PATH.
Use:
python -m pip install -r requirements.txtUse:
python -m streamlit run app/main.pyCheck that your CSV has:
- A date-like column
- At least one numeric target column
- Enough rows for forecasting
The SARIMA and Random Forest models can take longer on large datasets. Start with the default HHS CSV to verify your environment.
Choose Auto Best. Some statistical models may fail on unusual datasets, but the system will continue with the models that train successfully.
After running the dashboard, capture these screenshots for reports or presentations:
- Dashboard overview
- KPI cards
- Forecast chart with confidence intervals
- Model leaderboard
- Capacity risk indicator
- Discharge prediction panel
Suggested folder:
screenshots/
- Add XGBoost after confirming deployment environment support.
- Add automated hyperparameter tuning.
- Add holiday, policy-event, and operational-capacity features.
- Add drift monitoring.
- Add Docker deployment.
- Add authenticated dashboard access.
- Add automated report generation.
This project is released under the MIT License for academic, portfolio, and demonstration use.