I'm a Data Scientist and ML Engineer, and I recently finished my M.S. in Data Science at Rutgers. My work spans the full pipeline — designing ETL systems, fine-tuning ML models, running causal inference studies, and building dashboards that non-technical stakeholders can actually use and trust.
While finishing my degree, I taught Python, SQL, and ML to 120+ students at City College of New York. Explaining complex systems to people who are new to them sharpens your own understanding fast.
I'm actively looking for data science, ML engineering, or analytics engineering roles — ideally somewhere the problems are hard, and the data is messy.
📍 New York City · Open to hybrid and remote
| RAISE Finalist | LLM-Powered Sentiment Engine with BERT & Prompt Engineering — Feb 2025 |
| IEEE Published | Institutional Accreditation Data Analysis and Decision-Making System |
| World Happiness Prediction | Best Performer — Dec 2024 |
| Writing on Medium | SQL, Data Science & LLMs → medium.com/@rheagupta993 |
Also work with:
pandas numpy XGBoost HuggingFace Transformers BERT NLP matplotlib seaborn Power BI Tableau DAX T-SQL MLflow PyTest PySpark
Methods:
Causal Inference Propensity Score Matching OLS Regression A/B Testing Feature Engineering ETL Pipeline Design Data Warehousing Time Series
Fine-tuned BERT for multi-class sentiment classification with custom preprocessing pipelines, batched inference, and full evaluation tracking across F1, precision, and recall. Documented architecture and training methodology for full reproducibility.
Real-time route optimization system that cut bus clustering by ~35% on peak routes using GPS data, scheduling algorithms, and anomaly detection. Includes an interactive React + Leaflet dashboard for live position tracking.
End-to-end variance tracking across a $30M FY2024 media budget. Automated run-rate projections flagged 2 high-risk overspend channels (Digital Video +$374K, Programmatic +$114K) before quarter close — enabling proactive reallocation.
3-layer financial data warehouse built from scratch: raw ingestion → normalized schema → reporting mart. Includes a FICO-style credit scoring engine, a rule-based fraud detection system, and a 6-panel analytics dashboard across 50 synthetic customers.
Estimated the Average Treatment Effect (~+0.41) of a growth mindset intervention using OLS regression, Propensity Score Matching, and Inverse Probability Weighting. Validated assumptions through covariate balance, overlap, and ignorability checks.
I document what I'm learning as I build — currently working through AWS infrastructure from the ground up. The goal is practical clarity: what each service actually does, when to use it, and how the pieces connect.
AWS Series · ongoing
| Article | |
|---|---|
| AWS Complete Beginner to Advanced Guide (2026) | Pinned · full roadmap |
| Amazon Aurora | When RDS isn't enough |
| Amazon RDS: Managed Relational Databases on AWS | Setup, config, and tradeoffs |
| AWS RDS vs Aurora | Side-by-side comparison |
| What is Amazon Redshift? | Data warehousing on AWS |
| Amazon ElastiCache | In-memory caching at scale |
| Introduction to Amazon VPC | Networking fundamentals |
| How to Create a Custom VPC in AWS | Step-by-step walkthrough |
├── LLM agent architectures — tool use, memory, multi-step reasoning
├── Orchestration patterns with LangGraph and CrewAI
└── Writing about everything I learn on Medium as I go
If you're working on something in data science, ML, or analytics — or you're hiring — I'd like to hear about it.
rheavinod.gupta@rutgers.edu · LinkedIn · Medium
Based in New York City · Open to hybrid and remote roles

