Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

🧪 A/B Testing Case Study — Website Conversion Rate

A complete, end-to-end A/B testing analysis on an e-commerce landing page experiment — from raw data to a statistically-backed business recommendation. Built as part of my data analytics portfolio while transitioning from a Mechanical Engineering background into data analytics.

📌 Project Overview

A company tested a new landing page against its old landing page to see if it improves user conversion. This project analyzes the resulting experiment data to answer one question:

Should the company launch the new page, or keep the old one?

The analysis walks through the full statistical testing pipeline — data cleaning, conversion rate calculation, hypothesis testing, confidence intervals, effect size, and power analysis — before arriving at a clear, data-driven recommendation.

📂 Repository Structure

AB-Testing--Case-Study/
│
├── Data/
│   └── ab_data.csv          # Raw experiment data (294,479 user records)
│
├── python/
│   └── AB_testing.ipynb     # Full analysis notebook
│
└── README.md

📊 Dataset

The dataset (Data/ab_data.csv) contains 294,479 records of individual users who were randomly assigned to see either the old or new landing page.

Column Description
user_id Unique identifier for each user
timestamp Date and time the user visited the page
group Experiment group — control or treatment
landing_page Page shown to the user — old_page or new_page
converted Whether the user converted — 1 = yes, 0 = no

🛠️ Tools & Libraries

  • Python 3
  • Pandas & NumPy — data loading and cleaning
  • Matplotlib — visualization
  • Statsmodels — two-proportion z-test, confidence intervals, effect size, and power analysis

🔍 Methodology

The notebook follows a structured hypothesis-testing workflow:

  1. Data Cleaning — removed nulls, invalid group/page mismatches, and non-binary conversion values
  2. Group Balance Check — confirmed control and treatment groups were of comparable size
  3. Conversion Rate Calculation — computed conversion rate for each group
  4. Absolute Difference & Relative Lift — quantified the gap between treatment and control
  5. Two-Proportion Z-Test — tested whether the difference in conversion rates is statistically significant
  6. Hypothesis Decision — compared p-value against α = 0.05
  7. 95% Confidence Interval — estimated the plausible range for the true difference
  8. Effect Size (Cohen's h) — measured the practical size of the difference
  9. Power Analysis — checked whether the sample size was large enough to reliably detect an effect
  10. Visualization — bar chart comparing conversion rates
  11. Business Recommendation — final launch/no-launch decision based on the statistics

📈 Key Results

Metric Control (Old Page) Treatment (New Page)
Users 147,202 147,276
Conversions 17,723 17,514
Conversion Rate 12.04% 11.89%
Statistic Value
Absolute Difference -0.15%
Relative Lift -1.23%
Z-statistic 1.24
P-value 0.216
95% Confidence Interval [-0.38%, 0.09%]
Cohen's h (Effect Size) -0.0046 (negligible)
Statistical Power ~24%

✅ Conclusion

With a p-value of 0.216 (greater than α = 0.05), we fail to reject the null hypothesis — there is no statistically significant difference in conversion rate between the new page and the old page. The confidence interval for the difference includes zero, and the effect size is negligible.

Business takeaway: The new landing page does not improve conversions. The company should not launch the new page based on this data, and could instead invest in testing a more meaningfully different design.

Note: The observed statistical power (~24%) is well below the conventional 80% threshold, meaning the current sample is underpowered to detect a small effect even if one existed — a larger sample would be needed for a more conclusive result.

🎯 What I Learned

  • How to structure a full A/B test analysis the way it's done in industry, from raw data to a business decision
  • Difference between statistical significance (p-value) and practical significance (effect size)
  • Why confidence intervals matter more than a single p-value cutoff
  • How to run a power analysis and interpret whether a sample size is actually sufficient
  • Translating statistical output into a clear, non-technical business recommendation

🔮 Future Improvements

  • Segment the analysis by device type, region, or traffic source to check for hidden effects
  • Run a sequential/Bayesian A/B testing approach for comparison
  • Build an interactive Tableau dashboard to visualize the results

👤 Author

Punit Kumar B.Tech Mechanical Engineering, Delhi Technological University — transitioning into Data Analytics


⭐ If you found this project useful, consider giving it a star!

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages