The objective of this task is to perform Exploratory Data Analysis (EDA) on the Titanic dataset to understand the data through statistical summaries and visualizations. EDA helps identify patterns, trends, relationships, and anomalies before building Machine Learning models.
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- KaggleHub
Titanic Dataset
- Downloaded and loaded the Titanic dataset using KaggleHub.
- Imported the dataset into a Pandas DataFrame.
- Examined dataset shape and structure.
- Inspected column names and data types.
- Checked for missing values.
-
Generated summary statistics such as:
- Mean
- Median
- Standard Deviation
- Minimum and Maximum values
- Quartiles
Created the following visualizations:
- Histograms for numerical features
- Age Distribution Plot
- Fare Boxplot for Outlier Detection
- Survival Count Plot
- Survival by Gender Plot
- Correlation Heatmap
- Pairplot for Feature Relationships
- Encoded categorical variables.
- Generated a correlation matrix to identify relationships between features.
- Most passengers were between 20 and 40 years of age.
- Fare distribution contains several outliers.
- Female passengers had a higher survival rate compared to males.
- Passenger class showed a relationship with survival probability.
- Some features contained missing values that required attention.
- Fare distribution was positively skewed.
- task2.py
- Titanic-Dataset.csv
- README.md
- histograms.png
- age_distribution.png
- boxplot_fare.png
- survival_count.png
- survival_gender.png
- correlation_heatmap.png
- pairplot.png
Successfully performed Exploratory Data Analysis (EDA) on the Titanic dataset using statistical techniques and visualizations. The analysis provided insights into data distribution, feature relationships, survival patterns, and potential outliers.
Submitted as part of the AI & ML Internship Program.