Exploratory Data Analysis
EDA is the initial investigation of a dataset to understand its structure, spot anomalies, and generate hypotheses.
Definition
EDA uses summary statistics (mean, median, std, quantiles) and visualizations (histograms, scatter plots, box plots) to understand data.
There is no fixed procedure — EDA is open-ended curiosity-driven exploration before formal modeling.
Intuition
Understanding the data first prevents building models for the wrong problem or missing critical data issues.
Start broad: what is this data? Who/what are the rows? What does each column represent? Are there missing values?
Worked example
A scatter plot of income vs. spending by age group might reveal a spending peak at retirement age that suggests a new feature.
A bar chart of customer count by country might show 80% of data comes from one country — important for train-test splitting.
The math
Correlation matrices (sns.heatmap(df.corr())) reveal which numeric features move together.
Pair plots (sns.pairplot) show joint distributions for key variables — look for clusters, gaps, and outliers.
In practice
Every data science project should start with EDA; it guides cleaning, feature engineering, and model choice.
Revisit EDA on residuals after modeling — structure in residuals reveals model misspecification.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.