Data science

Exploratory Data Analysis

EDA is the initial investigation of a dataset to understand its structure, spot anomalies, and generate hypotheses.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

EDA uses summary statistics (mean, median, std, quantiles) and visualizations (histograms, scatter plots, box plots) to understand data.

There is no fixed procedure — EDA is open-ended curiosity-driven exploration before formal modeling.

Intuition

Understanding the data first prevents building models for the wrong problem or missing critical data issues.

Start broad: what is this data? Who/what are the rows? What does each column represent? Are there missing values?

Worked example

A scatter plot of income vs. spending by age group might reveal a spending peak at retirement age that suggests a new feature.

A bar chart of customer count by country might show 80% of data comes from one country — important for train-test splitting.

The math

Correlation matrices (sns.heatmap(df.corr())) reveal which numeric features move together.

Pair plots (sns.pairplot) show joint distributions for key variables — look for clusters, gaps, and outliers.

In practice

Every data science project should start with EDA; it guides cleaning, feature engineering, and model choice.

Revisit EDA on residuals after modeling — structure in residuals reveals model misspecification.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.