Data science

Descriptive Statistics

Descriptive statistics summarize the central tendency, spread, and shape of a dataset.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

Central tendency: mean (average), median (middle value), mode (most frequent). Spread: variance, standard deviation, range, IQR.

The shape is described by skewness (asymmetry) and kurtosis (tail weight relative to a normal distribution).

Intuition

Mean is sensitive to outliers — a single extreme value pulls the mean far from the median. Use median when your data is skewed.

Standard deviation is in the same units as the data, making it more interpretable than variance.

Worked example

np.std([2, 4, 4, 4, 5, 5, 7, 9]) ≈ 2.0 — most values cluster within 2 of the mean of 5.

df.describe() prints a table with count, mean, std, min, 25%, 50%, 75%, max for all numeric columns.

The math

Sample variance: s2=1n−1∑(xi−xˉ)2s^2 = \frac{1}{n-1}\sum(x_i - \bar{x})^2 (uses n−1n-1 for unbiased estimation).

IQR = Q3 - Q1; data points outside [Q1−1.5⋅IQR,  Q3+1.5⋅IQR][Q1 - 1.5\cdot IQR,\; Q3 + 1.5\cdot IQR] are considered outliers by Tukey's rule.

In practice

Always inspect descriptive stats before modeling — they guide outlier handling, transformation choice, and model selection.

Correlation coefficients (Pearson rr, Spearman ρ\rho) quantify linear and monotonic relationships between pairs of variables.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.