Descriptive Statistics
Descriptive statistics summarize the central tendency, spread, and shape of a dataset.
Definition
Central tendency: mean (average), median (middle value), mode (most frequent). Spread: variance, standard deviation, range, IQR.
The shape is described by skewness (asymmetry) and kurtosis (tail weight relative to a normal distribution).
Intuition
Mean is sensitive to outliers — a single extreme value pulls the mean far from the median. Use median when your data is skewed.
Standard deviation is in the same units as the data, making it more interpretable than variance.
Worked example
np.std([2, 4, 4, 4, 5, 5, 7, 9]) ≈ 2.0 — most values cluster within 2 of the mean of 5.
df.describe() prints a table with count, mean, std, min, 25%, 50%, 75%, max for all numeric columns.
The math
Sample variance: (uses for unbiased estimation).
IQR = Q3 - Q1; data points outside are considered outliers by Tukey's rule.
In practice
Always inspect descriptive stats before modeling — they guide outlier handling, transformation choice, and model selection.
Correlation coefficients (Pearson , Spearman ) quantify linear and monotonic relationships between pairs of variables.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.