Data science

Missing Values

Missing values (NaN in pandas, null in SQL) represent unknown or uncollected data that must be handled before modeling.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

df.isnull().sum() counts NaNs per column; df.dropna() removes rows with any NaN, df.fillna(value) replaces them.

Imputation replaces NaN with a substituted value — mean, median, mode, or a predicted value from other features.

Intuition

Dropping rows with NaN is safe if missingness is random and small; otherwise you lose data and potentially introduce bias.

Mean imputation preserves the mean but reduces variance; model-based imputation (e.g., using a regressor) can do better.

Worked example

df["age"].fillna(df["age"].median(), inplace=True) fills missing ages with the median age.

Sklearn's SimpleImputer(strategy="most_frequent") fills with the most common value for categorical columns.

The math

Missing at Random (MAR): the probability of missing depends on observed data; Missing Completely at Random (MCAR): no systematic relationship.

Multiple imputation (creating several imputed datasets and combining) accounts for imputation uncertainty.

In practice

Check if missingness in a column correlates with the target — if so, the missingness itself is a signal (informative missingness).

Some models like decision trees handle NaN natively; others like logistic regression require complete data.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.