Missing Values
Missing values (NaN in pandas, null in SQL) represent unknown or uncollected data that must be handled before modeling.
Definition
df.isnull().sum() counts NaNs per column; df.dropna() removes rows with any NaN, df.fillna(value) replaces them.
Imputation replaces NaN with a substituted value — mean, median, mode, or a predicted value from other features.
Intuition
Dropping rows with NaN is safe if missingness is random and small; otherwise you lose data and potentially introduce bias.
Mean imputation preserves the mean but reduces variance; model-based imputation (e.g., using a regressor) can do better.
Worked example
df["age"].fillna(df["age"].median(), inplace=True) fills missing ages with the median age.
Sklearn's SimpleImputer(strategy="most_frequent") fills with the most common value for categorical columns.
The math
Missing at Random (MAR): the probability of missing depends on observed data; Missing Completely at Random (MCAR): no systematic relationship.
Multiple imputation (creating several imputed datasets and combining) accounts for imputation uncertainty.
In practice
Check if missingness in a column correlates with the target — if so, the missingness itself is a signal (informative missingness).
Some models like decision trees handle NaN natively; others like logistic regression require complete data.
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.