Data science

Model Evaluation

Model evaluation quantifies how well a model performs, using metrics tailored to the problem type and business goal.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

Classification metrics: accuracy, precision, recall, F1, AUC-ROC. Regression metrics: R2R^2, RMSE, MAE, MAPE.

Always evaluate on held-out data — never on training data, which gives an optimistically biased estimate.

Intuition

No single metric tells the whole story; the right metric depends on class imbalance, asymmetric costs, and whether you need calibrated probabilities.

A model with 99% accuracy on a 1% fraud dataset might just predict "not fraud" for everything.

Worked example

In a cancer screening test, missing a cancer (false negative) is far costlier than a false alarm — use recall-oriented metrics.

For a recommendation system, ranking quality (AUC, NDCG) matters more than per-user prediction accuracy.

The math

Confusion matrix: rows = actual, columns = predicted. From it: Precision = TP/(TP+FP)TP/(TP+FP), Recall = TP/(TP+FN)TP/(TP+FN).

F1=2⋅Precision⋅RecallPrecision+RecallF_1 = 2\cdot\frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}, the harmonic mean favoring balance.

In practice

Use a validation set during development to compare models; report final metrics on the test set once.

For business metrics, translate model metrics (e.g., AUC) into expected revenue or cost impact when possible.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.