Model Evaluation
Model evaluation quantifies how well a model performs, using metrics tailored to the problem type and business goal.
Definition
Classification metrics: accuracy, precision, recall, F1, AUC-ROC. Regression metrics: , RMSE, MAE, MAPE.
Always evaluate on held-out data — never on training data, which gives an optimistically biased estimate.
Intuition
No single metric tells the whole story; the right metric depends on class imbalance, asymmetric costs, and whether you need calibrated probabilities.
A model with 99% accuracy on a 1% fraud dataset might just predict "not fraud" for everything.
Worked example
In a cancer screening test, missing a cancer (false negative) is far costlier than a false alarm — use recall-oriented metrics.
For a recommendation system, ranking quality (AUC, NDCG) matters more than per-user prediction accuracy.
The math
Confusion matrix: rows = actual, columns = predicted. From it: Precision = , Recall = .
, the harmonic mean favoring balance.
In practice
Use a validation set during development to compare models; report final metrics on the test set once.
For business metrics, translate model metrics (e.g., AUC) into expected revenue or cost impact when possible.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.