F1 Score
The F1 score is the harmonic mean of precision and recall, balancing both in a single metric.
Definition
.
F1 ranges from 0 to 1; 1 is perfect precision and recall.
Intuition
The harmonic mean penalizes extreme imbalance — a model with precision 1.0 and recall 0.1 gets F1 = 0.18, not 0.55.
Use F1 when you need a single number balancing precision and recall without tuning a threshold.
Worked example
Model A: precision=0.9, recall=0.5 → F1 = 0.64. Model B: precision=0.7, recall=0.7 → F1 = 0.70. Model B is better overall.
For imbalanced datasets, F1 is preferred over accuracy.
The math
Generalized to : ; (standard F1), weights precision more, weights recall more.
Macro F1 averages F1 per class; micro F1 aggregates TP, FP, FN globally before computing.
In practice
Standard metric for NLP tasks like named entity recognition and text classification on imbalanced corpora.
F1 is widely used in Kaggle competitions and academic benchmarks for classification.
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.