Data science

F1 Score

The F1 score is the harmonic mean of precision and recall, balancing both in a single metric.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

F1=2⋅Precision⋅RecallPrecision+Recall=2⋅TP2⋅TP+FP+FNF_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}} = \frac{2\cdot TP}{2\cdot TP + FP + FN}.

F1 ranges from 0 to 1; 1 is perfect precision and recall.

Intuition

The harmonic mean penalizes extreme imbalance — a model with precision 1.0 and recall 0.1 gets F1 = 0.18, not 0.55.

Use F1 when you need a single number balancing precision and recall without tuning a threshold.

Worked example

Model A: precision=0.9, recall=0.5 → F1 = 0.64. Model B: precision=0.7, recall=0.7 → F1 = 0.70. Model B is better overall.

For imbalanced datasets, F1 is preferred over accuracy.

The math

Generalized to FβF_\beta: Fβ=(1+β2)P⋅Rβ2P+RF_\beta = (1+\beta^2)\frac{P\cdot R}{\beta^2 P + R}; β=1\beta=1 (standard F1), β=0.5\beta=0.5 weights precision more, β=2\beta=2 weights recall more.

Macro F1 averages F1 per class; micro F1 aggregates TP, FP, FN globally before computing.

In practice

Standard metric for NLP tasks like named entity recognition and text classification on imbalanced corpora.

F1 is widely used in Kaggle competitions and academic benchmarks for classification.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.