Data science

ROC and AUC

The ROC curve plots true positive rate (recall) against false positive rate at every classification threshold; AUC summarizes it.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

AUC = probability that a randomly chosen positive instance ranks higher than a randomly chosen negative one.

AUC = 0.5 means random guessing; AUC = 1.0 means perfect separation; AUC < 0.5 means the model is worse than random (invert predictions).

Intuition

AUC is threshold-independent — it measures ranking quality across all possible thresholds.

Unlike accuracy, AUC is robust to class imbalance because it uses rank rather than raw counts.

Worked example

AUC = 0.85 means an 85% chance that a randomly chosen positive sample scores higher than a randomly chosen negative one.

Comparing two models: AUC 0.92 vs 0.87 — the first has better overall ranking ability.

The math

At threshold tt, TPR(t)=P(Y^=1∣Y=1)\text{TPR}(t) = P(\hat{Y}=1 | Y=1) and FPR(t)=P(Y^=1∣Y=0)\text{FPR}(t) = P(\hat{Y}=1 | Y=0).

AUC can be interpreted as ∫01TPR(FPR−1(x))dx\int_0^1 \text{TPR}(FPR^{-1}(x)) dx, the area under the ROC curve.

In practice

Used when you care about ranking (search, recommendation, fraud ranking) rather than a single class decision.

For calibrated probabilities, also report Brier score E[(Y−p^)2]E[(Y - \hat{p})^2] to measure probability estimation quality.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.