Data science

Classification

Classification predicts a discrete class label (spam/not spam, iris species) from input features.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

Binary classification: two classes (0/1, True/False). Multiclass: three or more mutually exclusive classes.

The model outputs a predicted class directly (logistic regression's threshold at 0.5) or a probability (predict_proba returns P(Y=1∣X)P(Y=1|X)).

Intuition

Classification draws a decision boundary through feature space — each point gets the label of the region it falls into.

The predicted probability lets you calibrate confidence and adjust the decision threshold based on the cost of false positives vs. false negatives.

Worked example

Email spam detection: input features are words/patterns, output is "spam" or "not spam".

Medical diagnosis: input is patient vitals, output is "disease present" or "absent".

The math

For KK classes, use one-vs-rest (OvR) or one-vs-one (OvO) to extend binary classifiers; softmax regression generalizes logistic regression directly.

Decision trees classify by recursively splitting on feature thresholds that maximize information gain or Gini impurity.

In practice

Fraud detection, churn prediction, loan approval, image labeling, and sentiment analysis are all classification tasks.

Choose accuracy for balanced classes; use precision-recall or F1 when classes are imbalanced.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.