Data science

Logistic Regression

Logistic regression models the probability of a binary outcome as a logistic (sigmoid) function of a linear combination of features.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

P(Y=1∣X)=σ(wTx+b)P(Y=1|X) = \sigma(w^Tx + b) where σ(z)=1/(1+e−z)\sigma(z) = 1/(1 + e^{-z}) squashes the linear output to [0, 1].

Parameters w,bw, b are fit by maximizing the likelihood — equivalent to minimizing binary cross-entropy loss.

Intuition

Despite "regression" in the name, it's a classification method — the sigmoid converts a raw score to a calibrated probability.

Log-odds log⁡(P/(1−P))\log(P/(1-P)) is linear in the features, so each unit increase in xjx_j multiplies the odds by ewje^{w_j}.

Worked example

Predicting loan default: P(default=1)=σ(−3.2+0.71⋅income−0.42⋅debt)P(\text{default}=1) = \sigma(-3.2 + 0.71 \cdot \text{income} - 0.42 \cdot \text{debt}).

If w1=0.71w_1 = 0.71, a one-unit increase in income multiplies the odds of default by e0.71≈2.0e^{0.71} \approx 2.0.

The math

Loss: L=−∑[yilog⁡y^i+(1−yi)log⁡(1−y^i)]\mathcal{L} = -\sum[y_i\log\hat{y}_i + (1-y_i)\log(1-\hat{y}_i)], minimized by gradient descent.

Multinomial logistic regression (softmax) extends to K>2K > 2 classes with softmax⁡k(z)=ezk/∑jezj\softmax_k(z) = e^{z_k}/\sum_j e^{z_j}.

In practice

Baseline classifier for binary problems — fast, interpretable (coefficients show direction and magnitude), and often competitive.

Used in epidemiology, credit scoring, and any domain where you need both predictions and calibrated probabilities.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.