Data science

Linear Regression

Linear regression predicts a continuous target as a weighted sum of features, fit by ordinary least squares (OLS).

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

y^=wTx+b\hat{y} = w^Tx + b where ww minimizes ∑i(yi−wTxi−b)2\sum_i (y_i - w^Tx_i - b)^2, the residual sum of squares (RSS).

Closed-form solution: w=(XTX)−1XTyw = (X^TX)^{-1}X^Ty (when XTXX^TX is invertible).

Intuition

Each coefficient wjw_j tells you the expected change in yy for a one-unit increase in xjx_j, holding all else constant.

OLS is the best linear unbiased estimator (BLUE) under Gauss-Markov assumptions: linearity, exogeneity, homoscedasticity, no perfect multicollinearity.

Worked example

Predicting weight from height: y^=−100+0.9×height(cm)\hat{y} = -100 + 0.9 \times \text{height(cm)}, so each extra cm adds about 0.9 kg.

If w2=3.5w_2 = 3.5, then feature x2x_2 has a positive association with yy after controlling for other features.

The math

Assumptions: y=Xw+ϵy = Xw + \epsilon, with E[ϵ∣X]=0E[\epsilon|X] = 0, Var(ϵ∣X)=σ2I\text{Var}(\epsilon|X) = \sigma^2 I, and ϵ∼N(0,σ2)\epsilon \sim N(0, \sigma^2) for inference.

Hypothesis test on wjw_j: H0:wj=0H_0: w_j = 0 (feature has no effect); p-value from tt-statistic t=wj/SE(wj)t = w_j / \text{SE}(w_j).

In practice

Baseline regression model; interpretable coefficients make it the go-to for explanatory modeling.

Check residual plots for non-linearity, heteroscedasticity, and outliers — violations indicate model misspecification.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.