Data science

Principal Component Analysis

PCA finds orthogonal directions of maximum variance in data, projecting onto a lower-dimensional subspace.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

PCA finds the kk orthogonal axes (principal components) that capture the most variance; projecting onto them loses as little information as possible.

Each component is a linear combination of the original features; components are ordered by explained variance.

Intuition

Dimensionality reduction helps when features outnumber samples, mitigates overfitting, and enables visualization (2-3 principal components).

PCA is unsupervised — it doesn't use the target variable, only the feature covariance structure.

Worked example

With 100 features but only 50 samples, PCA to 10-20 components before classification can dramatically improve generalization.

Visualizing high-dimensional data by plotting the first two principal components reveals clusters and outliers.

The math

PCA solves max⁡∣∣v∣∣=1Var(Xv)=max⁡vTXTXvvTv\max_{||v||=1} \text{Var}(Xv) = \max \frac{v^TX^TXv}{v^Tv}, the leading eigenvector of the covariance matrix XTXX^TX.

Variance explained by component jj: λj/∑iλi\lambda_j / \sum_i \lambda_i. A scree plot shows this to choose kk.

In practice

Used for noise filtering, compression, feature extraction, and as preprocessing before clustering or regression.

Note: PCA directions may not be interpretable as meaningful "factors" — they are purely statistical.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.