Data science

Probability Basics

Probability quantifies uncertainty, measuring how likely an event is on a scale from 0 (impossible) to 1 (certain).

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

P(A ∪ B) = P(A) + P(B) − P(A ∩ B); P(A ∩ B) = P(A) · P(B | A) for dependent events.

Bayes' theorem: P(A∣B)=P(B∣A)⋅P(A)P(B)P(A | B) = \frac{P(B | A) \cdot P(A)}{P(B)}, updating belief in A after observing B.

Intuition

Conditional probability P(A∣B)P(A | B) is the probability of A given that B has occurred — it restricts the sample space.

The "base rate fallacy" ignores P(A)P(A): a rare event with high false-positive rate often has low true positive predictive value.

Worked example

If 1% of people have a disease and the test is 99% accurate, P(positive∣sick)=0.99P(\text{positive} | \text{sick}) = 0.99 but P(sick∣positive)≈50%P(\text{sick} | \text{positive}) \approx 50\% due to the low base rate.

Drawing two aces without replacement from a deck: P(2nd ace∣1st ace)=3/51P(\text{2nd ace} | \text{1st ace}) = 3/51.

The math

Random variable XX maps outcomes to numbers; the PMF P(X=x)P(X = x) gives probabilities for discrete XX, the PDF f(x)f(x) for continuous XX.

Expectation E[X]=∑x⋅P(X=x)E[X] = \sum x \cdot P(X = x); variance Var(X)=E[(X−E[X])2]\text{Var}(X) = E[(X - E[X])^2].

In practice

Powers Bayes' theorem for spam filtering, medical diagnosis, and A/B testing — everywhere you update beliefs with evidence.

Probability distributions model real phenomena: binomial for coin flips, Poisson for rare events per time interval, normal for measurement errors.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.