Data science

Probability Distributions

A probability distribution assigns probabilities to outcomes of a random variable; key distributions model common data-generating processes.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

Binomial B(n,p)B(n, p): number of successes in nn independent Bernoulli trials, PMF P(k)=(nk)pk(1−p)n−kP(k) = \binom{n}{k}p^k(1-p)^{n-k}.

Normal (Gaussian) N(μ,σ2)N(\mu, \sigma^2): PDF 1σ2πe−(x−μ)2/2σ2\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}; symmetric, bell-shaped, defined by mean μ\mu and std σ\sigma.

Intuition

The central limit theorem says the sum/mean of i.i.d. random variables converges to normal as nn grows — this is why normal appears everywhere.

Poisson models "events per interval" — arrivals per minute, defects per item — with a single rate parameter λ\lambda.

Worked example

Heights of adult humans approximate N(170,92)N(170, 9^2) cm; IQ is centered at 100 with std 15.

Number of spam emails in an hour might follow Poisson(7) if you average 7 per hour.

The math

Normal properties: 68% of mass within 1σ1\sigma, 95% within 2σ2\sigma, 99.7% within 3σ3\sigma of μ\mu.

Standard normal Z=(X−μ)/σZ = (X - \mu)/\sigma; used to compute normal probabilities via ZZ-tables or scipy.stats.norm.cdf.

In practice

Assume normal errors in linear regression; check residuals visually to validate.

Use Poisson regression for count data; binomial for binary outcomes.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.