41 topics
Data science
Statistics, evaluation and machine-learning concepts with the math shown.
A
B
- Bias-Variance TradeoffBias-variance decomposition separates a model's error into bias (systematic error from wrong assumptions) and variance (sensitivity to…
- Browser and Local MLRunning ML models directly in the browser or on the local machine without a server enables privacy-preserving, low-latency inference.
C
- ClassificationClassification predicts a discrete class label (spam/not spam, iris species) from input features.
- ClusteringClustering groups similar data points without predefined labels — an unsupervised technique for discovering structure.
- Cross-ValidationCross-validation evaluates a learning procedure on held-out data. In k-fold CV, each fold takes a turn as the validation set while the…
D
- Data CleaningData cleaning is the process of fixing errors, inconsistencies, and structural issues in raw data before analysis.
- Data LeakageData leakage occurs when information from the test set sneaks into the training process, making evaluation overly optimistic.
- Data VisualizationVisualization translates data into graphics to reveal patterns, trends, and anomalies that numbers alone obscure.
- Dataset shiftDataset shift occurs when the distribution encountered during use differs from the training distribution. It can affect inputs, labels, or…
- Decision TreesA decision tree splits data recursively on feature values, creating a flowchart-like model that's easy to interpret.
- Descriptive StatisticsDescriptive statistics summarize the central tendency, spread, and shape of a dataset.
- Dimensionality ReductionDimensionality reduction transforms high-dimensional data into a lower-dimensional representation while preserving key structure.
E
F
- F1 ScoreThe F1 score is the harmonic mean of precision and recall, balancing both in a single metric.
- Feature EngineeringFeature engineering transforms raw data into model-friendly inputs, often determining model success more than algorithm choice.
- Feature ScalingFeature scaling standardizes the range or distribution of features so that no single feature dominates distance calculations.
G
K
L
M
- Missing ValuesMissing values (NaN in pandas, null in SQL) represent unknown or uncollected data that must be handled before modeling.
- Model DeploymentDeploying a model means making it available to make predictions on new data in a production environment.
- Model EvaluationModel evaluation quantifies how well a model performs, using metrics tailored to the problem type and business goal.
N
O
P
- Pandas DataFrameA pandas DataFrame is a two-dimensional labeled data structure resembling a table with named columns and a row index.
- Precision and RecallPrecision measures how trustworthy positive predictions are; recall measures how many actual positives the model captures.
- Principal Component AnalysisPCA finds orthogonal directions of maximum variance in data, projecting onto a lower-dimensional subspace.
- Probability BasicsProbability quantifies uncertainty, measuring how likely an event is on a scale from 0 (impossible) to 1 (certain).
- Probability calibrationCalibration asks whether predicted probabilities agree with observed frequencies. A model can rank cases well while assigning misleading…
- Probability DistributionsA probability distribution assigns probabilities to outcomes of a random variable; key distributions model common data-generating processes.
- Python for Data WorkPython's data ecosystem centers on pandas for tabular data and NumPy for numerical computing.
R
- RegressionRegression predicts a continuous numerical outcome (house price, GDP growth) from input features.
- RegularizationRegularization adds a penalty to the loss function to discourage overly complex models, reducing overfitting.
- ROC and AUCThe ROC curve plots true positive rate (recall) against false positive rate at every classification threshold; AUC summarizes it.