Data science

Data Cleaning

Data cleaning is the process of fixing errors, inconsistencies, and structural issues in raw data before analysis.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

Common issues: duplicate rows, mis-typed values ("N/A" vs NaN), column names with extra spaces, and date formats that don't parse.

Cleaning steps include parsing, deduplication, type coercion, and standardization.

Intuition

Garbage in, garbage out — even the best model fails if trained on messy data. Cleaning is where most data scientists spend the most time.

Aim for "tidy data": each column is a variable, each row is an observation, each cell is a single value.

Worked example

df.columns = df.columns.str.strip() removes leading/trailing spaces from column names.

df.drop_duplicates() removes duplicate rows; df["price"] = pd.to_numeric(df["price"], errors="coerce") converts strings to numbers.

The math

df.replace({"N/A": np.nan, "null": np.nan}) maps error strings to NaN in one pass.

df.duplicated().sum() counts duplicates; df.isnull().sum() counts NaNs per column.

In practice

Always check for outliers and impossible values (e.g., negative age, future dates) as part of cleaning.

Clean data enables reliable analysis and prevents subtle bugs in downstream model training.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.