Data Cleaning
Data cleaning is the process of fixing errors, inconsistencies, and structural issues in raw data before analysis.
Definition
Common issues: duplicate rows, mis-typed values ("N/A" vs NaN), column names with extra spaces, and date formats that don't parse.
Cleaning steps include parsing, deduplication, type coercion, and standardization.
Intuition
Garbage in, garbage out — even the best model fails if trained on messy data. Cleaning is where most data scientists spend the most time.
Aim for "tidy data": each column is a variable, each row is an observation, each cell is a single value.
Worked example
df.columns = df.columns.str.strip() removes leading/trailing spaces from column names.
df.drop_duplicates() removes duplicate rows; df["price"] = pd.to_numeric(df["price"], errors="coerce") converts strings to numbers.
The math
df.replace({"N/A": np.nan, "null": np.nan}) maps error strings to NaN in one pass.
df.duplicated().sum() counts duplicates; df.isnull().sum() counts NaNs per column.
In practice
Always check for outliers and impossible values (e.g., negative age, future dates) as part of cleaning.
Clean data enables reliable analysis and prevents subtle bugs in downstream model training.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.