Pandas DataFrame
A pandas DataFrame is a two-dimensional labeled data structure resembling a table with named columns and a row index.
Definition
A DataFrame is pd.DataFrame(data, index=rows, columns=cols) where each column is a pandas Series.
Key operations include select (df["col"]), filter (df[df["col"] > 0]), transform (df.apply), and aggregate (df.groupby).
Intuition
Every column is a Series sharing the same index; think of it as a dictionary of arrays that all line up by row.
The index makes alignment automatic — when you add two DataFrames, rows match by label, not by position.
Worked example
df[df["age"] > 30][["name", "city"]] selects adults and keeps only name and city columns.
df.groupby("department")["salary"].mean() computes average salary per department.
The math
df.merge(other, on="id", how="left") performs SQL-style joins; pandas aligns on the "id" column.
Method chaining df.pipe(clean).groupby("x").agg({"y": "mean"}) composes operations cleanly.
In practice
Used for loading CSVs, Excel files, SQL query results, and any rectangular dataset.
Essential for exploratory data analysis, data cleaning, and feature engineering in Python ML pipelines.
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.