Data science

Pandas DataFrame

A pandas DataFrame is a two-dimensional labeled data structure resembling a table with named columns and a row index.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

A DataFrame is pd.DataFrame(data, index=rows, columns=cols) where each column is a pandas Series.

Key operations include select (df["col"]), filter (df[df["col"] > 0]), transform (df.apply), and aggregate (df.groupby).

Intuition

Every column is a Series sharing the same index; think of it as a dictionary of arrays that all line up by row.

The index makes alignment automatic — when you add two DataFrames, rows match by label, not by position.

Worked example

df[df["age"] > 30][["name", "city"]] selects adults and keeps only name and city columns.

df.groupby("department")["salary"].mean() computes average salary per department.

The math

df.merge(other, on="id", how="left") performs SQL-style joins; pandas aligns on the "id" column.

Method chaining df.pipe(clean).groupby("x").agg({"y": "mean"}) composes operations cleanly.

In practice

Used for loading CSVs, Excel files, SQL query results, and any rectangular dataset.

Essential for exploratory data analysis, data cleaning, and feature engineering in Python ML pipelines.

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.