Data science

Python for Data Work

Python's data ecosystem centers on pandas for tabular data and NumPy for numerical computing.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

pandas provides DataFrames — two-dimensional tables with labeled rows and columns — for working with structured data.

NumPy supplies ndarrays, n-dimensional arrays that enable fast vectorized operations without explicit loops.

Intuition

Think of pandas as Excel with code — you get the flexibility of a spreadsheet but the repeatability of a script.

NumPy arrays are like Python lists but designed for math: they're homogeneous and support element-wise operations directly.

Worked example

pd.read_csv("sales.csv") loads a CSV into a DataFrame; df.groupby("region").sales.sum() aggregates by region.

np.array([1, 2, 3]) * 2 yields array([2, 4, 6]) — the multiplication applied element-wise automatically.

The math

A Series has an index and a dtype; its data can use NumPy arrays or extension arrays, including nullable and Arrow-backed types.

NumPy uses BLAS/LAPACK for optimized linear algebra under the hood.

In practice

The pandas/NumPy stack is the foundation for nearly all Python-based data analysis and machine learning.

Libraries like scikit-learn, statsmodels, and seaborn all consume pandas DataFrames or NumPy arrays.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.