Python for Data Work
Python's data ecosystem centers on pandas for tabular data and NumPy for numerical computing.
Definition
pandas provides DataFrames — two-dimensional tables with labeled rows and columns — for working with structured data.
NumPy supplies ndarrays, n-dimensional arrays that enable fast vectorized operations without explicit loops.
Intuition
Think of pandas as Excel with code — you get the flexibility of a spreadsheet but the repeatability of a script.
NumPy arrays are like Python lists but designed for math: they're homogeneous and support element-wise operations directly.
Worked example
pd.read_csv("sales.csv") loads a CSV into a DataFrame; df.groupby("region").sales.sum() aggregates by region.
np.array([1, 2, 3]) * 2 yields array([2, 4, 6]) — the multiplication applied element-wise automatically.
The math
A Series has an index and a dtype; its data can use NumPy arrays or extension arrays, including nullable and Arrow-backed types.
NumPy uses BLAS/LAPACK for optimized linear algebra under the hood.
In practice
The pandas/NumPy stack is the foundation for nearly all Python-based data analysis and machine learning.
Libraries like scikit-learn, statsmodels, and seaborn all consume pandas DataFrames or NumPy arrays.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.