Pandas
In short: The standard Python library for tabular data — loads, filters, groups and transforms data similar to a spreadsheet, just programmable.
In more detail: Pandas’s central data structure is the DataFrame (a table with named columns), built on NumPy. It lets you read in, clean, aggregate and export CSV/Excel/database data — the de-facto standard for data analysis in Python.
In Depth
Basic structure: DataFrame and Series
import pandas as pd
df = pd.read_csv("sales.csv")
revenue_per_category = df.groupby("category")["revenue"].sum()
# Filtering
big_sales = df[df["revenue"] > 1000]
# Computing a new column
df["revenue_with_tax"] = df["revenue"] * 1.19
# Handling missing values
df = df.dropna(subset=["customer_id"])A DataFrame consists of named columns (each column itself a Series object, technically a NumPy array of a uniform type) and an index for the rows — conceptually similar to a spreadsheet, but fully controllable programmatically and without the row limits of an Excel file (Pandas easily handles millions of rows, as long as memory is sufficient). A single column on its own (df["revenue"]) is a Series — one-dimensional, with the same index as the DataFrame it came from.
Typical operations
- Filtering: boolean indexing like
df[df["revenue"] > 1000]— reads almost like a SQLWHEREclause. - Grouping and aggregating:
groupbycorresponds to SQLGROUP BY, followed by an aggregation function (sum,mean,count). - Merging:
mergecorresponds to a SQLJOINbetween two DataFrames based on shared columns. - Reshaping: pivot tables (
pivot_table) turn rows into columns and vice versa, useful for cross-tabulation analyses. - Time series: built-in functions for date/time columns (resampling to daily/monthly level, timezone conversion) — one reason Pandas was originally developed for financial data analysis (the name stands for “panel data”).
Role in the data science workflow
In a typical data science workflow, Pandas usually handles the cleaning and preprocessing phase (“data wrangling”): handling missing values (dropping or filling them), fixing data types (e.g. converting text dates into real date objects), renaming or recomputing columns, removing duplicates — before the prepared data is passed on to more specialised libraries like SciPy (statistical tests) or machine learning tools (scikit-learn). In practice, data science teams often spend most of a project’s time exactly in this Pandas-heavy cleaning phase, not in the actual model training.
See also: NumPy, Data Science, SciPy