EMZETT.
Login

NumPy

In short: The fundamental Python library for numerical computing with arrays and matrices — the basis of almost all data science and machine learning libraries in Python.

In more detail: NumPy arrays are considerably faster than normal Python lists for mathematical operations, because they’re internally implemented in C and perform vector/matrix operations without Python loops (“vectorisation”). Libraries like Pandas or SciPy build directly on NumPy.

In Depth

import numpy as np
a = np.array([1, 2, 3, 4])
b = a * 2  # [2, 4, 6, 8] - without an explicit loop

The speed advantage over plain Python lists comes from two sources: NumPy arrays store their elements compactly and with a uniform type in memory (instead of referencing many individual, scattered objects like Python lists), and operations like a * 2 run internally as optimised, compiled C code over the entire array, instead of being processed element by element in a slow Python loop (“vectorisation”). For large amounts of data, this can be 10-100x faster than equivalent pure Python code.

NumPy forms the technical foundation for almost the entire Python data science ecosystem: Pandas internally stores table columns as NumPy arrays, SciPy and most machine learning libraries (scikit-learn, TensorFlow, PyTorch) build directly on NumPy’s array format and conventions — anyone using one of these libraries in production is usually also indirectly working with NumPy.

Broadcasting

A central concept that’s often initially unfamiliar to beginners is “broadcasting” — NumPy can automatically and sensibly combine operations between arrays of different shapes, without loops or explicit resizing being necessary. A single scalar, for example, is automatically applied to every element of an array (array * 2), and a smaller array can, under certain rules, be “virtually” expanded to the shape of a larger one, without memory actually being copied. This makes many mathematical operations compactly readable, but requires understanding the broadcasting rules to avoid unexpected behaviour with incompatible array shapes.

Multi-dimensional arrays and axes

Unlike simple Python lists, NumPy natively supports multi-dimensional arrays (matrices, tensors) with any number of dimensions (“axes”). Operations can be specifically performed along a particular axis — e.g. the sum of every column instead of the sum of all values:

matrix = np.array([[1, 2, 3], [4, 5, 6]])
matrix.sum(axis=0)  # sum per column: [5, 7, 9]
matrix.sum(axis=1)  # sum per row: [6, 15]

This axis concept is fundamental for working with image data (height × width × colour channels), time series, or the multi-dimensional tensors processed in machine learning frameworks like TensorFlow and PyTorch — whose own tensor objects deliberately adopt many conventions directly from NumPy, to make switching over easy.

Performance limits

Despite the speed advantages over pure Python, NumPy also hits limits: for extremely large amounts of data that no longer fit in memory, or for calculations that have to be distributed across several machines, complementary tools like Dask (distributed NumPy-compatible computing) or specialised GPU-accelerated libraries (CuPy) are used, offering the same API but running on different hardware or distributed across several machines.

See also: Pandas, SciPy, Data Science