EMZETT.
Login

Data Science

In short: The field of extracting insights from (often large) amounts of data using statistical and programming methods — an intersection of statistics, programming and domain expertise in the respective field.

In more detail: A typical data science workflow covers data acquisition, cleaning (often the most time-consuming step), analysis/visualisation, and frequently building predictive models (machine learning). In Python, the standard trio for this: Pandas (data), NumPy (numerics), SciPy/scikit-learn (advanced methods).

In Depth

The classic data science workflow can roughly be divided into phases: data acquisition (from databases, APIs, files), cleaning (missing values, outliers, duplicates — in practice often 60-80% of the actual working time), exploratory analysis (understanding distributions, recognising patterns, usually with visualisation), modelling (statistical models or machine learning methods), and finally communicating the results (dashboards, reports, presentations).

Python dominates the field thanks to a well-established toolset: Pandas for tabular data (DataFrames, similar to a spreadsheet, but programmable), NumPy for fast numerical array operations underneath, SciPy and scikit-learn for statistical tests and classic machine learning, and Jupyter Notebooks as an interactive development environment where code, output and visualisation live side by side in one document. For larger amounts of data that no longer fit in memory, distributed frameworks like Spark are used.

Data science overlaps heavily with, but isn’t identical to, AI/machine learning: data science is the broader process of data analysis and generating insight, machine learning is one of the possible tools within it (not every data science question needs a trained model — often a good visualisation or a statistical test is enough).

The data science team

In larger companies, the data science role is often split further: data engineers build and maintain the infrastructure that lets data flow reliably at all (pipelines, databases, data warehouses); data scientists analyse this data and build models; data analysts focus more on reporting and visualisation for business decisions rather than machine learning modelling. In smaller teams or startups, a single person often takes on all three roles at once.

See also: Pandas, NumPy, SciPy, AI