AI
In short: Artificial Intelligence — umbrella term for systems that solve tasks that normally require human intelligence (pattern recognition, language, decisions).
In more detail: Covers a broad field, from classic rule-based systems to modern neural networks and large language models (LLMs). Machine learning — systems that learn from data instead of following fixed programmed rules — is currently the dominant subfield of AI.
In Depth
AI can roughly be divided into several nested subfields: classic rule-based/symbolic AI (fixed if-then rules programmed by humans, e.g. early chess programs), machine learning (systems learn patterns themselves from training data, instead of every rule being manually programmed), and, within that, deep learning (machine learning specifically using multi-layered neural networks, responsible for most current breakthroughs in image and speech recognition).
A neural network consists of layers of artificial “neurons” that weight and combine input values — during training, these weights are gradually adjusted based on many examples, so the network’s predictions increasingly match the known correct answers. Large language models (see Gen AI) are a specific, very large type of such neural networks, trained on huge amounts of text.
In practical software development, AI today mostly appears via ready-made models/APIs (e.g. an image-recognition API, a language-model endpoint), less often through training a model from scratch yourself — the latter requires large amounts of data and computing power, while most use cases can be covered with pre-trained models.
Historical development
The term “artificial intelligence” was coined in 1956 at the Dartmouth Conference. The field’s history ran in waves: early euphoria (symbolic AI, expert systems in the 1960s-80s) was repeatedly followed by “AI winters” — periods of disappointed expectations and cut research funding, because the computing power and data available at the time weren’t enough for the ambitious goals. The current AI boom since the 2010s rests on three converging factors: massively increased computing power (mainly through GPUs, originally developed for graphics computation but ideal for the parallel matrix operations of neural networks), huge available amounts of training data (through the internet), and algorithmic advances like the transformer architecture (2017), which first made today’s language models practical.
Training phase vs. inference
A central concept is the distinction between training (the model learns from data, very computationally intensive, often takes days to weeks on specialised hardware) and inference (the fully trained model is applied to new inputs, considerably less computationally intensive, often runs in milliseconds). This distinction explains why most companies don’t train their own models: using a finished, already-trained model via API (pure inference) is orders of magnitude cheaper and faster to implement than training your own.
Limits and risks
AI systems, especially machine learning models, inevitably inherit biases from their training data — a model trained on historical application data, for example, can reproduce discriminatory patterns without that being explicitly programmed in. Many modern AI models are also “black boxes”: the exact reasoning behind a single prediction often can’t be fully traced, which raises legal and ethical questions in regulated areas (loan approval, medical diagnostics).
See also: Gen AI, API, Data Science