EMZETT.
Login

Data Sink

In short: The endpoint at which a data stream arrives and is processed or stored — the counterpart to the data source.

In more detail: In data pipelines, “source → sink” describes the path of the data from its origin (e.g. sensor, API, log file) to its destination (e.g. database, analytics tool, dashboard). The term makes it clear that data isn’t just transported but has to be consumed somewhere in the end.

In Depth

A role, not a product

The term comes from data processing (ETL pipelines, streaming systems such as Apache Kafka or Flink) and deliberately describes only the role in a data flow, not a specific product: the same database can act as a sink in one pipeline (data is stored there) and at the same time as a source in another pipeline (data is read from there and passed on to another system). A typical chain looks like this:

Source (sensor/API/log file) → processing/transformation → Sink (database/dashboard/data warehouse)

Push vs. pull sinks

The distinction between push and pull sinks is important: with a push sink, the source actively sends the data there without the sink having to ask on its own (e.g. an IoT sensor that sends measurements to an API at regular intervals, or a webhook that fires a message immediately on every event). With a pull sink, on the other hand, the sink actively fetches the data itself, usually at fixed intervals (e.g. an analytics tool or business intelligence dashboard that queries a database every night and imports the new records). Push models tend to deliver more up-to-date data with less delay, but put more load on the source; pull models are easier to control (the sink itself decides when it has capacity for new data), but lead to a delay between the creation and the actual processing of the data.

Multi-stage pipelines

Data sinks can themselves in turn serve as a source for the next processing stage — in more complex pipelines this creates multi-stage chains in which the same node is simultaneously the sink for the previous stage and the source for the next one. A typical example: raw data first lands in a data lake (sink, stage 1), is periodically read from there, cleaned and aggregated (the data lake becomes the source, a data warehouse the sink of stage 2), and a dashboard tool in turn reads from the data warehouse for the daily visualisation (the data warehouse becomes the source, the dashboard the final sink).

Error handling at the sink boundary

A practically important aspect of sinks is the handling of errors and retries: if writing to a sink fails (e.g. because the target database is briefly unreachable), the pipeline has to decide whether to discard the data, buffer it and retry later, or stop all upstream processing. Robust streaming systems often use a concept called a “dead letter queue” for this — a separate storage location for records that couldn’t be written successfully to the actual sink even after several attempts, so that they can be examined manually later without blocking the rest of the data flow.

See also: Transmission, API