Glossary · IIoT, data and AI
Data pipeline
Also known as: ETL pipeline
German: Datenpipeline
In data engineering, a data pipeline is an automated sequence of steps that moves data from one or more sources to a destination, extracting, validating, transforming and loading it along the way. Pipelines can run in batches or continuously as streams.
- IIoT
In one sentence
A data pipeline automatically moves data from sources to destinations, validating and transforming it in batches or as a stream.
Example
A pipeline reads machine states from an MQTT broker every second, enriches them with order data from the MES and writes them to a time-series database.
How it applies
- Engineering: Design pipelines for failure: buffering during connection loss, handling of late or duplicate values, and clear behavior when a source changes its format.
- Operation: Monitor throughput, latency and error rates. A silently broken pipeline produces dashboards that look normal but show old data.
- Documentation: Document sources, transformations, schedules, owners and error handling. Pipeline definitions stored as code can serve as part of the Data lineage.
Batch vs. streaming pipeline
A batch pipeline processes data in chunks at intervals, for example nightly reports. A streaming pipeline processes each event as it arrives, for example for near real-time monitoring. Neither is suited for closed-loop control, which stays in the controller. In IIoT architectures, the first pipeline stage often runs on an Edge device or IoT gateway close to the machine.