Glossary · IIoT, data and AI
Data lake
German: Data Lake
In data engineering, a data lake is a central repository that stores large amounts of raw data in its original format, structured, semi-structured or unstructured, until it is needed for analysis. The schema is applied when the data is read, not when it is stored.
- IIoT
In one sentence
A data lake stores large amounts of raw data in its original format and applies structure only when the data is read for analysis.
Example
A plant writes raw OPC UA values, camera images and MES events into object storage, where data scientists later select the parts they need.
How it applies
- Engineering: A data lake is flexible but not self-explanatory. Store data with asset identifiers, units, time zones and source information, otherwise it becomes a data swamp.
- Operation: Access rights, retention periods and data catalogs are needed from the start, especially when personal data or confidential recipes are included.
- Documentation: Maintain a catalog that describes each data set: origin, meaning, owner and update rate. Link it to Data lineage where possible.
Data lake vs. data warehouse
A Data warehouse stores cleaned, integrated data in a predefined schema optimized for reporting (schema on write). A data lake stores raw data and leaves structuring to the reader (schema on read). Many organizations combine both, sometimes in a so-called lakehouse architecture.