Glossary · AI engineering and governance
Data drift
Also known as: Covariate shift, Input drift
German: Datendrift
In machine learning, data drift is a change over time in the statistical distribution of a model's input data compared with the data the model was trained on, for example because of new products, sensors, operating modes or environmental conditions.
- Industrial AI
- AI
In one sentence
Data drift is a change in the statistical distribution of a model's input data compared with the data it was trained on.
Example
After a camera is replaced with a newer model, the images are brighter and sharper, and the input distribution of the inspection model shifts noticeably.
How it applies
- Engineering: Data drift can often be detected without ground truth by comparing input statistics in operation with those of the training data, which makes it a useful early warning.
- Operation: Not every data drift harms performance. Assess the impact before retraining, and check whether the cause is a process change or a measurement fault such as Sensor drift.
- Documentation: Record changes to sensors, cameras, materials and process settings, since they are common causes of data drift. Document the monitored statistics and thresholds.
Data drift vs. concept drift
Data drift changes the distribution of inputs. Concept drift changes the relationship between inputs and the target. A model can suffer from either or both; only concept drift can occur with unchanged input statistics.