Glossary · Automation fundamentals, platforms and components
Apache Kafka
Also known as: Kafka
German: Apache Kafka
In software engineering, Apache Kafka is an open-source distributed event streaming platform maintained by the Apache Software Foundation. Producers write records to topics, which are stored as partitioned, replicated logs that consumers read at their own pace.
- Automation components
- Vendor product
In one sentence
Apache Kafka is an open-source event streaming platform that stores records in partitioned, replicated topics for producers and consumers.
Example
Machine data from an MQTT broker is bridged into a Kafka topic, where a quality analytics service and a data lake loader each consume it independently.
How it applies
- Engineering: In manufacturing IT, Kafka often sits above the OT layer and distributes production events to MES, analytics and cloud services. It is not a real-time fieldbus and is not used for deterministic control.
- Operation: Retention settings decide how long events stay available for replay; consumer groups track what each application has read.
- Documentation: Describe topic names, message schemas, retention and ownership. When messages carry product or quality data, document which system is the Source of truth and how schema changes are versioned.
Kafka vs. MQTT
MQTT is a lightweight publish/subscribe protocol for devices and gateways, usually without long-term storage. Kafka is a storage-centric streaming platform for data centers and the cloud. Many architectures use both: MQTT at the edge, Kafka in the backbone.
Keep in mind: Product names, features and managed-service offerings around Kafka change. Check the current Apache Kafka documentation and your platform vendor's documentation before relying on a specific feature.