Kafka is a high-throughput distributed messaging system with publish and subscribe capabilities. It provides persistence with replication to disk for fault tolerance. Kafka is simple to implement and runs efficiently on large clusters with low latency and high throughput. It was created at LinkedIn to process streaming data from the LinkedIn website and has since been open sourced.