[!WARNING] As of May 2026 this project is undergoing significant rework! We do not advise using it until it is restablized and a formal release is announced. It has been marked as dormant by Apache Logging Services consensus on 2024-10-10. Users are advised to migrate to alternatives. For other inquiries, see the support policy.
Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of event data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tunable reliability mechanisms and many failover and recovery mechanisms. The system is centrally managed and allows for intelligent dynamic management. It uses a simple extensible data model that allows for online analytic application.
The Apache Flume Hadoop module provides Flume components that leverage Hadoop technologies.
Apache Flume Hadoop is open-sourced under the Apache Software Foundation License v2.0.
The Flume 2.x guide and FAQ are available here:
Compiling Flume Hadoop requires Java 17 or later; the Maven wrapper (./mvnw) downloads the right Maven version.
Bug and Issue tracker.