Apache Flume Hadoop provides various Flume components for the Hadoop ecosystem

Clone this repo:
  1. a5298ac Use the reusable CodeQL workflow from `logging-parent` by Piotr P. Karwasz · 3 weeks ago main
  2. 30d3e51 Bootstrap build and release infrastructure (#4) by Piotr P. Karwasz · 3 weeks ago
  3. 203a4bc Apply the `logging-parent` Spotless formatting (#6) by Piotr P. Karwasz · 3 weeks ago
  4. b9c0857 Route GitHub notifications away from the dev list (#5) by Piotr P. Karwasz · 3 weeks ago
  5. 2313700 Move Hive metastore to a test directory. Eliminate reload4j by Ralph Goers · 2 years, 5 months ago

Project status

[!WARNING] As of May 2026 this project is undergoing significant rework! We do not advise using it until it is restablized and a formal release is announced. It has been marked as dormant by Apache Logging Services consensus on 2024-10-10. Users are advised to migrate to alternatives. For other inquiries, see the support policy.

Welcome to Apache Flume Hadoop!

Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of event data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tunable reliability mechanisms and many failover and recovery mechanisms. The system is centrally managed and allows for intelligent dynamic management. It uses a simple extensible data model that allows for online analytic application.

The Apache Flume Hadoop module provides Flume components that leverage Hadoop technologies.

Apache Flume Hadoop is open-sourced under the Apache Software Foundation License v2.0.

Documentation

The Flume 2.x guide and FAQ are available here:

Compiling Flume Hadoop

Compiling Flume Hadoop requires Java 17 or later; the Maven wrapper (./mvnw) downloads the right Maven version.

Contact us!

Bug and Issue tracker.