commit	f2fcb80f8727c6b9620e2a84629d3e45d8c0e8f7	[log] [tgz]
author	Rich <jychen7@users.noreply.github.com>	Sun Apr 10 19:22:30 2022 -0400
committer	GitHub <noreply@github.com>	Sun Apr 10 19:22:30 2022 -0400
tree	dd8ffe7453549ccd41eefd6248a15c3980b453d7
parent	b1ef00b9f52124040506ca5ca158c50f503a4b98 [diff]

tree: dd8ffe7453549ccd41eefd6248a15c3980b453d7

README.md

DataFusion

DataFusion is an extensible query execution framework, written in Rust, that uses Apache Arrow as its in-memory format.

DataFusion supports both an SQL and a DataFrame API for building logical query plans as well as a query optimizer and execution engine capable of parallel execution against partitioned data sources (CSV and Parquet) using threads.

DataFusion also supports distributed query execution via the Ballista crate.

Use Cases

DataFusion is used to create modern, fast and efficient data pipelines, ETL processes, and database systems, which need the performance of Rust and Apache Arrow and want to provide their users the convenience of an SQL interface or a DataFrame API.

Why DataFusion?

High Performance: Leveraging Rust and Arrow's memory model, DataFusion achieves very high performance
Easy to Connect: Being part of the Apache Arrow ecosystem (Arrow, Parquet and Flight), DataFusion works well with the rest of the big data ecosystem
Easy to Embed: Allowing extension at almost any point in its design, DataFusion can be tailored for your specific usecase
High Quality: Extensively tested, both by itself and with the rest of the Arrow ecosystem, DataFusion can be used as the foundation for production systems.

Known Uses

Projects that adapt to or serve as plugins to DataFusion:

Here are some of the projects known to use DataFusion:

Ballista Distributed Compute Platform
Cloudfuse Buzz
Cube Store
delta-rs
InfluxDB IOx Time Series Database
ROAPI
Tensorbase
Squirtle
VegaFusion Server-side acceleration for the Vega visualization grammar

(if you know of another project, please submit a PR to add a link!)

Example Usage

Please see example usage to find how to use DataFusion.

Roadmap

Please see Roadmap for information of where the project is headed.

Architecture Overview

There is no formal document describing DataFusion's architecture yet, but the following presentations offer a good overview of its different components and how they interact together.

(March 2021): The DataFusion architecture is described in Query Engine Design and the Rust-Based DataFusion in Apache Arrow: recording (DataFusion content starts ~ 15 minutes in) and slides
(February 2021): How DataFusion is used within the Ballista Project is described in *Ballista: Distributed Compute with Rust and Apache Arrow: recording

User's guide

Please see User Guide for more information about DataFusion.

Developer's guide

Please see Developers Guide for information about developing DataFusion.