0.5.3 release tag
-----BEGIN PGP SIGNATURE-----

iQIzBAABCAAdFiEEr5uvedMRo9Mojlg/JKSZA3JiqqQFAl71kPkACgkQJKSZA3Ji
qqRnXA//ffD65nDNjy6bK5aOlLoQmNyFqoaAIHHOOhqJUljXmOJKlEorjCn3xJqr
REWeD2mEXPv5EXHJvnfJNQesaJxBrrb6YJ1p9QUtiqtHr9uulZfRdfMOBFTHafCC
lD5D7Q5Eg4mar3pPI3Uf2lmZPhLaoswmF0PjvocYX9l6Mj23YbnxIOUtp3XIbBMl
ev/xsVLIA8+sD3FCHZ7+o+T1JLfDN1hSY2/uZZfhaoxzcDrvwn5OICA3XV5EvID1
m3cCcWnAh62n0eiHFcVNtqNvNs2YlGfoNpSYmL9dPS/o/MsX4PlFo68VRrL87s6t
Fs5H7GhVLBUQNZM1NL7WQhqle7OYxOSaXq1mucSiZ7YPWx6y+js0igCAeFHISbtJ
VSDKqt6KbjInahK0vRVhbskMcWVMrbmRSJq596xskgWe3oar2OGa9l2Ubzzp9AKm
sOX/uK7hOc6QO8TRLy3JCQzo4OD6NyUvIxTKMu8i6PcMAXLfBf4DMNaAnz1FMnNb
Zti6jo4NHJfMnjuoYCuypNcUinEyvJJJrdvqgK30RIwa2nhumfNN6cd77IHkWUff
cEbtbvSxfuyxP1ozloEIQ9PkxJrebDZJ8wVqADZcNGatdK30/XkqlSwrSJSwIUJd
09/iYXniwXQs9M01vDSVUA7IiByFSOHCyIOEDFamz+MO/HdCFrs=
=GQDA
-----END PGP SIGNATURE-----
[MINOR] Update release version to reflect published version 0.5.3
27 files changed
tree: 46546126f8e3cdaeed0847c6059c71b626706cdb
  1. .github/
  2. .gitignore
  3. .travis.yml
  4. LICENSE
  5. NOTICE
  6. README.md
  7. doap_HUDI.rdf
  8. docker/
  9. hudi-cli/
  10. hudi-client/
  11. hudi-common/
  12. hudi-hadoop-mr/
  13. hudi-hive/
  14. hudi-integ-test/
  15. hudi-spark/
  16. hudi-timeline-service/
  17. hudi-utilities/
  18. packaging/
  19. pom.xml
  20. scripts/
  21. style/
README.md

Apache Hudi

Apache Hudi (pronounced Hoodie) stands for Hadoop Upserts Deletes and Incrementals. Hudi manages the storage of large analytical datasets on DFS (Cloud stores, HDFS or any Hadoop FileSystem compatible storage).

https://hudi.apache.org/

Build Status License Maven Central Join on Slack

Features

  • Upsert support with fast, pluggable indexing
  • Atomically publish data with rollback support
  • Snapshot isolation between writer & queries
  • Savepoints for data recovery
  • Manages file sizes, layout using statistics
  • Async compaction of row & columnar data
  • Timeline metadata to track lineage

Hudi supports three types of queries:

  • Snapshot Query - Provides snapshot queries on real-time data, using a combination of columnar & row-based storage (e.g Parquet + Avro).
  • Incremental Query - Provides a change stream with records inserted or updated after a point in time.
  • Read Optimized Query - Provides excellent snapshot query performance via purely columnar storage (e.g. Parquet).

Learn more about Hudi at https://hudi.apache.org

Building Apache Hudi from source

Prerequisites for building Apache Hudi:

  • Unix-like system (like Linux, Mac OS X)
  • Java 8 (Java 9 or 10 may work)
  • Git
  • Maven
# Checkout code and build
git clone https://github.com/apache/hudi.git && cd hudi
mvn clean package -DskipTests -DskipITs

To build the Javadoc for all Java and Scala classes:

# Javadoc generated under target/site/apidocs
mvn clean javadoc:aggregate -Pjavadocs

Build with Scala 2.12

The default Scala version supported is 2.11. To build for Scala 2.12 version, build using scala-2.12 profile

mvn clean package -DskipTests -DskipITs -Dscala-2.12

Quickstart

Please visit https://hudi.apache.org/docs/quick-start-guide.html to quickly explore Hudi's capabilities using spark-shell.