This project builds the Rust-powered core for PyPaimon while also providing DataFusion integration for querying Paimon tables.
Install via PyPI:
pip install pypaimon-rust
If you want to use the native Python DataFusion SessionContext, install datafusion as well.
The recommended way to query Paimon tables is through SQLContext, which supports multi-catalog registration, DDL, DML, and all Paimon-specific SQL extensions:
from pypaimon_rust.datafusion import SQLContext ctx = SQLContext() ctx.register_catalog("paimon", { "warehouse": "/path/to/warehouse", }) # DDL ctx.sql("CREATE SCHEMA paimon.my_db") ctx.sql("CREATE TABLE paimon.my_db.users (id INT, name STRING, PRIMARY KEY (id))") # DML ctx.sql("INSERT INTO paimon.my_db.users VALUES (1, 'alice'), (2, 'bob')") # Query tables via SQL (catalog.database.table) batches = ctx.sql("SELECT * FROM paimon.my_db.users")
SQLContext also registers Python scalar UDFs for common BLOB media and vector workflows. Install the optional media dependencies when you need image or video decoding:
pip install "pypaimon-rust[video]"
batches = ctx.sql(""" SELECT id, media_info(content) AS info_json, media_thumbnail(content, 160, 90) AS preview_png, video_snapshot(content, 1000) AS frame_png FROM paimon.my_db.assets """)
For vector search, vector_from_json converts JSON-encoded embeddings into Arrow List<Float32> values that can be used from a temporary table or subquery:
batches = ctx.sql(""" WITH queries AS ( SELECT id, vector_from_json(embedding_json) AS embedding FROM paimon.my_db.query_embeddings ) SELECT q.id AS query_id, r.id AS result_id FROM queries q CROSS JOIN LATERAL vector_search( 'paimon.my_db.items', 'embedding', q.embedding, 10 ) AS r """)
You can register temporary in-memory tables programmatically. Names support the same resolution rules as SQL: bare names use the current catalog and database, partially qualified names use the current catalog, and fully qualified names specify catalog.database.table.
Register a single PyArrow RecordBatch as a temporary table:
import pyarrow as pa batch = pa.record_batch([[1, 2], ["alice", "bob"]], names=["id", "name"]) ctx.register_batch("paimon.default.my_temp", batch) batches = ctx.sql("SELECT * FROM paimon.default.my_temp") # Drop it via SQL when no longer needed ctx.sql("DROP TEMPORARY TABLE paimon.default.my_temp")
Alternatively, if you want to use the native Python DataFusion SessionContext, install datafusion and register a PaimonCatalog:
from datafusion import SessionContext from pypaimon_rust.datafusion import PaimonCatalog catalog = PaimonCatalog({ "warehouse": "/path/to/warehouse", }) ctx = SessionContext() ctx.register_catalog_provider("paimon", catalog) # Query tables via SQL (catalog.database.table) df = ctx.sql("SELECT * FROM paimon.default.my_table LIMIT 10") df.show()
from datafusion import SessionContext from pypaimon_rust.datafusion import PaimonCatalog catalog = PaimonCatalog({ "metastore": "rest", "uri": "http://localhost:8080", "warehouse": "my_warehouse", }) ctx = SessionContext() ctx.register_catalog_provider("paimon", catalog)
Time travel queries are not supported in the Python binding at this time.