[feature](ivm) Support incremental view maintenance (IVM) for materialized tables (MTMV) (#62606)

### What problem does this PR solve?

Implement Incremental View Maintenance (IVM) for materialized tables
(MTMV) in Apache Doris. When a base table of an MTMV changes, only the
affected rows are derived from the base-table deltas (row binlog / Table
Stream) and transactionally applied to the materialized table, instead
of recomputing the whole result with a full refresh. This enables
maintaining near-real-time aggregation and join results at a fraction of
the cost of full refresh.

Supported operators: projection, filter, aggregate
(COUNT/SUM/AVG/MIN/MAX, bitmap aggregates, expression arguments, bare
GROUP BY, grouping sets), joins (INNER/CROSS/OUTER JOIN, nested join and
self join), UNION ALL, subquery aliases, and OneRowRelation.

Supported refresh: `REFRESH MATERIALIZED VIEW ...
INCREMENTAL/PARTITIONS`, refresh explain, dry-run, and automatic
full-refresh fallback when incremental maintenance is not valid (e.g.
binlog broken).

Lifecycle: stream creation/lifecycle integrated with MTMV, chained IVM
MTMV support.

This PR is the MTMV incremental-maintenance layer of the end-to-end
incremental computation stack tracked by #65418.

Issue Number: close #66719

Related issue: #65418 (tracking issue), #57921 (superseded umbrella
issue)

---------

Co-authored-by: seawinde <wusi@selectdb.com>
347 files changed
tree: 412f950ff70207e5e3db3ad4970521d0fb3c34e4
  1. .claude/
  2. .github/
  3. .idea/
  4. be/
  5. bin/
  6. build-support/
  7. cloud/
  8. common/
  9. conf/
  10. contrib/
  11. dist/
  12. docker/
  13. docs/
  14. extension/
  15. fe/
  16. fe_plugins/
  17. fs_brokers/
  18. gensrc/
  19. hooks/
  20. pytest/
  21. regression-test/
  22. samples/
  23. task_executor_simulator/
  24. thirdparty/
  25. tools/
  26. ui/
  27. webroot/
  28. .asf.yaml
  29. .clang-format
  30. .clang-format-ignore
  31. .clang-tidy
  32. .clangd
  33. .dockerignore
  34. .editorconfig
  35. .gitattributes
  36. .gitignore
  37. .gitleaks.toml
  38. .gitmodules
  39. .licenserc.yaml
  40. .rat-excludes
  41. .shellcheckrc
  42. AGENTS.md
  43. build-for-release.sh
  44. build-plugin.sh
  45. build.sh
  46. build_profile.sh
  47. CODE_OF_CONDUCT.md
  48. CONTRIBUTING.md
  49. CONTRIBUTING_CN.md
  50. doap_Doris.rdf
  51. env.sh
  52. generated-source.sh
  53. LICENSE.txt
  54. NOTICE.txt
  55. post-build.sh
  56. README.md
  57. reset_submodule.sh
  58. run-be-ut.sh
  59. run-cloud-ut.sh
  60. run-fe-ut.sh
  61. run-fs-env-test.sh
  62. run-regression-test.sh
  63. SECURITY.md
  64. sonar-project.properties
  65. threat-model.md
README.md

🌍 Read this in other language

English • العربية • বাংলা • Deutsch • Español • فارسی • Français • हिन्दी • Bahasa Indonesia • Italiano • 日本語 • 한국어 • Polski • Português • Română • Русский • Slovenščina • ไทย • Türkçe • Українська • Tiếng Việt • 简体中文 • 繁體中文

Apache Doris

License GitHub release Slack EN doc CN doc

Official Website Quick Download


Apache Doris is an open-source, real-time analytics and search database built on MPP architecture. It provides fast SQL analytics, lakehouse query acceleration, and hybrid search across structured, text, and vector data.

Explore the official website for the latest product overview, use cases, ecosystem updates, blogs, and user stories. For version updates, see all release notes.

📈 Use Cases

Use CaseWhat it provides
Customer-Facing AnalyticsShip sub-second interactive analytics to external users.
Data WarehousingBuild one real-time warehouse across business domains.
ObservabilityAnalyze high-throughput logs, events, and metrics with SQL.
Doris for AIUse vector, text, JSON, and structured search in one SQL engine.

🚀 Core Capabilities

Apache Doris is built around three core capabilities. The website is the source of truth for detailed product descriptions and examples.

CapabilityWhat it provides
Real-Time AnalyticsStreaming ingestion, incremental transformation, and sub-second queries under high concurrency.
Lakehouse AnalyticsFast SQL analytics over open table formats such as Iceberg, Delta Lake, and Hudi.
Hybrid SearchSQL-native analytics across JSON, full-text, and vector data for AI and search workloads.

🔌 Ecosystem

Doris sits at the center of the modern data stack. It connects upstream databases, streaming systems, and lakehouse storage with downstream BI, AI, analytics, and observability tools.

For the latest ecosystem coverage, visit the official website and the connection and integration documentation.

👣 Get Started

🧱 Architecture

Apache Doris supports both compute-storage coupled and compute-storage decoupled deployments. In decoupled mode, stateless compute groups run over shared object storage, so you can scale compute on demand and isolate workloads.

Learn more in the deployment guide and deployment mode guide.

📣 Project Updates

ResourceWhat it provides
Community ReportWeekly updates on community activity, merged PRs, contributors, and feature progress.
Roadmap 2026The 2026 planning discussion for AI and hybrid search, query engine, storage, and data lake work.

🧩 Components

Doris provides connectors and tools for common data engineering workflows.

👨‍👩‍👧‍👦 Users

Apache Doris is used in production by thousands of companies worldwide across internet services, finance, retail, logistics, manufacturing, energy, telecommunications, AI, and other industries.

🙌 Contributors

Apache Doris graduated from the Apache Incubator and became an Apache Top-Level Project in June 2022. Thanks to all community contributors who help build Doris.

contrib graph

🌈 Community and Support

💬 Contact Us

🧰 Links

📜 License

Apache License, Version 2.0

Note Some licenses of the third-party dependencies are not compatible with Apache 2.0 License. So you need to disable some Doris features to comply with Apache 2.0 License. For details, refer to the thirdparty/LICENSE.txt