[fix](fe) Fix TopN lazy materialization for queries ordered by an alias (#68019)

### What problem does this PR solve?

Issue Number: None

Related PR: None

Problem Summary:

`SELECT lazy_col AS x, lazy_col AS y FROM t ORDER BY x LIMIT 1` failed
planning with
`A expression contains slot not from children`.

The TopN order key is the alias slot, so that alias has to be computed
below the TopN.
`MaterializeProbeVisitor` only protects the order key slot itself (an
order key slot is in
`TopN.getInputSlots()`) and never resolves an identity alias down to the
column the alias reads.
The probe of the other output (`lazy_col AS y`) therefore resolved to
the base column `lazy_col`
and classified it as lazily materialized, so `LazySlotPruning` removed
`lazy_col` from the scan while
`lazy_col AS x` below the TopN still read it. The final `Validator`
rejected the resulting plan and
the query returned an error. With `fe_debug=true` the failure was caught
inside `LazyMaterializeTopN`
instead, which silently skipped lazy materialization (the query
succeeded but lost the optimization).

Reproduction (master, `fe_debug=false`):

```sql
create table t(sort_col int, lazy_col int) duplicate key(sort_col)
  distributed by hash(sort_col) buckets 1 properties('replication_num'='1');
select lazy_col as x, lazy_col as y from t order by x limit 1;
-- ERROR 1105: A expression contains slot not from children
--   Slot: lazy_col#1  Children Output:{0, 4}
--   Plan: PhysicalProject[lazy_col#1 AS x#2, __DORIS_GLOBAL_ROWID_COL__t#4]
--         +--PhysicalLazyMaterializeOlapScan[PhysicalOlapScan[t]]
```

Fix: `LazyMaterializeTopN` resolves the TopN order keys through the
identity alias chain of the
Projects under the TopN and adds the resolved slots (plus the
intermediate alias slots) to
`requiredMaterializedSlots`, so the probe rejects every lazy candidate
backed by a column an order
key reads. The resolution stops at set operations, which the probe never
materializes through
(lazy materialization is not supported through set operations today; if
that ever changes, order
keys have to be resolved per branch).

Effect: affected plans now either keep only the ordering column
materialized (other columns are
still fetched lazily) or skip lazy materialization, and the plan stays
valid. Plans that order by a
plain column are unchanged.

### Release note

TopN lazy materialization no longer builds an invalid plan (no more
`A expression contains slot not from children`) when a query orders by
an alias of a column.
The column that feeds the order key is materialized during the scan,
while other columns keep using
lazy materialization.

### Check List (For Author)

- Test
- [x] Regression test
(`regression-test/suites/query_p0/topn_lazy/order_by_alias`)
- [x] Unit Test (`TopnLazyMaterializeTest`, `LazyMaterializeTopNTest`)
- Behavior changed:
- [x] Yes. Queries that order by an alias of a projected column no
longer fail planning; the
ordering column is kept materialized instead of being pruned from the
scan.
- Does this need documentation?
    - [x] No.

### Check List (For Reviewer who merge this PR)

- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label
7 files changed
tree: 3b0918dd7df39c7f8b99af049f1ceba110f957cb
  1. .claude/
  2. .github/
  3. .idea/
  4. be/
  5. bin/
  6. build-support/
  7. cloud/
  8. common/
  9. conf/
  10. contrib/
  11. dist/
  12. docker/
  13. docs/
  14. extension/
  15. fe/
  16. fe_plugins/
  17. fs_brokers/
  18. gensrc/
  19. hooks/
  20. pytest/
  21. regression-test/
  22. samples/
  23. task_executor_simulator/
  24. thirdparty/
  25. tools/
  26. ui/
  27. webroot/
  28. .asf.yaml
  29. .clang-format
  30. .clang-format-ignore
  31. .clang-tidy
  32. .clangd
  33. .dockerignore
  34. .editorconfig
  35. .gitattributes
  36. .gitignore
  37. .gitleaks.toml
  38. .gitmodules
  39. .licenserc.yaml
  40. .rat-excludes
  41. .shellcheckrc
  42. AGENTS.md
  43. build-for-release.sh
  44. build-plugin.sh
  45. build.sh
  46. build_profile.sh
  47. CODE_OF_CONDUCT.md
  48. CONTRIBUTING.md
  49. CONTRIBUTING_CN.md
  50. doap_Doris.rdf
  51. env.sh
  52. generated-source.sh
  53. LICENSE.txt
  54. NOTICE.txt
  55. post-build.sh
  56. README.md
  57. reset_submodule.sh
  58. run-be-ut.sh
  59. run-cloud-ut.sh
  60. run-fe-ut.sh
  61. run-fs-env-test.sh
  62. run-regression-test.sh
  63. SECURITY.md
  64. sonar-project.properties
  65. threat-model.md
README.md

🌍 Read this in other language

EnglishالعربيةবাংলাDeutschEspañolفارسیFrançaisहिन्दीBahasa IndonesiaItaliano日本語한국어PolskiPortuguêsRomânăРусскийSlovenščinaไทยTürkçeУкраїнськаTiếng Việt简体中文繁體中文

Apache Doris

License GitHub release Slack EN doc CN doc

Official Website Quick Download


Apache Doris is an open-source, real-time analytics and search database built on MPP architecture. It provides fast SQL analytics, lakehouse query acceleration, and hybrid search across structured, text, and vector data.

Explore the official website for the latest product overview, use cases, ecosystem updates, blogs, and user stories. For version updates, see all release notes.

📈 Use Cases

Use CaseWhat it provides
Customer-Facing AnalyticsShip sub-second interactive analytics to external users.
Data WarehousingBuild one real-time warehouse across business domains.
ObservabilityAnalyze high-throughput logs, events, and metrics with SQL.
Doris for AIUse vector, text, JSON, and structured search in one SQL engine.

🚀 Core Capabilities

Apache Doris is built around three core capabilities. The website is the source of truth for detailed product descriptions and examples.

CapabilityWhat it provides
Real-Time AnalyticsStreaming ingestion, incremental transformation, and sub-second queries under high concurrency.
Lakehouse AnalyticsFast SQL analytics over open table formats such as Iceberg, Delta Lake, and Hudi.
Hybrid SearchSQL-native analytics across JSON, full-text, and vector data for AI and search workloads.

🔌 Ecosystem

Doris sits at the center of the modern data stack. It connects upstream databases, streaming systems, and lakehouse storage with downstream BI, AI, analytics, and observability tools.

For the latest ecosystem coverage, visit the official website and the connection and integration documentation.

👣 Get Started

🧱 Architecture

Apache Doris supports both compute-storage coupled and compute-storage decoupled deployments. In decoupled mode, stateless compute groups run over shared object storage, so you can scale compute on demand and isolate workloads.

Learn more in the deployment guide and deployment mode guide.

📣 Project Updates

ResourceWhat it provides
Community ReportWeekly updates on community activity, merged PRs, contributors, and feature progress.
Roadmap 2026The 2026 planning discussion for AI and hybrid search, query engine, storage, and data lake work.

🧩 Components

Doris provides connectors and tools for common data engineering workflows.

👨‍👩‍👧‍👦 Users

Apache Doris is used in production by thousands of companies worldwide across internet services, finance, retail, logistics, manufacturing, energy, telecommunications, AI, and other industries.

🙌 Contributors

Apache Doris graduated from the Apache Incubator and became an Apache Top-Level Project in June 2022. Thanks to all community contributors who help build Doris.

contrib graph

🌈 Community and Support

💬 Contact Us

🧰 Links

📜 License

Apache License, Version 2.0

Note Some licenses of the third-party dependencies are not compatible with Apache 2.0 License. So you need to disable some Doris features to comply with Apache 2.0 License. For details, refer to the thirdparty/LICENSE.txt