[SPARK-58941][SDP] Sort schema inference flows by identifier parts to avoid dotted-name collisions

### What changes were proposed in this pull request?

`SchemaInferenceUtils.inferSchemaFromFlows` merges the flows that write to a table in a
deterministic order so that, when two flows emit a column whose names differ only in case, the
surviving spelling is well-defined and every caller agrees on it. It established that order by
sorting on `flow.identifier.unquotedString`.

`TableIdentifier.unquotedString` joins the identifier's name parts with an **unescaped** `.`. That
mapping is not one-to-one: two structurally distinct identifiers whose parts contain dots (a dot is
legal inside a back-tick-quoted schema or flow name) can render to the same string -- e.g.
`` `c`.`a.b`.`x` `` and `` `c`.`a`.`b.x` `` both render to `c.a.b.x`. Because `sortBy` is stable,
colliding keys fall back to the incoming `flows` order, which is the nondeterministic
flow-resolution completion order the sort was meant to remove, so the surviving column casing could
flip between runs.

This PR sorts on the identifier's parts `(catalog, database, table)` instead. Comparing the parts
tuple is injective -- two distinct identifiers are two distinct tuples, so they can never collide.
It also preserves the existing order for dotless identifiers (a shorter part sorts first, exactly
as the `.` separator, 0x2E, ordered a joined string), unlike `TableIdentifier.quotedString`, which
would additionally reorder some ordinary pairs because `` ` `` (0x60) sorts after name characters.
Flow identifiers are always fully qualified (`assertIsFullyQualifiedForCreate`), so all three parts
are present and the ordering never depends on an absent catalog/database.

This is a follow-up to SPARK-58517, which added the sort.

### Why are the changes needed?

The sort exists to guarantee a deterministic surviving column spelling. `unquotedString` is a lossy
key, so under identifiers that contain dots the guarantee silently breaks: the merge order reverts
to the nondeterministic completion order of concurrent flow resolution, and a case-only column
spelling can flip between runs of an unchanged pipeline. On the non-merging evolution paths
`diffSchemas` keys column identity on the exact name, so a run-to-run flip surfaces as a
`deleteColumn` + `addColumn` for a column that only changed case. Sorting on the identifier parts is
injective and removes the collision.

### Does this PR introduce _any_ user-facing change?

No. The merge order changes only for identifiers with dotted (back-tick-quoted) name parts, and the
change is confined to the unreleased `master` / `branch-4.x`. For ordinary (dotless) identifiers the
order is unchanged.

### How was this patch tested?

Added a unit test to `SchemaInferenceUtilsSuite` that builds two resolved flows whose identifiers
(`` `c`.`a.b`.`x` `` and `` `c`.`a`.`b.x` ``) would collide under a dot-joined key but stay distinct
when sorted on their parts, each carrying a case-only-differing column, and asserts the inferred
schema is identical regardless of the order the flows are passed in. The test fails with the
previous `unquotedString` key and passes with the parts key.

Ran `build/sbt 'pipelines/testOnly *SchemaInferenceUtilsSuite'` (13 tests, all passed).

### Was this patch authored or co-authored using generative AI tooling?

Generated-by: Claude Opus 4.8

Closes #58223 from anew/sdp-sort-flows-quoted-identifier.

Authored-by: Andreas Neumann <anew@apache.org>
Signed-off-by: Szehon Ho <szehon.apache@gmail.com>
3 files changed
tree: 704abb8a28cba7477b5c2234dcd342a648e829a4
  1. .github/
  2. .mvn/
  3. assembly/
  4. bin/
  5. binder/
  6. build/
  7. common/
  8. conf/
  9. connector/
  10. core/
  11. data/
  12. dev/
  13. docs/
  14. examples/
  15. graphx/
  16. hadoop-cloud/
  17. launcher/
  18. licenses/
  19. licenses-binary/
  20. mllib/
  21. mllib-local/
  22. project/
  23. python/
  24. R/
  25. repl/
  26. resource-managers/
  27. sbin/
  28. sql/
  29. streaming/
  30. tools/
  31. udf/
  32. ui-test/
  33. .asf.yaml
  34. .gitattributes
  35. .gitignore
  36. .nojekyll
  37. .pre-commit-config.yaml
  38. .sbtopts
  39. AGENTS.md
  40. CONTRIBUTING.md
  41. LICENSE
  42. LICENSE-binary
  43. NOTICE
  44. NOTICE-binary
  45. pom.xml
  46. pyproject.toml
  47. README.md
  48. scalastyle-config.xml
  49. SECURITY.md
README.md

Apache Spark

Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R (Deprecated), and an optimized engine that supports general computation graphs for data analysis. It also supports a rich set of higher-level tools including Spark SQL for SQL and DataFrames, pandas API on Spark for pandas workloads, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing.

License Maven Central Java GitHub Actions Build PySpark Coverage PyPI Downloads

Online Documentation

You can find the latest Spark documentation, including a programming guide, on the project web page. This README file only contains basic setup instructions.

Build Pipeline Status

BranchStatus
masterGitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-4.xGitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-4.3GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-4.2GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-4.1GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-4.0GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
GitHub Actions Build
branch-3.5GitHub Actions Build
GitHub Actions Build
GitHub Actions Build

Building Spark

Spark is built using Apache Maven. To build Spark and its example programs, run:

./build/mvn -DskipTests clean package

(You do not need to do this if you downloaded a pre-built package.)

More detailed documentation is available from the project site, at “Building Spark”.

For general development tips, including info on developing Spark using an IDE, see “Useful Developer Tools”.

Interactive Scala Shell

The easiest way to start using Spark is through the Scala shell:

./bin/spark-shell

Try the following command, which should return 1,000,000,000:

scala> spark.range(1000 * 1000 * 1000).count()

Interactive Python Shell

Alternatively, if you prefer Python, you can use the Python shell:

./bin/pyspark

And run the following command, which should also return 1,000,000,000:

>>> spark.range(1000 * 1000 * 1000).count()

Example Programs

Spark also comes with several sample programs in the examples directory. To run one of them, use ./bin/run-example <class> [params]. For example:

./bin/run-example SparkPi

will run the Pi example locally.

You can set the MASTER environment variable when running examples to submit examples to a cluster. This can be spark:// URL, “yarn” to run on YARN, and “local” to run locally with one thread, or “local[N]” to run locally with N threads. You can also use an abbreviated class name if the class is in the examples package. For instance:

MASTER=spark://host:7077 ./bin/run-example SparkPi

Many of the example programs print usage help if no params are given.

Running Tests

Testing first requires building Spark. Once Spark is built, tests can be run using:

./dev/run-tests

Please see the guidance on how to run tests for a module, or individual tests.

There is also a Kubernetes integration test, see resource-managers/kubernetes/integration-tests/README.md

A Note About Hadoop Versions

Spark uses the Hadoop core library to talk to HDFS and other Hadoop-supported storage systems. Because the protocols have changed in different versions of Hadoop, you must build Spark against the same version that your cluster runs.

Please refer to the build documentation at “Specifying the Hadoop Version and Enabling YARN” for detailed guidance on building for a particular distribution of Hadoop, including building for particular Hive and Hive Thriftserver distributions.

Configuration

Please refer to the Configuration Guide in the online documentation for an overview on how to configure Spark.

Contributing

Please review the Contribution to Spark guide for information on how to get started contributing to the project.