[SPARK-58941][SDP] Sort schema inference flows by identifier parts to avoid dotted-name collisions ### What changes were proposed in this pull request? `SchemaInferenceUtils.inferSchemaFromFlows` merges the flows that write to a table in a deterministic order so that, when two flows emit a column whose names differ only in case, the surviving spelling is well-defined and every caller agrees on it. It established that order by sorting on `flow.identifier.unquotedString`. `TableIdentifier.unquotedString` joins the identifier's name parts with an **unescaped** `.`. That mapping is not one-to-one: two structurally distinct identifiers whose parts contain dots (a dot is legal inside a back-tick-quoted schema or flow name) can render to the same string -- e.g. `` `c`.`a.b`.`x` `` and `` `c`.`a`.`b.x` `` both render to `c.a.b.x`. Because `sortBy` is stable, colliding keys fall back to the incoming `flows` order, which is the nondeterministic flow-resolution completion order the sort was meant to remove, so the surviving column casing could flip between runs. This PR sorts on the identifier's parts `(catalog, database, table)` instead. Comparing the parts tuple is injective -- two distinct identifiers are two distinct tuples, so they can never collide. It also preserves the existing order for dotless identifiers (a shorter part sorts first, exactly as the `.` separator, 0x2E, ordered a joined string), unlike `TableIdentifier.quotedString`, which would additionally reorder some ordinary pairs because `` ` `` (0x60) sorts after name characters. Flow identifiers are always fully qualified (`assertIsFullyQualifiedForCreate`), so all three parts are present and the ordering never depends on an absent catalog/database. This is a follow-up to SPARK-58517, which added the sort. ### Why are the changes needed? The sort exists to guarantee a deterministic surviving column spelling. `unquotedString` is a lossy key, so under identifiers that contain dots the guarantee silently breaks: the merge order reverts to the nondeterministic completion order of concurrent flow resolution, and a case-only column spelling can flip between runs of an unchanged pipeline. On the non-merging evolution paths `diffSchemas` keys column identity on the exact name, so a run-to-run flip surfaces as a `deleteColumn` + `addColumn` for a column that only changed case. Sorting on the identifier parts is injective and removes the collision. ### Does this PR introduce _any_ user-facing change? No. The merge order changes only for identifiers with dotted (back-tick-quoted) name parts, and the change is confined to the unreleased `master` / `branch-4.x`. For ordinary (dotless) identifiers the order is unchanged. ### How was this patch tested? Added a unit test to `SchemaInferenceUtilsSuite` that builds two resolved flows whose identifiers (`` `c`.`a.b`.`x` `` and `` `c`.`a`.`b.x` ``) would collide under a dot-joined key but stay distinct when sorted on their parts, each carrying a case-only-differing column, and asserts the inferred schema is identical regardless of the order the flows are passed in. The test fails with the previous `unquotedString` key and passes with the parts key. Ran `build/sbt 'pipelines/testOnly *SchemaInferenceUtilsSuite'` (13 tests, all passed). ### Was this patch authored or co-authored using generative AI tooling? Generated-by: Claude Opus 4.8 Closes #58223 from anew/sdp-sort-flows-quoted-identifier. Authored-by: Andreas Neumann <anew@apache.org> Signed-off-by: Szehon Ho <szehon.apache@gmail.com>
Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Scala, Java, Python, and R (Deprecated), and an optimized engine that supports general computation graphs for data analysis. It also supports a rich set of higher-level tools including Spark SQL for SQL and DataFrames, pandas API on Spark for pandas workloads, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for stream processing.
You can find the latest Spark documentation, including a programming guide, on the project web page. This README file only contains basic setup instructions.
| Branch | Status |
|---|---|
| master | |
| branch-4.x | |
| branch-4.3 | |
| branch-4.2 | |
| branch-4.1 | |
| branch-4.0 | |
| branch-3.5 | |
Spark is built using Apache Maven. To build Spark and its example programs, run:
./build/mvn -DskipTests clean package
(You do not need to do this if you downloaded a pre-built package.)
More detailed documentation is available from the project site, at “Building Spark”.
For general development tips, including info on developing Spark using an IDE, see “Useful Developer Tools”.
The easiest way to start using Spark is through the Scala shell:
./bin/spark-shell
Try the following command, which should return 1,000,000,000:
scala> spark.range(1000 * 1000 * 1000).count()
Alternatively, if you prefer Python, you can use the Python shell:
./bin/pyspark
And run the following command, which should also return 1,000,000,000:
>>> spark.range(1000 * 1000 * 1000).count()
Spark also comes with several sample programs in the examples directory. To run one of them, use ./bin/run-example <class> [params]. For example:
./bin/run-example SparkPi
will run the Pi example locally.
You can set the MASTER environment variable when running examples to submit examples to a cluster. This can be spark:// URL, “yarn” to run on YARN, and “local” to run locally with one thread, or “local[N]” to run locally with N threads. You can also use an abbreviated class name if the class is in the examples package. For instance:
MASTER=spark://host:7077 ./bin/run-example SparkPi
Many of the example programs print usage help if no params are given.
Testing first requires building Spark. Once Spark is built, tests can be run using:
./dev/run-tests
Please see the guidance on how to run tests for a module, or individual tests.
There is also a Kubernetes integration test, see resource-managers/kubernetes/integration-tests/README.md
Spark uses the Hadoop core library to talk to HDFS and other Hadoop-supported storage systems. Because the protocols have changed in different versions of Hadoop, you must build Spark against the same version that your cluster runs.
Please refer to the build documentation at “Specifying the Hadoop Version and Enabling YARN” for detailed guidance on building for a particular distribution of Hadoop, including building for particular Hive and Hive Thriftserver distributions.
Please refer to the Configuration Guide in the online documentation for an overview on how to configure Spark.
Please review the Contribution to Spark guide for information on how to get started contributing to the project.