)]}'
{
  "log": [
    {
      "commit": "8b701473f465e4d7eb05335b0529bad9327db8b3",
      "tree": "47fc984a59812682af8705413b26bf50d197c7c9",
      "parents": [
        "dd0c34c521826f85b8dc072b70bf7b9d53aea1f5"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Sat Jul 25 07:25:48 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Sat Jul 25 07:25:48 2026 +0800"
      },
      "message": "[GLUTEN-12608][CORE] Make ShuffleManagerRouter cache tolerate the executor lifecycle (#12609)"
    },
    {
      "commit": "dd0c34c521826f85b8dc072b70bf7b9d53aea1f5",
      "tree": "b797792a90f09de7158e89abcd6ee0a5cbf2b078",
      "parents": [
        "63bd7c28e2a19ef5e2c524b5dbdd703e02711b6c"
      ],
      "author": {
        "name": "Gluten Performance Bot",
        "email": "137994563+GlutenPerfBot@users.noreply.github.com",
        "time": "Fri Jul 24 18:21:38 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 18:21:38 2026 +0100"
      },
      "message": "[GLUTEN-6887][VL] Daily Update Velox Version (2026_07_24) (#12615)\n\n* [GLUTEN-6887][VL] Daily Update Velox Version (dft-2026_07_24)\n\nUpstream Velox\u0027s New Commits:\nce0173ba1 by Ping Liu, docs: Improve description of building requirement on CPUs\nec5d52c27 by oerling, feat(torchwave): Add ATen ops, reductions, and search (#18187)\n499c96260 by Krishnakanth, feat(exec): Add kRightAnti join type (#18138)\n75a5a5fae by Avinash Raj, fix(cudf): Register Presto functions in TPCH GPU benchmark\n054106586 by Maria Basmanova, fix(text): Honor serde field.delim when writing TEXT files (#18240)\nb986bd7ea by Jimmy Lu, fix: Stop reading ahead past a back-pressured operator to prevent stale LazyVector loads (#18139)\nf57780b87 by Krishna Pai, build(ci): Reject personal container images in workflows\nf7bd0add0 by Krishna Pai, fix(ci): Use official velox-dev image for adapters jobs\n2b58ad1e8 by LingBin, refactor: Refactor iotaData initialization to use IIFE\n0ffc16474 by Chandrashekhar Kumar Singh, fix(exec): Surface requested cancel as OperationCancelled (#18222)\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\n\n* [GLUTEN-12599][VL] Add actions/cache fallback when Stash misses\n\nPer review: Stash has no restore-keys, so a changed ep/build-velox/src\ntree misses its exact key and cold-builds. Keep the Stash restore first\non the x86/arm/enhanced consumers, and fall back to actions/cache/restore\n(with restore-keys, warm from main) when stash-hit is false; the\nactions/cache save is retained so both paths stay populated. Producers,\nnightly and bundle keep the legacy actions/cache method unchanged. Also\nadd a default branch to the gh/jq arch case per review.\n\nCo-Authored-By: Claude Opus 4.8 (1M context) \u003cnoreply@anthropic.com\u003e\n\n---------\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: Smallfu666 \u003cnick20350@gmail.com\u003e\nCo-authored-by: Claude Opus 4.8 (1M context) \u003cnoreply@anthropic.com\u003e"
    },
    {
      "commit": "63bd7c28e2a19ef5e2c524b5dbdd703e02711b6c",
      "tree": "bb312fc9f48c5880443ae50521f8a4e8b6e590f4",
      "parents": [
        "d10e043141efe06a95e61b6d2447996b7f72501f"
      ],
      "author": {
        "name": "Yuan",
        "email": "yuanzhou@apache.org",
        "time": "Fri Jul 24 09:34:47 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 17:34:47 2026 +0100"
      },
      "message": "[VL] Fix password leak in debug-mode config logging (#12606)\n\n* [CORE] Fix password leak in debug-mode config logging\n\nprintConfig() already had redaction logic keyed on spark.redaction.regex,\nbut when that config key was absent (the common case) getRedactionRegex()\nreturned std::nullopt and every config value — including passwords,\ntokens, and secrets — was logged in plain text.\n\nThis patch adds a hard-coded default redaction pattern to guard on this case\n\n\n---------\n\nSigned-off-by: Yuan \u003cyuanzhou@apache.org\u003e"
    },
    {
      "commit": "d10e043141efe06a95e61b6d2447996b7f72501f",
      "tree": "7ecae700a04e89f0214c6dea150a990f0ff3bed2",
      "parents": [
        "246a2f0526cbb0f46dd70f1704bc7eb896447732"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Fri Jul 24 08:41:29 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 08:41:29 2026 +0100"
      },
      "message": "Bump actions/setup-java from 4 to 5 (#12515)\n\nBumps [actions/setup-java](https://github.com/actions/setup-java) from 4 to 5.\n- [Release notes](https://github.com/actions/setup-java/releases)\n- [Commits](https://github.com/actions/setup-java/compare/v4...v5)\n\n---\nupdated-dependencies:\n- dependency-name: actions/setup-java\n  dependency-version: \u00275\u0027\n  dependency-type: direct:production\n  update-type: version-update:semver-major\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "246a2f0526cbb0f46dd70f1704bc7eb896447732",
      "tree": "9517032cb5b7b59b2b33089f5fe6d436a5bcb55a",
      "parents": [
        "86df8560c61ec43994d7edbed7224835fc248e1e"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Fri Jul 24 08:39:59 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 08:39:59 2026 +0100"
      },
      "message": "Bump actions/github-script from 7 to 9 (#12519)\n\nBumps [actions/github-script](https://github.com/actions/github-script) from 7 to 9.\n- [Release notes](https://github.com/actions/github-script/releases)\n- [Commits](https://github.com/actions/github-script/compare/v7...v9)\n\n---\nupdated-dependencies:\n- dependency-name: actions/github-script\n  dependency-version: \u00279\u0027\n  dependency-type: direct:production\n  update-type: version-update:semver-major\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "86df8560c61ec43994d7edbed7224835fc248e1e",
      "tree": "2edb24593e93e1e3647b29f2d1c3644c01e0d777",
      "parents": [
        "a9638bfbf9c964546cbe3dc198b0fe502bea9ea2"
      ],
      "author": {
        "name": "Han-Yin Chang",
        "email": "nick20350@gmail.com",
        "time": "Fri Jul 24 15:11:20 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 08:11:20 2026 +0100"
      },
      "message": "[GLUTEN-12599][VL] Migrate Velox CI ccache restore to Apache Stash (part 2) (#12610)\n\n* [GLUTEN-12599][VL] Part 2: restore Velox CI ccache from Apache Stash\n\nSwitch the ccache restore steps in velox_backend_arm.yml,\nvelox_backend_x86.yml and velox_backend_enhanced.yml from\nactions/cache to the Apache Stash entries seeded by\nvelox_backend_cache.yml (#12602). Stash has no restore-keys\nmechanism, so restores use the same\nhashFiles(\u0027ep/build-velox/src/**\u0027) key the producer saves under;\na changed Velox build tree falls back to a cold ccache build.\nConsumer-side actions/cache saves are kept for instant rollback\nand will be removed in the final cleanup once the Stash restore\npath has been validated in CI.\n\nCo-Authored-By: Claude Opus 4.8 (1M context) \u003cnoreply@anthropic.com\u003e\n\n* [GLUTEN-12599][VL] Install gh and jq for Stash restore in container jobs\n\nThe stash/restore action shells out to gh and jq, which the\nvcpkg-centos-9 and centos-9-jdk8 job containers do not provide\n(host-runner jobs already have them). Install pinned static\nbinaries before the restore step in the three container-based\njobs.\n\nCo-Authored-By: Claude Opus 4.8 (1M context) \u003cnoreply@anthropic.com\u003e\n\n---------\n\nCo-authored-by: Claude Opus 4.8 (1M context) \u003cnoreply@anthropic.com\u003e"
    },
    {
      "commit": "a9638bfbf9c964546cbe3dc198b0fe502bea9ea2",
      "tree": "eb04eecd2a7af2709dd45cf0907674b72c0c4918",
      "parents": [
        "26f04d815c28c4253fe693a02c3b69f4b385ba58"
      ],
      "author": {
        "name": "kevinwilfong",
        "email": "kevinwilfong@fb.com",
        "time": "Thu Jul 23 23:53:32 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 07:53:32 2026 +0100"
      },
      "message": "[VL] Include headers from INSTALL_PREFIX as system includes (#12423)\n\n#12105 introduced the flag -DCMAKE_NO_SYSTEM_FROM_IMPORTED\u003dON to Velox builds. This has the unintended side effect of including dependencies\u0027 header files as non-system includes which means that warnings in the build are promoted to errors, e.g.\n\n  In file included from t.cc:1:\n  In file included from …/deps-install/include/re2/re2.h:220:\n  In file included from …/deps-install/include/absl/base/call_once.h:40:\n  In file included from …/deps-install/include/absl/base/nullability.h:153:\n  In file included from …/deps-install/include/absl/base/internal/nullability_impl.h:22:\n  …/deps-install/include/absl/meta/type_traits.h:511:36: error: builtin __is_trivially_relocatable is deprecated; use __builtin_is_cpp_trivially_relocatable instead\n  [-Werror,-Wdeprecated-builtins]\n    511 |     : std::integral_constant\u003cbool, __is_trivially_relocatable(T)\u003e {};\n        |                                    ^\n  1 error generated.\nTo fix this, this PR proposes explicitly including the header files under ${INSTALL_PREFIX}/include from dependencies as system includes, suppressing these errors which we have limited control over."
    },
    {
      "commit": "26f04d815c28c4253fe693a02c3b69f4b385ba58",
      "tree": "e5bb10bfd41ac9b8d8ba90914e7d6d136cfade71",
      "parents": [
        "302e086d0971fdb0b8c36bdf1a552e7ce8054bad"
      ],
      "author": {
        "name": "wangxinshuo-bolt",
        "email": "wangxinshuo.db@bytedance.com",
        "time": "Fri Jul 24 14:10:04 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 24 14:10:04 2026 +0800"
      },
      "message": "[GLUTEN-12535][CORE] Support backend-specific JNI input adapters (#12535)\n\n* [CORE] Add reusable JNI input adapters for backend runtimes\n\n* Remove useless test\n\nfix\n\n* fix\n\n* [CORE] Decouple JNI input adapters from Runtime\n\n* [CORE] Move shuffle stream JNI bridge to JniCommon"
    },
    {
      "commit": "302e086d0971fdb0b8c36bdf1a552e7ce8054bad",
      "tree": "b954449396d13b0ec2992cc854b7b682f106160d",
      "parents": [
        "995983d4f2bb23a6afc9007c9a65d7177ba8abcb"
      ],
      "author": {
        "name": "Han-Yin Chang",
        "email": "nick20350@gmail.com",
        "time": "Thu Jul 23 22:16:47 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 23 15:16:47 2026 +0100"
      },
      "message": "[GLUTEN-12599][VL] Seed Apache Stash CI caches (#12602)\n\nThis is part 1 of the Apache Stash migration for #12599.\n\nsave a second copy of each cache produced by velox_backend_cache.yml to Apache Stash\nkeep all existing actions/cache restore/save steps unchanged\nleave the ARM, x86, and enhanced Velox consumer workflows unchanged, so other PRs continue using the existing cache path\nuse stable architecture/configuration-specific Stash keys\ninclude hidden files when saving the .ccache directory\npin Apache Stash to an immutable commit SHA\nAfter this merges, the Apache main cache workflow can run and confirm that all Stash artifacts are stored correctly. Part 2 will be submitted separately to enable Stash restores in CI jobs only after that verification."
    },
    {
      "commit": "995983d4f2bb23a6afc9007c9a65d7177ba8abcb",
      "tree": "356969a4af47baa87f74ae4ac16b4cd98c80b7e4",
      "parents": [
        "ed81935c40db39b4d23f7d7f420113eccfbaa243"
      ],
      "author": {
        "name": "Niels Pardon",
        "email": "mail@niels-pardon.de",
        "time": "Thu Jul 23 16:15:54 2026 +0200"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 23 15:15:54 2026 +0100"
      },
      "message": "[GLUTEN-12597][CORE] Remove unused vendored Substrait proto files and dead proto-based type derivation (#12598)\n\nDelete four vendored Substrait proto files that carry no live message types in Gluten and were removed upstream (substrait-io/substrait#952 and #940): capabilities.proto, function.proto, parameterized_types.proto, type_expressions.proto. Their generated headers were only referenced by dangling #include lines in cpp/velox/substrait/SubstraitParser.h (the substrait::Capabilities/FunctionSignature/ParameterizedType/DerivationExpression types are used nowhere in the native code), which are removed here as well.\n\ntype_expressions.proto\u0027s DerivationExpression had one JVM consumer, the org.apache.gluten.substrait.derivation package (DerivationExpressionNode, DerivationExpressionBuilder, BinaryOPNode, DerivationFP64TypeNode). That package is orphaned (no callers in the producer, native backends, or tests), so it is removed together with the proto. Upstream expresses output-type derivation via the ANTLR grammar in the extension YAMLs, not DerivationExpression protos.\n\nAlso remove the deprecated Expression.Enum message and its rex_type oneof field (10), reserving the number and name to match upstream (substrait-io/substrait#1086). Expression.Enum is not built by the producer nor parsed by the Velox/ClickHouse backends; enum function arguments use FunctionArgument.enum instead.\n\nNo functional change. Preparatory cleanup toward rebasing the vendored proto onto Substrait 0.98.0."
    },
    {
      "commit": "ed81935c40db39b4d23f7d7f420113eccfbaa243",
      "tree": "370221c01a8caee3d0c67c20e5405530c5dbeb07",
      "parents": [
        "37f518a095caf94ce99e49d14dbfc3e1c7f55475"
      ],
      "author": {
        "name": "Ismaël Mejía",
        "email": "iemejia@gmail.com",
        "time": "Thu Jul 23 15:17:13 2026 +0200"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 23 21:17:13 2026 +0800"
      },
      "message": "[VL] Enable file handle cache by default with TTL-based eviction (#12400)"
    },
    {
      "commit": "37f518a095caf94ce99e49d14dbfc3e1c7f55475",
      "tree": "02ce3b54009529d5d107a53a003dd3b2a9e17778",
      "parents": [
        "1b545dfac89cd5cd21e6d2ff621bdf957d39a74d"
      ],
      "author": {
        "name": "Zhen Wang",
        "email": "643348094@qq.com",
        "time": "Thu Jul 23 10:15:52 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 23 10:15:52 2026 +0800"
      },
      "message": "[GLUTEN-12594][VL] Fix incorrect task attempt ID used for Uniffle block IDs"
    },
    {
      "commit": "1b545dfac89cd5cd21e6d2ff621bdf957d39a74d",
      "tree": "393b3828614bb4c8613ef791df0220ab4eb32e1d",
      "parents": [
        "9c00fa03162dd3a4871304a030905aa2ff6906f8"
      ],
      "author": {
        "name": "Zhen Wang",
        "email": "643348094@qq.com",
        "time": "Thu Jul 23 09:24:14 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 23 09:24:14 2026 +0800"
      },
      "message": "[GLUTEN-12474][CORE] Preserve V1 write ordering for dynamic partition writes (#12514)\n\n* test GLUTEN-12474\n\n* Preserve V1 write ordering when inserting local sorts\n\n* add unit tests for other spark version\n\n* fix\n\n* fix spotless check\n\n* Enhance local sort handling for GlutenPlan support"
    },
    {
      "commit": "9c00fa03162dd3a4871304a030905aa2ff6906f8",
      "tree": "303bd30ccfa2db7a008f8f680956d54dd8bf868d",
      "parents": [
        "7c3503d94e18c08b907c8c86b0922b0d190a1381"
      ],
      "author": {
        "name": "Gluten Performance Bot",
        "email": "137994563+GlutenPerfBot@users.noreply.github.com",
        "time": "Wed Jul 22 18:32:17 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 18:32:17 2026 +0100"
      },
      "message": "[GLUTEN-6887][VL] Daily Update Velox Version (2026_07_22) (#12596)\n\n* [GLUTEN-6887][VL] Daily Update Velox Version (dft-2026_07_22)\n\nUpstream Velox\u0027s New Commits:\n4487e4409 by Jiaan Geng, fix(spark): Map ORC/DWRF files by position per-file for Hive placeholder names and forced positional evolution\nac6b2d3cd by Mohammad Linjawi, feat(function): Add Spark ANSI interval unary minus, add, and subtract\ne3a8856a8 by Shruti Shivakumar, fix(cudf): Correct left/full join with multiple build batches\n35580128c by Maria Basmanova, fix(dwio): Stabilize exception message templates (#18210)\nd888e41e2 by Maria Basmanova, refactor(dwio): Add context to region bounds check (#18209)\n08c2e67d0 by Devavret Makkar, feat(cuDF): Rebatch before HashJoinProbe to improve compute utilization in large clusters\n2987d093a by RindsSchei225e, feat(fuzzer): Selectively project columns (#17976)\n88b281981 by Jianjian Xie, fix(gcs): Improve GCS error diagnostics\n8da089e08 by Kyle Edwards, build: Bump CMake to 4.3.2\n042c1b633 by Mariam AlMesfer, feat: Add Spark CAST(varchar as timestamp_utc)\n9a52d0ea2 by Richard Barnes, refactor: Migrate folly::format to fmt::format in velox (#18198)\nfb11b7d7a by Krishnakanth, docs: Add MarkSortedNode to the operators reference\n14779a1a9 by Pedro Pedreira, fix(exec): Copy LocalExchangeMemoryManager under task lock to avoid teardown race\nfb9788df6 by Pedro Pedreira, fix(type): Add explicit constructor to TestCustomType for libstdc++ 12 (#18175)\nfbcf1134a by Yabin Ma, fix(dwio): Fix flaky DirectBufferedInputTest.duplicateRegionsShareCoalescedRead\n22be234ba by Maria Basmanova, fix(memory): Record real exception message templates\n06dec49a3 by Matt Gara, feat(cudf): Add support for `DATE_ADD` and `DATE_TRUNC`\n62f49074a by Harthi7, feat: Make Spark CAST(string as timestamp) ANSI-compliant\nab67e6826 by Henry Dikeman, feat(dwio): Add currentStripe() accessor to RowReader (#18146)\n163f2c43f by Maria Basmanova, fix(type): Treat custom types as opaque in coercion (#18164)\n4661eb90e by Krishna Pai, fix: Revert Update FBOS to v2026.07.13.00\nc021c2e4a by Maria Basmanova, fix(core): Handle shared expression DAGs in linear time (#18169)\n5462c2030 by Henry Dikeman, build: Update FBOS to v2026.07.13.00 (#17791)\n39a3c7ed9 by Richard Barnes, Migrate `folly::bit_cast` callers to C++20 `std::bit_cast`\n97915b7e0 by oerling, feat(torchwave): Fused elementwise index_select (#18148)\na77ee8b91 by Sreeni Viswanadha, perf: Optimize MAP_UNION_SUM hot loop (decode-once + single-probe) (#18004)\n111df56a2 by Zhenyuan Zhao, fix: Fix heap-buffer-overflow reads in Parquet reader (#18105)\nc036aaaff by Henry Dikeman, fix(expr): Dedup values in tryMergeBigintRanges (#18134)\n56907faf2 by Henry Dikeman, fix(type): Guard NegatedBigintRange::mergeWith overflow (#18133)\n572423ff9 by Chandrashekhar Kumar Singh, fix(exec): LocalExchangeSource use-after-free at teardown (#18157)\nf82cf06d6 by Maria Basmanova, refactor: Make AggregateCallExpr::dropAlias overridable (#18151)\n57ef1b8c1 by Philo He, fix(sparksql): Enable ANSI-compliant cast from INTEGRAL to DECIMAL\nb7cded725 by Shruti Shivakumar, fix(cudf): Return NULL instead of NaN for avg of empty groups\n0b1741281 by Krishna Pai, fix(ci): Use Velox-owned fuzzer images\nf4042228a by Karthikeyan, fix(dwio): Fix dictionary filter cache-mask extraction on ARM NEON\n43a1fb41b by zhichenxu-meta, feat(rpc): Add kInvalidRequest error kind for non-retryable failures (#18095)\n94d87c285 by Bradley Dice, build: Upgrade bundled gflags to 2.3.0\n77d1607a0 by Krishnakanth, fix(core): Preserve markerName in UnnestNode::Builder\nadfc53cae by Krishnakanth, fix(core): Validate UnnestNode unnestNames count\ne567359ec by Arpit Porwal, refactor: Allow registering collect_list under arbitrary names (#18132)\nad186be92 by BRIJ RAJ KISHORE, feat(sparksql): Make unaryminus ANSI-compliant for integer overflow detection (#18096)\n56a570229 by Sreeni Viswanadha, fix: Report the actual JSON element type in JSON cast error messages (#18017)\n3637c0cf9 by James Lamb, build(ci): Update to zizmor 1.26.1\n6609329d8 by Deepak Majeti, fix(stats): Add more runtime metrics to aggregate per operator (#18022)\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\n\n* fix gpu build\n\nvelox-gpu requres cmake 4.0\n\nSigned-off-by: Yuan \u003cyuanzhou@apache.org\u003e\n\n---------\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nSigned-off-by: Yuan \u003cyuanzhou@apache.org\u003e\nCo-authored-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: Yuan \u003cyuanzhou@apache.org\u003e"
    },
    {
      "commit": "7c3503d94e18c08b907c8c86b0922b0d190a1381",
      "tree": "85b582cacce3642df5a55f19abaa658e51c6969a",
      "parents": [
        "d76a4f81851e07028d485a78a105528571b73d4c"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Wed Jul 22 16:05:12 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 16:05:12 2026 +0100"
      },
      "message": "Bump org.apache.parquet:parquet-hadoop from 1.15.2 to 1.17.1 (#12579)\n\nBumps [org.apache.parquet:parquet-hadoop](https://github.com/apache/parquet-mr) from 1.15.2 to 1.17.1.\n- [Release notes](https://github.com/apache/parquet-mr/releases)\n- [Changelog](https://github.com/apache/parquet-java/blob/master/CHANGES.md)\n- [Commits](https://github.com/apache/parquet-mr/compare/apache-parquet-1.15.2...apache-parquet-1.17.1)\n\n---\nupdated-dependencies:\n- dependency-name: org.apache.parquet:parquet-hadoop\n  dependency-version: 1.17.1\n  dependency-type: direct:production\n  update-type: version-update:semver-minor\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "d76a4f81851e07028d485a78a105528571b73d4c",
      "tree": "c7d78e2f58528e2991bacc74ada948629ae3575e",
      "parents": [
        "2511881d8aff11040e35c2562dbc53a2fd71aaf2"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Wed Jul 22 16:04:29 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 16:04:29 2026 +0100"
      },
      "message": "Bump org.xerial:sqlite-jdbc from 3.50.3.0 to 3.53.2.0 (#12577)\n\nBumps [org.xerial:sqlite-jdbc](https://github.com/xerial/sqlite-jdbc) from 3.50.3.0 to 3.53.2.0.\n- [Release notes](https://github.com/xerial/sqlite-jdbc/releases)\n- [Changelog](https://github.com/xerial/sqlite-jdbc/blob/master/CHANGELOG)\n- [Commits](https://github.com/xerial/sqlite-jdbc/compare/3.50.3.0...3.53.2.0)\n\n---\nupdated-dependencies:\n- dependency-name: org.xerial:sqlite-jdbc\n  dependency-version: 3.53.2.0\n  dependency-type: direct:development\n  update-type: version-update:semver-minor\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "2511881d8aff11040e35c2562dbc53a2fd71aaf2",
      "tree": "e0ae8bcec9ff8f9f050f3e72e4f6fcdcf6bb6e11",
      "parents": [
        "1562bc27e9b1141d1ef633d3dd4f447090443056"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Wed Jul 22 16:03:52 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 16:03:52 2026 +0100"
      },
      "message": "Bump org.apache.maven.plugins:maven-surefire-plugin from 3.3.0 to 3.5.6 (#12578)\n\nBumps [org.apache.maven.plugins:maven-surefire-plugin](https://github.com/apache/maven-surefire) from 3.3.0 to 3.5.6.\n- [Release notes](https://github.com/apache/maven-surefire/releases)\n- [Commits](https://github.com/apache/maven-surefire/compare/surefire-3.3.0...surefire-3.5.6)\n\n---\nupdated-dependencies:\n- dependency-name: org.apache.maven.plugins:maven-surefire-plugin\n  dependency-version: 3.5.6\n  dependency-type: direct:production\n  update-type: version-update:semver-minor\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "1562bc27e9b1141d1ef633d3dd4f447090443056",
      "tree": "71249b1a9d28760f691835b501ebc11aedbc22a4",
      "parents": [
        "5906016217114ab21f95aa2381c4162c8c0a79dd"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Wed Jul 22 16:03:21 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 16:03:21 2026 +0100"
      },
      "message": "Bump org.apache.maven.plugins:maven-compiler-plugin from 3.8.1 to 3.15.0 (#12576)\n\nBumps [org.apache.maven.plugins:maven-compiler-plugin](https://github.com/apache/maven-compiler-plugin) from 3.8.1 to 3.15.0.\n- [Release notes](https://github.com/apache/maven-compiler-plugin/releases)\n- [Commits](https://github.com/apache/maven-compiler-plugin/compare/maven-compiler-plugin-3.8.1...maven-compiler-plugin-3.15.0)\n\n---\nupdated-dependencies:\n- dependency-name: org.apache.maven.plugins:maven-compiler-plugin\n  dependency-version: 3.15.0\n  dependency-type: direct:production\n  update-type: version-update:semver-minor\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "5906016217114ab21f95aa2381c4162c8c0a79dd",
      "tree": "c9c7b464e44c2f60f7cccbdd76afb23fbfdce9d1",
      "parents": [
        "717d3e47f10988a49c2da79ed41c943bef08fe21"
      ],
      "author": {
        "name": "dependabot[bot]",
        "email": "49699333+dependabot[bot]@users.noreply.github.com",
        "time": "Wed Jul 22 15:59:43 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 15:59:43 2026 +0100"
      },
      "message": "Bump pillow in /tools/workload/benchmark_velox/analysis (#12591)\n\nBumps [pillow](https://github.com/python-pillow/Pillow) from 12.2.0 to 12.3.0.\n- [Release notes](https://github.com/python-pillow/Pillow/releases)\n- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)\n- [Commits](https://github.com/python-pillow/Pillow/compare/12.2.0...12.3.0)\n\n---\nupdated-dependencies:\n- dependency-name: pillow\n  dependency-version: 12.3.0\n  dependency-type: direct:production\n...\n\nSigned-off-by: dependabot[bot] \u003csupport@github.com\u003e\nCo-authored-by: dependabot[bot] \u003c49699333+dependabot[bot]@users.noreply.github.com\u003e"
    },
    {
      "commit": "717d3e47f10988a49c2da79ed41c943bef08fe21",
      "tree": "82dc0b030e1b5d1e90b6fa095c464a91ea0f5f0e",
      "parents": [
        "3df21cad0c2713214e72f6a0626c9790ed61c1a9"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Wed Jul 22 15:02:07 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 15:02:07 2026 +0100"
      },
      "message": "[VL] Remove gpu hash shuffle writer type (#12593)"
    },
    {
      "commit": "3df21cad0c2713214e72f6a0626c9790ed61c1a9",
      "tree": "471b0cacabbe4a759c937551c764a610952a26fd",
      "parents": [
        "36beabd4118501e686f1c5ae700ce6b18a0e1b3b"
      ],
      "author": {
        "name": "Hongze Zhang",
        "email": "hongze.zzz123@gmail.com",
        "time": "Wed Jul 22 13:19:25 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 22 13:19:25 2026 +0100"
      },
      "message": "[CORE][VL] Add configuration for maximum input partitions in V2 batch scans (#12589)"
    },
    {
      "commit": "36beabd4118501e686f1c5ae700ce6b18a0e1b3b",
      "tree": "76b0ba0e29da6b51f36893128c0c3f3fe31a87e1",
      "parents": [
        "ce7aa6307239098418e5be235a2e8343a29e0bb6"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Tue Jul 21 21:14:50 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 21 14:14:50 2026 +0100"
      },
      "message": "[GLUTEN-12462][CORE] Expose shuffle reader metrics iterator delegate (#12587)"
    },
    {
      "commit": "ce7aa6307239098418e5be235a2e8343a29e0bb6",
      "tree": "e6fce30f479cd9f2d0e008ff1632a7168b686325",
      "parents": [
        "352e05a070a5ea849e1fb912902d4309bf9e5b0e"
      ],
      "author": {
        "name": "JiaKe",
        "email": "ke.jia@ibm.com",
        "time": "Tue Jul 21 08:56:31 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 21 08:56:31 2026 +0100"
      },
      "message": "[VL]Add hash table memory usage metric (#12528)"
    },
    {
      "commit": "352e05a070a5ea849e1fb912902d4309bf9e5b0e",
      "tree": "a52fba6d46cb3130f65b5b6fbc24564d67659d76",
      "parents": [
        "25845255fb1e24a4bc037c054d9bb7fe58e61c52"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Tue Jul 21 15:51:09 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 21 15:51:09 2026 +0800"
      },
      "message": "[GLUTEN-11550][VL] Disable GlutenExplainSuite in Spark 4.x (#12546)"
    },
    {
      "commit": "25845255fb1e24a4bc037c054d9bb7fe58e61c52",
      "tree": "f9a424c1fa06cd61f554c9dd9d790a2d05fe7602",
      "parents": [
        "bff929c0bc672a6b5de9049d88ccfda4c7c64f42"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Tue Jul 21 13:25:48 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 21 13:25:48 2026 +0800"
      },
      "message": "[MINOR] Sync LICENSE with actual on-disk paths (#12584)"
    },
    {
      "commit": "bff929c0bc672a6b5de9049d88ccfda4c7c64f42",
      "tree": "c6ca3c909e4eb974704b85e5ce997646575e8e11",
      "parents": [
        "ac19a60da66ded142537b056ecc776d13670f5ec"
      ],
      "author": {
        "name": "Philo He",
        "email": "philohe@amazon.com",
        "time": "Tue Jul 21 10:57:39 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 19:57:39 2026 -0700"
      },
      "message": "[VL] Fix build issues reported by weekly jobs (#12321)\n\nCo-authored-by: Yuan \u003cyuanzhou@apache.org\u003e"
    },
    {
      "commit": "ac19a60da66ded142537b056ecc776d13670f5ec",
      "tree": "4c8a49809a752de7ca1d8144dcaf2b22602b0102",
      "parents": [
        "2b49fc040ea9e6410c57f70fc5475c0dbdd578f1"
      ],
      "author": {
        "name": "Yuan",
        "email": "yuanzhou@apache.org",
        "time": "Mon Jul 20 12:40:51 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 12:40:51 2026 -0700"
      },
      "message": "[VL][DOC] Update Velox backend limitations (#12575)"
    },
    {
      "commit": "2b49fc040ea9e6410c57f70fc5475c0dbdd578f1",
      "tree": "21e2c988bd415cc0beebcebe23e54f4fe65d2e34",
      "parents": [
        "a88d3c1e47caee04e1e862a5eb1a26c4b294ae9e"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Mon Jul 20 15:14:18 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 15:14:18 2026 +0100"
      },
      "message": "[VL] Support setting stage execution mode for hybrid execution (#12511)"
    },
    {
      "commit": "a88d3c1e47caee04e1e862a5eb1a26c4b294ae9e",
      "tree": "c199ab19dbeeadd1c346369d75a2122920ead41a",
      "parents": [
        "f734ed79f36b46fe05a5f21e2f010a80fbb1aa59"
      ],
      "author": {
        "name": "Yuming Wang",
        "email": "yumwang@ebay.com",
        "time": "Mon Jul 20 21:55:31 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 21:55:31 2026 +0800"
      },
      "message": "[GLUTEN-12555][CORE] Propagate child outputPartitioning and outputOrdering in ColumnarToColumnarExec (#12556)"
    },
    {
      "commit": "f734ed79f36b46fe05a5f21e2f010a80fbb1aa59",
      "tree": "7db414b9fe6912a8821085de867e79a567c3ec38",
      "parents": [
        "f2cda79d9161bcc7e90a1b619e31740a906e8214"
      ],
      "author": {
        "name": "inf",
        "email": "ialhazmim@gmail.com",
        "time": "Mon Jul 20 08:56:37 2026 +0000"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 09:56:37 2026 +0100"
      },
      "message": "[VL] test(iceberg): Fix disabled iceberg row group bytes test (#12509)\n\nPreviously, the row group bytes value was a hard boundary of the uncompressed value. Now the implementation causes the buffer to reach or slightly exceed the threshold. The tests are updated to guard bigger written row group size than configuration."
    },
    {
      "commit": "f2cda79d9161bcc7e90a1b619e31740a906e8214",
      "tree": "34101d62e06c9403b65858ae851a8e9beaec1f96",
      "parents": [
        "6fd6df38a8129cb29084a3c0e76521f4fe9b0beb"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Mon Jul 20 15:39:34 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 15:39:34 2026 +0800"
      },
      "message": "[MINOR] Clean up Spark 3.2 build/config/docs leftovers (#12550)"
    },
    {
      "commit": "6fd6df38a8129cb29084a3c0e76521f4fe9b0beb",
      "tree": "a22dcb67048b2753fa1eb3d3614c2f2cbe3e5db7",
      "parents": [
        "a9c410da2c33b9c3f2b73a6bb69325ee89ebde31"
      ],
      "author": {
        "name": "jackylee",
        "email": "qcsd2011@gmail.com",
        "time": "Mon Jul 20 11:51:26 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 20 11:51:26 2026 +0800"
      },
      "message": "[VL][CI] Quick Fix Spark 4.1 CI crash: isolate libgluten\u0027s bundled libstdc++ (#12557)"
    },
    {
      "commit": "a9c410da2c33b9c3f2b73a6bb69325ee89ebde31",
      "tree": "3a228b586ae07e22cc103f05c7f8deb7bdbae5c1",
      "parents": [
        "c2ff19bed9de43ab44e296abe0d0843017029869"
      ],
      "author": {
        "name": "Gluten Performance Bot",
        "email": "137994563+GlutenPerfBot@users.noreply.github.com",
        "time": "Sun Jul 19 06:56:55 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Sat Jul 18 22:56:55 2026 -0700"
      },
      "message": "[GLUTEN-6887][VL] Daily Update Velox Version (2026_07_16) (#12527)\n\n* [GLUTEN-6887][VL] Daily Update Velox Version (dft-2026_07_16)\n\nUpstream Velox\u0027s New Commits:\n8bece69ac by Linsong Wang, Merge branch \u0027ci-fix-pr\u0027 into staging-rebase-pr\n669c9f2e4 by wforget, [OAP] Change SpillPartitionId::kMaxSpillLevel to 7\na6e61e605 by Ke Jia, [OAP]feat: Implement se/dser method for HashTable\nf0b511b59 by Yuan, [OAP] [11771] Fix smj result mismatch issue in semi, anit and full outer join\n7ab4ec425 by Rui Mo, feat: Allow subfield rename and deletion for Parquet format\n53a990726 by Zhenyuan Zhao, fix(dwrf): Bound PATCHED_BASE gap chain to avoid OOB read (#18106)\ne1dd48045 by Hongze Zhang, refactor: Split `dwio/common/Statistics.h` to `.h` and `.cpp`\n5219cafbd by oerling, feat(torchwave): Free intermediate tensors at last use (#18119)\nd8ceba11c by Wechar Yu, fix(spark): Support VARBINARY type in to_json function\n1861f170f by Krishna Pai, fix: Revert folly double/float formatting goldens (#18116)\nea4ae633b by Apurva Kumar, feat: Replace (union) an existing deletion vector on repeated mutation (#18117)\na3c8b8657 by Zhenyuan Zhao, fix: Fix test failure (#18103)\n949c7309b by Deepak Majeti, feat(parquet): Add support for Timestamp type rowgroup filtering\n0b01ba0ad by Deepak Majeti, fix(build): FBThrift warnings for parquet.thrift\n1e7952ade by Apurva Kumar, test: NIMBLE Iceberg field-id end-to-end write+read test (#18056)\n332b91f61 by PRASHANT GOLASH, perf: Use DirectBufferedInput for Nimble (enable loadQuantum) (#17948)\ndb15b91d3 by Liangcai Li, feat(cudf): Add Spark substring support\nfe6aae45d by Reema, fix(cudf): Fix multi column `hash_with_seed` on GPU\n049fe62d6 by Philo He, fix: Enable ANSI-compliant cast from VARCHAR to DECIMAL\nc59294974 by Amit Dutta, fix(window): Fail gracefully on oversized window partitions (#18102)\nefb391ea4 by Richard Barnes, Update test goldens for folly double/float string formatting change\nd12dbd60e by jkhaliqi, feat: Add CompressionKind_LZ4_HADOOP\n7f6858340 by Daniel Bauer, fix(cudf): Support ms/s timestamp scalars on GPU\n4991078e0 by Ping Liu, docs: Scope BMI/BMI2/F16C requirements to x86\n4920e8bf3 by Ismaël Mejía, perf(dwio): Cross-platform bit-unpacking and level conversion\n4cf8f85db by Ismaël Mejía, fix(hive): Wire file-handle-expiration-duration-ms to SimpleLRUCache (#18072)\nc1287a247 by meta-codesync[bot], misc: Add Thrift AllowLegacyMissingUris annotation to Velox files (#18012)\nc759d48ad by Raymond Lin, fix(dwio): Correct flat-map prefetch eligibility (#18038)\n039614eda by Ismaël Mejía, fix(dwio): Advance pointers in unpackNaive after decoding loop\na8c470526 by Durvesh Pilankar, docs: Fix duplicated words and typos\naffccb3dc by Kyle Hubert, fix(ucx): Preserve shared exchange client on operator close (#17752)\nd1f09a430 by binwei yang, upgrade to gcc-133 + spark-3.5 build\n6841cd103 by binwei yang, remove spark3.4\n7531da59e by binwei yang, add cache key\ne962632aa by Yuan, fix: Enable CI on IBM repo\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\n\n* Add support for existing deletion vectors in IcebergWriter\n\n* Format parameters for IcebergInsertTableHandle constructor\n\n* Update IcebergWriter.cc\n\n* Fix type in IcebergWriter for ExistingDeletionVector\n\n* remove unused profile\n\n* Update VeloxWriterUtils.cc\n\n---------\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: BInwei Yang \u003cfelixybw@apache.org\u003e\nCo-authored-by: Rong Ma \u003cmarong@apache.org\u003e"
    },
    {
      "commit": "c2ff19bed9de43ab44e296abe0d0843017029869",
      "tree": "2ab017707454a1404bf049372aba85bd0c0cea70",
      "parents": [
        "23ed0c15445c56aba91f90512573b061d415f90b"
      ],
      "author": {
        "name": "BInwei Yang",
        "email": "felixybw@apache.org",
        "time": "Fri Jul 17 09:29:46 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 17 09:29:46 2026 -0700"
      },
      "message": "[VL] Set ccache size to 1G (#12540)\n\nThis PR standardizes the ccache maximum size to 1G across additional Velox CI workflows that were still using ccache’s default size, reducing disk usage during builds while keeping cache behavior consistent with other existing Velox workflows."
    },
    {
      "commit": "23ed0c15445c56aba91f90512573b061d415f90b",
      "tree": "3b8cd0d5f76494ef74805e48c0200cf6dc501165",
      "parents": [
        "4286da6e90c94d9309743a09bf752f77603b5548"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Fri Jul 17 17:39:26 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 17 17:39:26 2026 +0800"
      },
      "message": "[MINOR][CH] Drop Spark 3.2 / 3.1-era compat workarounds in ClickHouse backend (#12541)"
    },
    {
      "commit": "4286da6e90c94d9309743a09bf752f77603b5548",
      "tree": "2211ceac5782ae2fe0a212f6e00c16abf32ef693",
      "parents": [
        "c823f816a32c51130a7e70793c2d459a4854cb41"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Fri Jul 17 14:00:56 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 17 14:00:56 2026 +0800"
      },
      "message": "[VL] Fix GlutenSetCommandSuite in Spark 4.x (#12533)"
    },
    {
      "commit": "c823f816a32c51130a7e70793c2d459a4854cb41",
      "tree": "84a703bb29be7f1fcae1de8eb236e4622a696e96",
      "parents": [
        "82261019346f2a860778d12d2b320bb0dacc63f3"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Fri Jul 17 10:16:30 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 17 10:16:30 2026 +0800"
      },
      "message": "[MINOR] Align supported Spark version references (#12525)"
    },
    {
      "commit": "82261019346f2a860778d12d2b320bb0dacc63f3",
      "tree": "c1f440cd458224c6a89157d8d6bed220e9e67281",
      "parents": [
        "0dd1d29b4a6ed214ab80fcd5c77dbebc6cc8c4f5"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Thu Jul 16 19:39:42 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 19:39:42 2026 +0800"
      },
      "message": "[MINOR][CH] Remove residual Spark 3.2 branches from clickhouse tests (#12532)\n\n[MINOR][CH] Remove residual Spark 3.2 branches from clickhouse tests (#12532)\n\nSpark 3.2 was dropped by #11351 / #11687 / #11731 / #11887; the\n`spark32` protected def in `GlutenClickHouseWholeStageTransformerSuite`\nwas defined as `sparkVersion.equals(\"3.2\")` and has been dead since.\nThis PR removes the definition and prunes every reachable\n`if (spark32) ...` / `if (!spark32) ...` / `${if (spark32) ... else ...}`\nbranch to keep only the Spark 3.3+ path.\n\nTest-code changes (all in `backends-clickhouse/src/test/`):\n\n- `GlutenClickHouseWholeStageTransformerSuite.scala`: drop the\n  `protected def spark32` definition (always false).\n- `GlutenClickHouseTPCHBucketSuite.scala`: `hasSortByCol \u003d !spark32`\n  collapses to `true` on all supported Sparks; version-gated\n  `if (spark32) ...` branches removed.\n- `GlutenClickHouseTPCHParquetBucketSuite.scala`,\n  `GlutenClickHouseDeltaParquetWriteSuite.scala`,\n  `GlutenClickHouseMergeTreeWriteSuite.scala`,\n  `GlutenClickHouseMergeTreeOptimizeSuite.scala`,\n  `GlutenClickHouseMergeTreeWriteOnHDFSSuite.scala`,\n  `GlutenClickHouseMergeTreeWriteOnHDFSWithRocksDBMetaSuite.scala`,\n  `GlutenClickHouseMergeTreeWriteOnS3Suite.scala`,\n  `GlutenClickHouseMergeTreePathBasedWriteSuite.scala`: every reachable\n  `if (spark32) ... else ...` block, `if (!spark32) ...` guard, and\n  inline `${if (spark32) \"\" else \"SORTED BY (...)\"}` interpolation is\n  reduced to the Spark 3.3+ path (always emit `SORTED BY`).\n- `GlutenClickHouseTPCDSParquetAQESuite.scala`,\n  `GlutenClickHouseTPCDSParquetColumnarShuffleAQESuite.scala`:\n  comments narrowed from \"On Spark 3.2, ... on Spark 3.3, ...\" to\n  describe only the surviving Spark 3.3+ shape.\n- `hive/GlutenClickHouseNativeWriteTableSuite.scala`: drop stale\n  `// spark 3.2 without orc or parquet suffix` comment.\n\nMain-code change (one file):\n\n- `RowToCHNativeColumnarExec.scala`: drop the `// For spark 3.2.`\n  comment above `withNewChildInternal`. The override is required by\n  `TreeNode`\u0027s API on every Spark version currently supported by\n  Gluten, not a Spark 3.2-only quirk. Mirrors the same cleanup for\n  `RowToVeloxColumnarExec` included in #12525.\n\nExplicitly kept for a separate follow-up PR:\n\n- `backends-clickhouse/.../ExtendedColumnPruning.scala:66-72` — the\n  local `getAttributeToExtractValues` re-implementation exists because\n  Spark 3.2\u0027s upstream signature was 2-arg. On 3.3+ it is 3-arg; the\n  local copy could be replaced with a delegate. That is a real\n  refactor, not a comment fix.\n- `backends-clickhouse/.../CHColumnarWrite.scala:157` — the\n  `bucketSpec` reflection was needed for Spark 3.2, may be replaceable\n  with direct access on 3.3+. Also a real refactor.\n- `CustomSum.scala:28` — historical provenance of a copied file, not\n  a version gate; keep as-is.\n\nVerified with `mvn -pl backends-clickhouse -am install -Pspark-3.3,\nbackends-clickhouse,delta` (SUCCESS) and `scalastyle:check spotless:check`\n(SUCCESS)."
    },
    {
      "commit": "0dd1d29b4a6ed214ab80fcd5c77dbebc6cc8c4f5",
      "tree": "15a4ee2d9274a0f825d09210fefbb63f8ab088f2",
      "parents": [
        "580835de86cc18058bba94a0a709c00cd91d28ba"
      ],
      "author": {
        "name": "wangxinshuo-bolt",
        "email": "wangxinshuo.db@bytedance.com",
        "time": "Thu Jul 16 17:39:08 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 17:39:08 2026 +0800"
      },
      "message": "[MINOR][CORE] Remove stale libhdfs install rule (#12531)"
    },
    {
      "commit": "580835de86cc18058bba94a0a709c00cd91d28ba",
      "tree": "570deac754804851615ce94ad3785dccbf343e36",
      "parents": [
        "8a60e5901ba70842a5ddfc3fd5a70f0a1a80d4c2"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 16 16:25:50 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 16:25:50 2026 +0800"
      },
      "message": "[VL] Fix macOS libvelox linking with Velox dbgen (#12530)"
    },
    {
      "commit": "8a60e5901ba70842a5ddfc3fd5a70f0a1a80d4c2",
      "tree": "19f3eefbfd8187fa40f49c1f19cc522f60ffb4c6",
      "parents": [
        "522b45d62b27a0df5b880e82d0b23427bbe9559f"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 16 15:18:37 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 15:18:37 2026 +0800"
      },
      "message": "[GLUTEN-11550][VL] Fix GlutenRemoveRedundantProjectsSuite in Spark 4.x (#12506)"
    },
    {
      "commit": "522b45d62b27a0df5b880e82d0b23427bbe9559f",
      "tree": "3a22a41d5af4bb77ff9113e2a4111d3d95cd06d1",
      "parents": [
        "53893cd75ae24260a3e8386ec9795a2817aef854"
      ],
      "author": {
        "name": "GGboom",
        "email": "guojinhong2@huawei.com",
        "time": "Thu Jul 16 14:04:03 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 14:04:03 2026 +0800"
      },
      "message": "[GLUTEN-12426][FLINK] Feat: Add columnar StreamRecordTimestampInserter operator (#12428)\n\n* feat(flink): add columnar StreamRecordTimestampInserter operator"
    },
    {
      "commit": "53893cd75ae24260a3e8386ec9795a2817aef854",
      "tree": "99dc9ee875a813e2c1d1a892a94eec8756df4ea4",
      "parents": [
        "b939e826b2d751552ef657197b92fb769986b4ca"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 16 11:59:35 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 11:59:35 2026 +0800"
      },
      "message": "[CORE] Add explicit tracking for disabled test suites (#12512)"
    },
    {
      "commit": "b939e826b2d751552ef657197b92fb769986b4ca",
      "tree": "885537876de9e002f56d5f5260c0f6be7b871904",
      "parents": [
        "49ef3fd9d694a78de0fcec122ba85f22a50464bd"
      ],
      "author": {
        "name": "lgbo",
        "email": "lgbo.ustc@gmail.com",
        "time": "Thu Jul 16 11:16:14 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 11:16:14 2026 +0800"
      },
      "message": "[GLUTEN-12468][FLINK] Handle WatermarkStatus elements in GlutenSourceFunction (#12461)\n\n* fix: Handle WatermarkStatus elements in GlutenSourceFunction\n\nAdd processing for WatermarkStatus elements from native idle detection.\nWhen IDLE is received, call sourceContext.markAsTemporarilyIdle() to notify\nFlink that this source is temporarily idle, allowing watermark progress to\ncontinue from other sources.\n\n* test: Add unit tests for GlutenSourceFunction WatermarkStatus handling\n\nAdds GlutenSourceFunctionWatermarkStatusTest with 5 test cases covering:\n- IDLE status triggers markAsTemporarilyIdle()\n- ACTIVE status is a no-op on SourceContext\n- IDLE→ACTIVE transition does not call markAsTemporarilyIdle()\n- Repeated IDLE calls are idempotent\n- ACTIVE status produces no invocations\n\nUses reflection to invoke private processWatermarkStatus() and a custom\nTrackingSourceContext spy — no Mockito or native session required.\n\n* test: Add integration test for idle watermark status in WatermarkAssigner\n\n* fix: Update velox4j reference to feature/idle-source-handling branch\n\n* fix: Add shouldCallNoMoreSplits option to GlutenSourceFunction for unbounded test scenarios\n\n* feat: E2E test for idle WatermarkStatus detection with Kafka + MiniCluster\n\n- Add GlutenStreamSource.isShouldCallNoMoreSplits() delegating to source\n- Add GlutenSourceFunction.isShouldCallNoMoreSplits() getter\n- Extend OffloadedJobGraphGenerator to preserve shouldCallNoMoreSplits\n  when creating a new GlutenSourceFunction during offloading\n- Rewrite GlutenSourceFunctionWatermarkStatusE2ETest as a real E2E test\n  using embedded Kafka broker + Flink MiniCluster, verifying that\n  WatermarkStatus.IDLE is emitted after idle timeout\n- Fix EmptyNode output type in WatermarkPushDownSpec project to match\n  table scan schema (avoids FieldNotFound error during plan init)\n\n* chore: remove obsolete GlutenSourceFunctionWatermarkStatusTest\n\nCovered by GlutenSourceFunctionWatermarkStatusE2ETest which tests\nthe same behavior end-to-end with Kafka + MiniCluster.\n\n* test: verify idle inputs are excluded from combined min-watermark\n\nAdd testIdleInputExcludedFromMinWatermark to\nGlutenStreamTwoInputWatermarkStatusTest: when one input is marked\nIDLE, its watermark is excluded from min-watermark calculation so\nthe other active input can advance freely.\n\n* fix: upgrade surefire in gluten-flink-ut from 3.0.0-M5 to 3.3.0\n\n* fix: revert gluten-flink surefire upgrade\n\n* fix: clean up idle source E2E test\n\n* Update velox4j reference for idle source handling\n\n* fix: Remove noMoreSplits test toggle\n\n* Update velox4j reference for watermark destructor fix\n\n* Update velox4j reference for idle timer tests\n\n* Update velox4j reference for idle source handling\n\n* Update velox4j reference for callback bridge fix\n\n* Update velox4j reference to gluten branch"
    },
    {
      "commit": "49ef3fd9d694a78de0fcec122ba85f22a50464bd",
      "tree": "86d1a9258b496914b1ca9da0e4bd19745f6d8fe4",
      "parents": [
        "0ff76818b6ae7bbc61d6e623a44e8084099c7dbc"
      ],
      "author": {
        "name": "PerumalsamyR",
        "email": "perumalsamy.ravindran@gmail.com",
        "time": "Wed Jul 15 19:49:04 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 10:49:04 2026 +0800"
      },
      "message": "[GLUTEN-12394][CORE] Consolidate JniLibLoader.load() as a thin delegate of loadAndCreateLink() (#12520)"
    },
    {
      "commit": "0ff76818b6ae7bbc61d6e623a44e8084099c7dbc",
      "tree": "160f246c7caa47d6a4040f813858e394e50cdfb6",
      "parents": [
        "7c7b5b18bcf0e33dfed1e66d76636adfaa9439e7"
      ],
      "author": {
        "name": "jianzhenwu",
        "email": "117174379+jianzhenwu@users.noreply.github.com",
        "time": "Thu Jul 16 10:40:18 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 03:40:18 2026 +0100"
      },
      "message": "[GLUTEN-12008][VL] Align Expand projection types with output (#12009)"
    },
    {
      "commit": "7c7b5b18bcf0e33dfed1e66d76636adfaa9439e7",
      "tree": "eaebd080015be72537fcc20d8460189f55a4161c",
      "parents": [
        "b1e00723b0f131a16e271e31673cd0a371d8a32a"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Thu Jul 16 10:25:23 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 10:25:23 2026 +0800"
      },
      "message": "[MINOR][CH] Remove orphan Spark 3.2 / Delta 2.0 source directories (#12524)\n\n[MINOR][CH] Remove orphan Spark 3.2 / Delta 2.0 source directories (#12524)\n\nThe `\u003cdelta.binary.version\u003e20\u003c/delta.binary.version\u003e` property lived only\nin the `spark-3.2` Maven profile, which was removed by #11351. Since then,\n`src-delta20/` under both `backends-clickhouse/` and `gluten-delta/` has\nbeen unreachable dead code:\n\n- No profile currently sets `delta.binary.version\u003d20`\n  (spark-3.3 -\u003e 23, spark-3.4 -\u003e 24, spark-3.5 -\u003e 33,\n   spark-4.0/4.1 -\u003e 40).\n- No CI job, doc, or script activates the directories.\n- Every deleted class has a live replacement in `src-delta23/` and\n  `src-delta33/` (and, for `gluten-delta`, `src-delta24/` /\n  `src-delta40/` as well), so no user-visible symbol is lost.\n\nVerified with a clean rebuild on spark-3.3 + backends-clickhouse + delta."
    },
    {
      "commit": "b1e00723b0f131a16e271e31673cd0a371d8a32a",
      "tree": "d49d6e4841ebd88e662e2c8ccb3057608bc19b6d",
      "parents": [
        "204b8d71288a7f709e22331d28f704fc4cb78593"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Thu Jul 16 09:23:58 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 16 09:23:58 2026 +0800"
      },
      "message": "[MINOR][CORE] Remove residual Spark 3.2 compatibility code in gluten-core (#12522)"
    },
    {
      "commit": "204b8d71288a7f709e22331d28f704fc4cb78593",
      "tree": "444e706d93ecc75a6bf57a08f76b28bc5b4722c8",
      "parents": [
        "d707d66a2d52369b2603e6299f76d0dec99833c9"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Wed Jul 15 14:16:00 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 15 14:16:00 2026 +0100"
      },
      "message": "[VL] Unify gpu hash shuffle writer output (#12499)"
    },
    {
      "commit": "d707d66a2d52369b2603e6299f76d0dec99833c9",
      "tree": "c59dc28c03a9639f6e5a2b11154aa20a9d4622a0",
      "parents": [
        "010ba431063daac214d6d04b72d330d2d4a7d562"
      ],
      "author": {
        "name": "Yuan",
        "email": "yuanzhou@apache.org",
        "time": "Wed Jul 15 15:12:30 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 15 08:12:30 2026 +0100"
      },
      "message": "[CI]Enable updates from depedency bot (#12508)\n\nThis patch enabled the auto update from github dependency bot.\n\nGithub bot will try to make a PR to fix the security issues:\nhttps://github.com/apache/gluten/security/dependabot\n\nThis can help to fix all CVE issues in Gluten quickly. Committers will need to decide to reject the patch if it\u0027s not working for Gluten in some cases. If the update PR is closed, the bot will not raise the PR on the same issue again.\n\nSigned-off-by: Yuan \u003cyuanzhou@apache.org\u003e"
    },
    {
      "commit": "010ba431063daac214d6d04b72d330d2d4a7d562",
      "tree": "5773f7db4b236ee9a2b6fdb59d1ba481d142cfae",
      "parents": [
        "0209ec8e16bcbdb9f6ac8d1ae87faf434b351b20"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Wed Jul 15 15:11:39 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 15 15:11:39 2026 +0800"
      },
      "message": "[GLUTEN-11550][VL] Fix GlutenWholeStageCodegenSuite in Spark 4.x (#12507)"
    },
    {
      "commit": "0209ec8e16bcbdb9f6ac8d1ae87faf434b351b20",
      "tree": "f2fc662474020a315836c582ab89c0b3bd84ee3c",
      "parents": [
        "f02988114dd41dacde1d9744f9ddc04a932ad1ce"
      ],
      "author": {
        "name": "xumanbu",
        "email": "manbu.feng@163.com",
        "time": "Wed Jul 15 14:27:33 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 15 07:27:33 2026 +0100"
      },
      "message": "[CORE] Log error when -Xms equals -Xmx makes JVM heap shrinking ineffective (#12480)\n\nAdd a check in DynamicOffHeapSizingMemoryTarget static initializer to detect when -Xms equals or is very close to -Xmx. In this scenario, the JVM heap size is fixed and System.gc() cannot shrink totalMemory (return committed pages to OS), making the shrinkOnHeapMemory mechanism ineffective. An error log is emitted to alert users, consistent with the existing checks for -XX:+ExplicitGCInvokesConcurrent and -XX:+DisableExplicitGC.\n\nCo-authored-by: jam.xu \u003cjamxu@shein.com\u003e"
    },
    {
      "commit": "f02988114dd41dacde1d9744f9ddc04a932ad1ce",
      "tree": "5eb06c25722435976f387dba418ef7da836f80ae",
      "parents": [
        "7f8940c0f6ea44a26f3c1fb369d54faeb4bf023d"
      ],
      "author": {
        "name": "kevinyhzou",
        "email": "37431499+KevinyhZou@users.noreply.github.com",
        "time": "Wed Jul 15 11:54:28 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 15 11:54:28 2026 +0800"
      },
      "message": "[GLUTEN-12206][FLINK]Support filesystem writer bucket state snapshot and restore (#12472)\n\n* [GLUTEN-12206][FLINK] Support filesystem writer bucket state snapshot and restore"
    },
    {
      "commit": "7f8940c0f6ea44a26f3c1fb369d54faeb4bf023d",
      "tree": "174dd6175760f6efd001808f21dc3e51b10cbde4",
      "parents": [
        "14683bb9303ceb5f962e0615434413f70ea88ca2"
      ],
      "author": {
        "name": "Philo He",
        "email": "philohe@amazon.com",
        "time": "Wed Jul 15 01:09:21 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 14 10:09:21 2026 -0700"
      },
      "message": "[GLUTEN-12288][VL] Remove redundant Spark 3.5 SMJ-enabled full test jobs to cut GHA usage (#12505)"
    },
    {
      "commit": "14683bb9303ceb5f962e0615434413f70ea88ca2",
      "tree": "7b11878f9d16456e4747da420fafdec440960786",
      "parents": [
        "d8ca277bbb8cac9cd01b68b599368b7313c900a0"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Tue Jul 14 13:30:56 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 14 13:30:56 2026 +0100"
      },
      "message": "[VL] Add timestamp to shuffle test (#12503)"
    },
    {
      "commit": "d8ca277bbb8cac9cd01b68b599368b7313c900a0",
      "tree": "3651b3dfdce9cf805b8cf57fc33e89687ad43f57",
      "parents": [
        "061b799fbb7dd7fbe7952daa375c1958d84d2df0"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Tue Jul 14 13:47:42 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 14 13:47:42 2026 +0800"
      },
      "message": "[VL] Fix GlutenSimpleSQLViewSuite in Spark 4.x (#12502)"
    },
    {
      "commit": "061b799fbb7dd7fbe7952daa375c1958d84d2df0",
      "tree": "1e53694649595686772b288f889c25f2c15ff5ba",
      "parents": [
        "e6cc336eeaecee387b367675255fc6f74d79fac6"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Tue Jul 14 13:30:45 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 14 13:30:45 2026 +0800"
      },
      "message": "[VL] Fix GlutenRemoveRedundantSortsSuite in Spark 4.x (#12497)"
    },
    {
      "commit": "e6cc336eeaecee387b367675255fc6f74d79fac6",
      "tree": "2014265e7ec87fb0b30f8223d49d3dc9fc484226",
      "parents": [
        "dd512bcbec5a7057023fea013115e18e5fadd127"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Tue Jul 14 11:38:52 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 14 11:38:52 2026 +0800"
      },
      "message": "[VL] Fix GlutenGroupBasedUpdateTableSuite in Spark 4.x (#12493)"
    },
    {
      "commit": "dd512bcbec5a7057023fea013115e18e5fadd127",
      "tree": "389325c1c5f073c538844185fbb0f86326c3762d",
      "parents": [
        "b45c1d87dafe08a6f8be39eae9a3f9f566efa058"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Mon Jul 13 12:32:03 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 13 12:32:03 2026 +0800"
      },
      "message": "[MINOR][CORE] Fix self-suppress in TaskResources.runUnsafe error handler (#12496)"
    },
    {
      "commit": "b45c1d87dafe08a6f8be39eae9a3f9f566efa058",
      "tree": "a69adfbb4d3faf05d56480f03b2a675213b77e49",
      "parents": [
        "9685185b9478b4ba1071fd2f03204942284bcad5"
      ],
      "author": {
        "name": "Gluten Performance Bot",
        "email": "137994563+GlutenPerfBot@users.noreply.github.com",
        "time": "Mon Jul 13 04:54:24 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 13 04:54:24 2026 +0100"
      },
      "message": "[GLUTEN-6887][VL] Daily Update Velox Version (2026_07_10) (#12494)\n\n* [GLUTEN-6887][VL] Daily Update Velox Version (dft-2026_07_10)\n\nUpstream Velox\u0027s New Commits:\n267246d0c by Philo He, refactor: Refactor cast error handling to build error message lazily\n899cbfa91 by Ismaël Mejía, perf(parquet): Add Parquet writer benchmark\n27c7f3d4a by Matt Gara, feat(cudf): Add NotFunction, IsNullFunction, IsNotNullFunction\nb3280cd37 by Muhammad Haseeb, refactor(cudf): Rework Iceberg schema evolution\n0f35c0ba6 by Pradeep Vaka, feat(exec): Fire stateChangeFuture on scan split queue drain (#17998)\n980395430 by Mariam AlMesfer, feat(sparksql): Add casts between TIMESTAMP  and TIMESTAMP_UTC\n1ffe42ed4 by Jiaan Geng, fix(spark): Fix out-of-bounds write in Spark `map()` when a key repeats 3+ times under LAST_WIN\nec17d38c7 by Krishna Pai, fix(ci): Isolate main sanitizer builds per commit\n20a1e5eda by Simon Eves, feat(cudf): GPU Decimal (Part 4)\n6d2cebaff by Yabin Ma, fix: Fix thread-local runtime stat writer leak in SortBufferTest\n3a542c2eb by Jimmy Lu, fix(torchwave): Correct and speed up ads-preproc on wave (#17893)\n080b050e3 by Pedro Pedreira, feat(vector): Add DecodedVector::hasNulls() (#18054)\n6987b0fe4 by Karthikeyan, refactor(exec): Rename OutputBufferManager to DefaultOutputBufferManager\n4c9a221f3 by Wechar Yu, fix(parquet): Flush row group by buffered bytes in the writer\n54978b5e2 by Hongze Zhang, build: Add ColumnReaderStatisticsTests.cpp to cmake\n974ccd669 by Suryadev Sahadevan Rajesh, fix: Handle nested leaf columns by dotted path (#17979)\n7d36cc363 by Pedro Pedreira, feat(vector): Add DecodedVector::forEachNull() and forEachNotNull() (#18051)\n7ca0a98f2 by zhichenxu-meta, fix(rpc): Degrade whole-batch RPC failures to per-row errors instead of failing the query (#18049)\nf962082ca by RindsSchei225e, feat(fuzzer): Add no-evolution mode and fix test flatmap column consistency (#17975)\nf8e59f971 by Patrick Wilson, feat(cudf): Add GPU-accelerated Window operator\na581bbda9 by Maria Basmanova, feat(cudf): Revert Optimize Velox-cudf Iceberg schema evolution (#18055)\n5cd0ef714 by zhichenxu-meta, feat(rpc): Add admission-controlled adaptive RPC rate limiter (#18039)\nea72605c3 by Muhammad Haseeb, Fix synthetic row count override\nc1f27ed9b by Muhammad Haseeb, MInor\nf8302e336 by Muhammad Haseeb, More fixes\n3f346d649 by Muhammad Haseeb, cleanup artifacts\nbd91b86a5 by Muhammad Haseeb, Address empty column projection case.\nf413d9699 by Muhammad Haseeb, Address comments from @kjmph\ndafb89b3e by Muhammad Haseeb, Apply suggestions from code review\n860f52594 by Muhammad Haseeb, Address comment\n50c0dd364 by Muhammad Haseeb, Use default constructor for `CudfIcebergConnectorFactory`\n9fee89f67 by Muhammad Haseeb, Minor\n829d4c9e8 by Muhammad Haseeb, Add a consumed counter\ndebc3ddb0 by Muhammad Haseeb, Minor syntax fix\n6e3c5be24 by Muhammad Haseeb, Rework schema evolution in Velox-cudf\n10e8b81c1 by Bikramjeet Singh Vig, fix: Prevent VectorFuzzer crash when normalizing encoded map keys\n6bacc2b52 by Apurva Kumar, feat: Iceberg V3 field-id + NIMBLE stats on the Velox Iceberg connector (#18041)\na4b60c92a by Krishna Pai, fix(ci): Remove custom import gate workflow\nd2f87f5e8 by Jiaan Geng, fix(spark): Escape supplementary-plane chars in get_json_object object/array results\n1251ff9f6 by Lukas Krenz, build: Add NVIDIA Vera compiler flags\n3c20437c8 by Apurva Kumar, feat: Iceberg V3 field-id resolver + NIMBLE reader/writer (FB-internal)[nimble] (#17895)\na9d808961 by Suryadev Sahadevan Rajesh, docs(blog): Add \"Making OpenZL available in Nimble OSS\" post (#18033)\nadbbe7538 by Shruti Shivakumar, fix(cudf): Handle zero-column inputs in CudfFromVelox\nc67a26a3b by Bradley Dice, fix(cudf): Update benchmark flags\n6461c38b4 by Reetika Agrawal, fix(iceberg): Filter pushdown with initial-default columns\n1a733d647 by Yuxuan Chen, perf: Call vector reserve to avoid unnecessary growth during insertion\nc7c4adbfc by Philo He, fix(spark): Align timestamp to integral cast with Spark overflow semantics under ANSI\n76cc6401e by RindsSchei225e, feat(fuzzer): Add random chunking mode in fuzzer (#17970)\nc90260d94 by Jimmy Lu, feat: Add fixed point operator plan nodes (#17814)\nfe7031c8e by oerling, test(torchwave): Synthetic-data end-to-end test for the ROO preproc graph (#18025)\n149468e77 by oerling, fix(torchwave): Fix ig pipelines\n79c57a486 by Christian Zentgraf, feat(abfs): Allow registration of multiple file systems (#1809)\n2f33d52db by Kyle Hubert, fix(cudf): Order packed table release on materialization stream\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\n\n* Fix compile issue\n\n* fix cast validation for TimestampNTZ\n\n* Disable failed unit tests\n\n---------\n\nSigned-off-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: glutenperfbot \u003cglutenperfbot@glutenproject-internal.com\u003e\nCo-authored-by: Ke Jia \u003cke.jia@ibm.com\u003e\nCo-authored-by: BInwei Yang \u003cfelixybw@apache.org\u003e\nCo-authored-by: Rong Ma \u003cmarong@apache.org\u003e"
    },
    {
      "commit": "9685185b9478b4ba1071fd2f03204942284bcad5",
      "tree": "d1a2f4c0c36b97ff1673dbb7637b97b0733ce9e0",
      "parents": [
        "cd37d7aef914350f66f0faebf34bbd071f5cfbea"
      ],
      "author": {
        "name": "kevinyhzou",
        "email": "37431499+KevinyhZou@users.noreply.github.com",
        "time": "Mon Jul 13 09:40:30 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 13 09:40:30 2026 +0800"
      },
      "message": "[GLUTEN-12206][FLINK] Support native checkpoint snapshot and restore for source operator (#12490)\n\n* feat: persist native source checkpoint state\n\n* fix: sync Gluten task output metrics\n\nfix: keep Gluten compatible with older velox4j APIs\n\nRevert \"fix: keep Gluten compatible with older velox4j APIs\"\n\nThis reverts commit 67a3cccbf6013af0e91c3a5b56e85bc4d869af2b.\n\n[FLINK] Drain Velox pipeline in prepareSnapshotPreBarrier for checkpoint consistency\n\nremove changes in source metrics\n\nremove metric changes\n\nremove metrics\n\n* update fink.md\n\n---------\n\nCo-authored-by: zhangzhibiao \u003czhangzhibiao@bigo.sg\u003e"
    },
    {
      "commit": "cd37d7aef914350f66f0faebf34bbd071f5cfbea",
      "tree": "60940329d9df236689b2653c966ed5377b5eedd8",
      "parents": [
        "022bc4b3794028e99d07960e314e531a09d0316e"
      ],
      "author": {
        "name": "Felipe Pessoto",
        "email": "felipepessoto@hotmail.com",
        "time": "Sun Jul 12 18:32:22 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 13 09:32:22 2026 +0800"
      },
      "message": "[VL] Fall back to vanilla Delta write for unsupported schemas (#12444)\n\nVelox cannot write every Spark type. When native Delta write is enabled, Gluten\noffloads certain Delta write commands to the native writer via OffloadDeltaCommand\n-- DataFrameWriter.save (append/overwrite), CREATE/REPLACE TABLE AS SELECT,\nUPDATE, DELETE and OPTIMIZE. The native write path inserts a RowToVeloxColumnarExec\ntransition whose SparkArrowUtil.toArrowSchema call throws\n`UnsupportedOperationException: Unsupported data type: \u003ctype\u003e` at runtime for any\ntype it has no Arrow mapping for (VariantType is the motivating example). (Plain\nSQL INSERT INTO and MERGE are not offloaded and were unaffected.)\n\nGuard GlutenOptimisticTransaction.writeFiles with the backend\u0027s existing schema\nvalidator: if VeloxValidatorApi.validateSchema reports the input schema is not\nsupported, delegate to super.writeFiles (the vanilla Delta write path) instead of\noffloading. This reuses the same check Velox already applies to scan schemas, so\nunsupported types fall back consistently before Arrow schema conversion rather\nthan adding a one-off VariantType guard. Supported writes are unaffected, and the\nvalidator\u0027s reason is logged when it falls back.\n\nAdds DeltaVariantWriteSuite, which exercises offloaded writes: top-level\nand struct-nested variant columns via DataFrameWriter.save, plus an UPDATE."
    },
    {
      "commit": "022bc4b3794028e99d07960e314e531a09d0316e",
      "tree": "d1372286520d46ef93d4d444269341776200ddc7",
      "parents": [
        "b57060effb744bb8a773cc766a0166dde2c879c6"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Mon Jul 13 08:53:28 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 13 08:53:28 2026 +0800"
      },
      "message": "[VL] Adjust Spark 4.x inherited suite coverage (#12484)"
    },
    {
      "commit": "b57060effb744bb8a773cc766a0166dde2c879c6",
      "tree": "e98d578cec01b30934c4c3a893539fe3774b1b38",
      "parents": [
        "5ecef3c0e16255f691e7df84e8269098128e8dfe"
      ],
      "author": {
        "name": "Joey",
        "email": "joey.ljy@alibaba-inc.com",
        "time": "Sat Jul 11 16:05:26 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Sat Jul 11 16:05:26 2026 +0800"
      },
      "message": "[MINOR] Fix Scala test cache invalidation for suffixed source roots (#12498)"
    },
    {
      "commit": "5ecef3c0e16255f691e7df84e8269098128e8dfe",
      "tree": "99bd98ef85d3aae6799b32328fecaa30cd87bc5b",
      "parents": [
        "7f5ff0def7916ad9d7504ac07b0f093c32409052"
      ],
      "author": {
        "name": "Rui Mo",
        "email": "rui@apache.org",
        "time": "Fri Jul 10 16:27:52 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 16:27:52 2026 +0100"
      },
      "message": "[VL] Refactor Velox metrics transport to use JSON payloads (#12126)"
    },
    {
      "commit": "7f5ff0def7916ad9d7504ac07b0f093c32409052",
      "tree": "fcc443aef47efb99923d27638ab4afaabf3fa5a1",
      "parents": [
        "26c2954f6bed910dbe6d050c99ad7dff38f73801"
      ],
      "author": {
        "name": "JiaKe",
        "email": "ke.jia@ibm.com",
        "time": "Fri Jul 10 12:30:01 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 12:30:01 2026 +0100"
      },
      "message": "[VL] Disable useHashTableCache if the hash table is not pre-built during HashJoinNode creation (#12487)"
    },
    {
      "commit": "26c2954f6bed910dbe6d050c99ad7dff38f73801",
      "tree": "6e2ec0caeaa21b8dce944e1584ac89454c5b9faf",
      "parents": [
        "fadf00629d66033609736ea6bbd9de695de51691"
      ],
      "author": {
        "name": "Mariam AlMesfer",
        "email": "Mariamalmesfer22@gmail.com",
        "time": "Fri Jul 10 13:52:36 2026 +0300"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 11:52:36 2026 +0100"
      },
      "message": "[GLUTEN-11622][VL] Enable minute/second(timestamp_ntz) native execution (#12295)\n\nCo-authored-by: Mariam-Almesfer \u003cmariam.almesfer@ibm.com\u003e"
    },
    {
      "commit": "fadf00629d66033609736ea6bbd9de695de51691",
      "tree": "10830aa38fee158107ac33d6a8345cb6fcdf0f62",
      "parents": [
        "c357f6b811dfc9de29cdde576cb6ec6795d994ec"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Fri Jul 10 17:53:18 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 10:53:18 2026 +0100"
      },
      "message": "[VL] Fix SPARK-54439 in GlutenKeyGroupedPartitioningSuite for Spark 4.1 (#12482)\n\nThis PR updates GlutenKeyGroupedPartitioningSuite for Spark 4.1 to handle the new SPARK-54439 test coverage.\n\nSpark 4.1 added tests for KeyGroupedPartitioning cases where the reported partitioning key size does not match the join key size. Since Gluten does not currently support these cases with native columnar shuffle, this PR excludes the original Spark-side SPARK-54439 tests from the Velox settings and adds Gluten-aware versions that validate the expected fallback behavior."
    },
    {
      "commit": "c357f6b811dfc9de29cdde576cb6ec6795d994ec",
      "tree": "597da7239fbc097dcd3d8c8a171677d8ae0fa6d0",
      "parents": [
        "fc756ffae8ad5fb9d6c33248e975d17336b2cde7"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Fri Jul 10 17:51:47 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 10:51:47 2026 +0100"
      },
      "message": "[VL] Fix GlutenVariantInferShreddingSuite in Spark 4.1 (#12481)"
    },
    {
      "commit": "fc756ffae8ad5fb9d6c33248e975d17336b2cde7",
      "tree": "a4be12e7242f5b5666f2d0b3cc143dad0a24f532",
      "parents": [
        "c5cb134c792008844084c1b68b40ee4db0407d81"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Fri Jul 10 10:01:53 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 10:01:53 2026 +0100"
      },
      "message": "[VL] Support GPU async native shuffle read (#12370)"
    },
    {
      "commit": "c5cb134c792008844084c1b68b40ee4db0407d81",
      "tree": "76ca0286d0ff4620cbed4a8554b7a83cfe8da3b4",
      "parents": [
        "2479a6f168ddba627f5aee10187bcf6fc01f1014"
      ],
      "author": {
        "name": "Ismaël Mejía",
        "email": "iemejia@gmail.com",
        "time": "Fri Jul 10 10:06:48 2026 +0200"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 09:06:48 2026 +0100"
      },
      "message": "[GLUTEN][VL] Optimize Delta Lake DV materialization and plan rule performance (#12390)\n\n* [GLUTEN][VL] Eliminate per-file I/O and allocation in DV materialization\n\nCache the resolved table path and Hadoop Configuration across all files\nin a partition during normalize(). Previously, each file triggered\nindependent filesystem exists() checks (to find the _delta_log\ndirectory) and allocated a fresh Hadoop Configuration clone. For a\npartition with N files on object storage, this produced N+ redundant\nHTTP HEAD requests on the driver critical path.\n\nFor on-disk DVs, read the raw bytes directly from the DV file using\nDelta\u0027s DeletionVectorStore.readRangeFromStream (which includes\nchecksum verification) instead of going through StoredBitmap.load()\n+ serializeAsByteArray(). The on-disk format is already Portable\nRoaring Bitmap Array -- the same format the native Velox side expects\n-- so this eliminates the expensive deserialize-into-Java-Roaring-\nobjects + re-serialize round-trip per file.\n\nChanges:\n- Resolve table path once using the first file, reuse for all others\n- Create one Hadoop Configuration per normalize() call\n- Read raw DV bytes directly for on-disk DVs (skip deser+reser)\n- Fall back to load+serialize for inline DVs (small, in-metadata)\n- (delta40) Cache the reflective method lookup for parseDescriptor\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* [GLUTEN][VL] Optimize Delta post-transform rules for non-Delta queries\n\nReduce plan traversal overhead from 5 full passes to effectively 1 for\nDelta queries and 0 for non-Delta queries:\n\n- Add early-exit guard: check plan.exists(DeltaScanTransformer) once and\n  skip all Delta-specific rules if no Delta scan is present. This\n  eliminates all overhead for non-Delta queries.\n- Replace quadratic containsNativeDeltaScan (full subtree .exists() per\n  Filter/Project node) with a shallow 2-level child check that is O(1),\n  safe because transformUp processes bottom-up.\n- Pre-compute inputFileRelatedNames as a static Set[String] instead of\n  allocating 3 Expression objects + 2 Seqs per call per column.\n- Batch createPhysicalAttributes: single call with full attribute list\n  instead of per-column invocation that walks the reference schema N\n  times for a table with N columns.\n- Fuse nativeDeletionVectorRule, pushDownInputFileExprRule, and\n  columnMappingRule into a single registered rule to reduce the number\n  of injected post-transforms from 4 to 2.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* [GLUTEN][VL] Reduce allocation in Delta scan and DV serialization\n\n- DeltaScanTransformer.scanFilters: change from def to lazy val to\n  avoid rebuilding the physicalByExprId map and re-traversing filter\n  expression trees on every call (invoked 3+ times per scan node).\n\n- DeltaLocalFilesNode: use UnsafeByteOperations.unsafeWrap() instead of\n  ByteString.copyFrom() for the DV byte array. This is a zero-copy\n  wrap since the byte[] lifetime is guaranteed by DeltaFileReadOptions,\n  eliminating an O(DV_size) memcpy per file on the driver.\n\n- LocalFilesNode: improve documentation on the copy constructor noting\n  that the original is discarded after construction.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* [GLUTEN][VL] Add DeltaPlanningBenchmark for JVM-side planning perf\n\nAdds a Spark Benchmark that measures the two hot paths optimized in\nthis patch series:\n\n1. DV Materialization (DeltaDeletionVectorScanInfo.normalize):\n   Creates a Delta table with N DV-bearing files and times the\n   normalize() call that resolves table paths, loads DV bitmaps, and\n   builds split metadata. Directly measures the impact of caching\n   table path + Hadoop conf + DV store (\"Eliminate per-file I/O\" commit).\n\n2. Post-transform rule application (DeltaPostTransformRules.rules):\n   Applies the Delta post-transform rules to a plan containing\n   DeltaScanTransformer nodes. Measures rule traversal overhead\n   including the early-exit guard, shallow containsNativeDeltaScan,\n   pre-computed names, and batched attribute mapping (\"Optimize Delta post-transform rules\" commit).\n\n3. Non-Delta plan overhead (control):\n   Applies the same rules to a plain parquet plan to verify the\n   early-exit guard produces zero overhead for non-Delta queries.\n\nConfigurable via spark.gluten.benchmark.delta.numFiles (default 100)\nand spark.gluten.benchmark.delta.rowsPerFile (default 10000).\n\nMeasured results (local filesystem, 100 DV-bearing files):\n\n  Benchmark                         Before    After     Speedup\n  -------                           ------    -----     -------\n  DV Materialization (100 files)    22 ms     7 ms      3.3x\n  Post-transform rules (Delta)     37 us     20 us      1.8x\n  Post-transform rules (parquet)   4908 ns   220 ns    22.3x\n\nCall count reduction for 100 DV-bearing files:\n\n  Operation                         Before    After    Eliminated\n  ---------                         ------    -----    ----------\n  FileSystem.exists() (HEAD reqs)   100-300   1        99-299\n  newHadoopConf() (deep clone)      100-300   1        99-299\n  new HadoopFileSystemDVStore()     100       1        99\n  Plan tree traversals (non-Delta)  5         0        5\n  Plan tree traversals (Delta)      5         1        4\n  containsNativeDeltaScan subtree   O(n^2)    O(1)     --\n  createPhysicalAttributes calls    N cols    1        N-1\n\nProjected DV materialization time by storage backend (100 files):\n\n  Storage    exists() latency    Before         After       Speedup\n  -------    ----------------    ------         -----       -------\n  Local FS   ~67 us/call         22 ms          7 ms        3.3x\n  ABFS       20-80 ms/call       2-24 sec       1.0-1.1 s   2-22x\n  GCS        30-100 ms/call      3-30 sec       1.0-1.1 s   3-27x\n  S3         50-150 ms/call      5-45 sec       1.1-1.2 s   5-38x\n\n  After \u003d 1 exists() call + 100 DV loads (~10 ms each on object stores)\n  Before \u003d 100-300 exists() calls + 100 DV loads\n\nRemote object storage impact analysis:\n\nThe dominant cost in DV materialization is resolveTablePath(), which\ncalls FileSystem.exists() to locate the _delta_log directory. On local\nFS this is ~67us per call; on object stores each exists() is an HTTP\nHEAD request with the latencies shown above.\n\nBefore this patch, resolveTablePath() was called per-file, plus\nisDeltaTablePath() could walk up 1-3 parent directories per file.\nAfter: a single exists() call resolves the table path for all files.\n\nThe DV bitmap load (StoredBitmap.load) remains per-file but benefits\nfrom connection pooling via the shared HadoopFileSystemDVStore (the\nFS instance is reused across all files since the same Configuration\nobject hits Hadoop\u0027s FileSystem cache).\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* [MINOR] Add Eclipse project files to .gitignore\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Address review comments: defensive guards, test improvements, fallback parse\n\n- Fix misleading \u0027direct list reference transfer\u0027 comment in LocalFilesNode\n  to accurately describe the shallow list copy behavior.\n- Add empty partitionFiles guard in normalize() for both delta33 and delta40\n  to prevent NoSuchElementException on empty input.\n- Strengthen test assertions: use \u0027eq\u0027 identity check for early-exit guard,\n  rename test to match actual behavior, replace silent \u0027if\u0027 with assert.\n- Fix --jars doc syntax to use comma-separated format in benchmark.\n- Remove orphaned parseDescriptor Scaladoc block.\n- Cache all available parse methods and try in order, preserving fallback\n  semantics while avoiding per-call getMethod overhead.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6"
    },
    {
      "commit": "2479a6f168ddba627f5aee10187bcf6fc01f1014",
      "tree": "db027c11ded4dfa065a12ce4b8e93126292014e2",
      "parents": [
        "268ae68667fcfa0eb0c9848efddf080c03ee0219"
      ],
      "author": {
        "name": "Hw-TinY",
        "email": "yangyi261@huawei.com",
        "time": "Fri Jul 10 10:19:32 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 10 10:19:32 2026 +0800"
      },
      "message": "[GLUTEN-12439] [FLINK] Add CONCAT / CONCAT_WS expression registration (#12452)\n\n* Add UT for CONCAT / CONCAT_WS functions\n\n---------\n\nCo-authored-by: Hw-TinY \u003chw-tiny@users.noreply.github.com\u003e"
    },
    {
      "commit": "268ae68667fcfa0eb0c9848efddf080c03ee0219",
      "tree": "ebdf22e7d897c53c5ee4e03751c4b9352741b1fc",
      "parents": [
        "85d3ca109ebe38168b2a2e0503490f554df3ed3e"
      ],
      "author": {
        "name": "Wechar Yu",
        "email": "yuwq1996@gmail.com",
        "time": "Thu Jul 09 18:40:27 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 11:40:27 2026 +0100"
      },
      "message": "[GLUTEN-12449] Remove unsupported functions in ArrowWritableColumnVector (#12450)"
    },
    {
      "commit": "85d3ca109ebe38168b2a2e0503490f554df3ed3e",
      "tree": "f8cf337a4a00c6c47ff594e0109d8736bbcb179b",
      "parents": [
        "0f8074c8fd0bb5960242b601ffaa5c7d01d4bc3d"
      ],
      "author": {
        "name": "Reema",
        "email": "80041251+ReemaAlzaid@users.noreply.github.com",
        "time": "Thu Jul 09 10:02:05 2026 +0000"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 18:02:05 2026 +0800"
      },
      "message": "[GLUTEN-12280][VL] Fix Spark 4 Arrow Python UDF stream writer (#12345)\n\nWhat changes are proposed in this pull request?\nFixes #12280.\n\nFix Spark 4 Arrow Python UDF execution with the Velox backend by keeping the Arrow stream writer alive across input batches instead of reopening the IPC stream per batch.\n\nAlso adds a regression test for Arrow Python UDF over Parquet scan\n\nHow was this patch tested?\nAdded ArrowEvalPythonExecSuite coverage.\n\nVerified locally on Spark 4.0.2 / Scala 2.13 / linux aarch64. The repro uses ColumnarArrowPythonRunner, returns max(ship_len) \u003d 7, and no longer fails with Invalid IPC stream"
    },
    {
      "commit": "0f8074c8fd0bb5960242b601ffaa5c7d01d4bc3d",
      "tree": "5dbe8cd663951f46082f88c3bb634b90bd24cc46",
      "parents": [
        "a1caefecb4c5b855c47cee4903c05e97eaa4b67f"
      ],
      "author": {
        "name": "JiaKe",
        "email": "ke.jia@ibm.com",
        "time": "Thu Jul 09 10:02:39 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 10:02:39 2026 +0100"
      },
      "message": "[VL]Use Velox\u0027s HashTableCache to cache the BHJ\u0027s HashTable (#12163)"
    },
    {
      "commit": "a1caefecb4c5b855c47cee4903c05e97eaa4b67f",
      "tree": "67c0cdba9a16cf136959cd890d46fe840c9e153a",
      "parents": [
        "31adaf8756f0017cceb15e608d40e3b53152b6b1"
      ],
      "author": {
        "name": "Rong Ma",
        "email": "marong@apache.org",
        "time": "Thu Jul 09 09:37:48 2026 +0100"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 09:37:48 2026 +0100"
      },
      "message": "[MINOR][VL] Fix explain with ColumnarPartialGenerateExec (#12478)"
    },
    {
      "commit": "31adaf8756f0017cceb15e608d40e3b53152b6b1",
      "tree": "2b18f7d587569a3e58ead272bddbc36772ba59af",
      "parents": [
        "5d9bced5d4c80115e4adfa5c82948615a5a31277"
      ],
      "author": {
        "name": "Navaneeth Sujith",
        "email": "43727786+navaneethsujith09@users.noreply.github.com",
        "time": "Wed Jul 08 23:11:14 2026 -0700"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 14:11:14 2026 +0800"
      },
      "message": "[GLUTEN-4730][CH] feat: Support date_from_unix_date function: Add ClickHouse backend implementation (#12410)\n\n* Addressed comments.\n\n* removed trailing spaces\n\n* added a curly brace\n\n* test to see if this is the fix"
    },
    {
      "commit": "5d9bced5d4c80115e4adfa5c82948615a5a31277",
      "tree": "32ac8059a62e8b0f73d8dce133b90472c39b860b",
      "parents": [
        "ef647cac78e216ab475b4ab87a2c2c8694ccb1df"
      ],
      "author": {
        "name": "BRIJ RAJ KISHORE",
        "email": "22271048+brijrajk@users.noreply.github.com",
        "time": "Thu Jul 09 11:04:02 2026 +0530"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 08 22:34:02 2026 -0700"
      },
      "message": "[GLUTEN-12157][VL] Fix silently-skipped math/scalar test suites and add Velox native tests for sin, tan, tanh, radians, ln (#12158)"
    },
    {
      "commit": "ef647cac78e216ab475b4ab87a2c2c8694ccb1df",
      "tree": "dced246f7f7ad4319f0e96c3e5c36eba4a33b7ba",
      "parents": [
        "75d6d1f0d0265535a6aafde544400649dbe7e9f7"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 09 13:33:13 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 13:33:13 2026 +0800"
      },
      "message": "[CORE] Limit SparkSessionExtension suite to Gluten test (#12475)"
    },
    {
      "commit": "75d6d1f0d0265535a6aafde544400649dbe7e9f7",
      "tree": "759f20ceb8fa5799d44dccbbb0600cc507967daf",
      "parents": [
        "065d2e30de86a174a11ee26cf2c5b0df04ff1523"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 09 13:32:39 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 13:32:39 2026 +0800"
      },
      "message": "[VL] Fix SPARK-53322 in GlutenKeyGroupedPartitioningSuite for Spark 4.1 (#12469)"
    },
    {
      "commit": "065d2e30de86a174a11ee26cf2c5b0df04ff1523",
      "tree": "b5564cbb6b687dc51beb263773cf3859b63d2bc9",
      "parents": [
        "8aeff651e538b4912d079066fc4ab5cd967b051a"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Thu Jul 09 13:32:03 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 13:32:03 2026 +0800"
      },
      "message": "[MINOR][CORE] Avoid O(n^2) stack walk in CallerInfo.inBloomFilterStatFunctionCall (#12477)"
    },
    {
      "commit": "8aeff651e538b4912d079066fc4ab5cd967b051a",
      "tree": "e402b76317ea263a9d7b97119ab9b87f33bba9c4",
      "parents": [
        "6f179f5814fee88d13f712ebcecdc7b470d856fb"
      ],
      "author": {
        "name": "jackylee",
        "email": "qcsd2011@gmail.com",
        "time": "Thu Jul 09 10:12:49 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 10:12:49 2026 +0800"
      },
      "message": "[VL] Support test/debug/benchmark native builds on macOS (follow-up to #12331) (#12470)"
    },
    {
      "commit": "6f179f5814fee88d13f712ebcecdc7b470d856fb",
      "tree": "86553ba17c6a993a29cca49cc8f69ec4d2c72588",
      "parents": [
        "630d5ace8af5287e49c0942835450198babc121c"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Thu Jul 09 10:10:28 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Thu Jul 09 10:10:28 2026 +0800"
      },
      "message": "[MINOR] Remove stale complex type scan fallback config from tests (#12476)"
    },
    {
      "commit": "630d5ace8af5287e49c0942835450198babc121c",
      "tree": "1a40ca9ce3094dd2815504664c3eb2c597102c42",
      "parents": [
        "a6202db2856fff1952fc1fc0d62df5edfa1b3168"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Wed Jul 08 20:56:25 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 08 20:56:25 2026 +0800"
      },
      "message": "[CORE] optimize driver heap memory by lazily store metadata columns and fix mismatch result when input_file_name is only in filter condition (#11899)\n\n* avoid driver oom caused by unnecessary metadata columns in splitinfo\n\n* support input_file_name only in filter\n\n* fix code style"
    },
    {
      "commit": "a6202db2856fff1952fc1fc0d62df5edfa1b3168",
      "tree": "bede1f4e14340f655a86396a4736fc619851a595",
      "parents": [
        "9baa47c4240bd761878c9f99dc2fd96d9749b83b"
      ],
      "author": {
        "name": "Mingliang Zhu",
        "email": "zhuml1206@gmail.com",
        "time": "Wed Jul 08 16:15:58 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Wed Jul 08 16:15:58 2026 +0800"
      },
      "message": "[VL] Fix GlutenSubquerySuite in Spark 4.x (#12473)"
    },
    {
      "commit": "9baa47c4240bd761878c9f99dc2fd96d9749b83b",
      "tree": "b5e03c136b37df401d3773baf31efb6c32a95157",
      "parents": [
        "cd994505b08c6ae37ab6813a56c0139b226f4e66"
      ],
      "author": {
        "name": "Reema",
        "email": "80041251+ReemaAlzaid@users.noreply.github.com",
        "time": "Wed Jul 08 03:08:37 2026 +0000"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 20:08:37 2026 -0700"
      },
      "message": "[GLUTEN-12303][VL] Support async multipart upload for S3 writes (#12305)\n\n* feat:Support multi threaded async data upload to object storage\n\nThe PR is ported from facebookincubator/velox#14472. The original author is @weixiuli"
    },
    {
      "commit": "cd994505b08c6ae37ab6813a56c0139b226f4e66",
      "tree": "27a3acc537dcff48e744b5eb04ebf750f3ac4bf1",
      "parents": [
        "4017ec94d34cf2bdaa80b0be61af980a8cdc1959"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Wed Jul 08 01:20:52 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 18:20:52 2026 +0100"
      },
      "message": "[GLUTEN-12465][CORE] Support optional stage InputStats plumbing in scan/input-iterator transformers (#12433)\n\n* [GLUTEN][CORE] Support optional stage InputStats plumbing in scan/input-iterator transformers\n\nIntroduce an opt-in framework to pass stage input statistics (scan / shuffle /\nbroadcast) to the native engine as estimated row-size hints.\n\n- Add InputStats / InputStatsKind / ShuffleStageWrapper and ApplyStageInputStatsRule.\n- Add RelNode.childNode() with implementations across all rel nodes; carry\n  rowSize / inputStats on ReadRelNode and InputIteratorRelNode; expose\n  PlanNode.getRelNodes() and AdvancedExtensionNode optimization/enhancement getters.\n- Add BasicScanExecTransformer.getInputStats hook and optional inputStats field on\n  FileSourceScanExecTransformer; thread wsContext through genFirstStageIterator and\n  the whole-stage RDDs; inject shuffle/broadcast stats in InputIteratorTransformer.\n- Gate everything behind spark.gluten.sql.enablePassStageInputStats (default false),\n  so velox / clickhouse behavior is unchanged.\n\nGenerated-by: TraeCli openrouter-3o\nCo-Authored-By: Aime \u003caime@bytedance.com\u003e\nChange-Id: I7cd0a5abc98b6c41857230131513bfcb5ec6bf95\n\n* change as request\n\n---------\n\nCo-authored-by: Aime \u003caime@bytedance.com\u003e\nCo-authored-by: liyang.127 \u003cliyang.127@bytedance.com\u003e"
    },
    {
      "commit": "4017ec94d34cf2bdaa80b0be61af980a8cdc1959",
      "tree": "b140290ef9f2f66f33a973b3e6f2e9aaeb9f226a",
      "parents": [
        "ef309d210edd8e2ee47a1a8c20785e65d5b6a276"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Tue Jul 07 18:29:58 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 18:29:58 2026 +0800"
      },
      "message": "[GLUTEN][CORE] Expose ColumnarBatches.isLightBatch/ensureOffloaded as public (#12431)\n\nThese two helpers are needed by out-of-package callers that produce or\nconsume native-based columnar batches (e.g. a custom parquet batch\nwriter under org.apache.spark.sql.execution.datasources), which cannot\nreach the package-private methods. Promote them to public so backends\ncan reuse the batch type check and offload logic without duplicating it.\n\nGenerated-by: TraeCli openrouter-3o\n\nChange-Id: I9ec9927cff51f967ae807f9b755c63776b3f8127\n\nCo-authored-by: Aime \u003caime@bytedance.com\u003e"
    },
    {
      "commit": "ef309d210edd8e2ee47a1a8c20785e65d5b6a276",
      "tree": "e2d84d12e6b5f4d150b4e59a8b20237c7d2603d3",
      "parents": [
        "cfef50944aa7e10db4f6263c5530dff6d32c0dad"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Tue Jul 07 18:13:21 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 18:13:21 2026 +0800"
      },
      "message": "[GLUTEN][CORE] Add backend hooks for native config forwarding (#12437)\n\nAllow backends to contribute extra non-prefixed session and backend config keys so native config forwarding can be extended without hard-coding backend-specific entries in GlutenConfig."
    },
    {
      "commit": "cfef50944aa7e10db4f6263c5530dff6d32c0dad",
      "tree": "eef2a402636edc6c5c7cd102ca449deb253107de",
      "parents": [
        "bd2c1ce6bee5a88a706c57e3c24dd12198f6bfb7"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Tue Jul 07 17:50:15 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 17:50:15 2026 +0800"
      },
      "message": "[GLUTEN][CORE] Add TaskContext JNI callback for reading Spark task attempt id from native (#12435)\n\nInstead of extending every JNI entry point (createRuntime / MemoryManager\ncreate/hold/release) to plumb the Spark task attempt id from Java down to\nC++ as an extra parameter, expose a small callback surface that native code\nuses on demand:\n\n- Java side: org.apache.gluten.task.TaskContextJniWrapper#currentTaskAttemptId()\n  reads TaskContext.get().taskAttemptId() on the current thread and returns\n  -1 when there is no task context.\n- C++ side: gluten::getCurrentSparkTaskAttemptId() attaches the current\n  thread to the JVM as a daemon on demand and calls back into the Java\n  helper via JNI. The class ref and method id are cached in function-local\n  statics on first use.\n\nBecause Spark\u0027s TaskContext is a per-thread ThreadLocal, this returns a\nmeaningful value whenever the native call is running on an executor task\nthread (or any thread inheriting that ThreadLocal), which is exactly when\nbackends need it.\n\nNo signature change to Runtime / MemoryManager / RuntimeJniWrapper /\nNativeMemoryManagerJniWrapper. No behavior change for existing backends\n(Velox, ClickHouse) that do not query the task attempt id from native.\n\n\nChange-Id: I3185249796b0c396813dc39f54bd8e8b8589ca2a\n\nCo-authored-by: Aime \u003caime@bytedance.com\u003e"
    },
    {
      "commit": "bd2c1ce6bee5a88a706c57e3c24dd12198f6bfb7",
      "tree": "e50804098981d6f3fa9a7cf4c49042514e474976",
      "parents": [
        "9f9c2482a08e542d6ebdf6b5526ab047896ac14f"
      ],
      "author": {
        "name": "Reema",
        "email": "80041251+ReemaAlzaid@users.noreply.github.com",
        "time": "Tue Jul 07 08:11:30 2026 +0000"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Tue Jul 07 09:11:30 2026 +0100"
      },
      "message": "[VL] Support cudf.allowCpuFallback config for Velox cuDF backend (#12451)"
    },
    {
      "commit": "9f9c2482a08e542d6ebdf6b5526ab047896ac14f",
      "tree": "1123fe1cc800e6ed2e61d716aa85eb34da100c00",
      "parents": [
        "8a987f8b81f61b429123431e01d90996c131ad04"
      ],
      "author": {
        "name": "James Jenkins",
        "email": "108480513+Jenkins-J@users.noreply.github.com",
        "time": "Mon Jul 06 12:01:42 2026 -0400"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 06 17:01:42 2026 +0100"
      },
      "message": "[MINOR][VL] Fix arrow 18 build errors on Power (#12408)\n\n* Fix arrow 18 build errors\n\nCorrect the path to openssl source location. Change openssl install directory to allow arrow to resolve openssl. Include missing patch file for ppc46le support in arrow 18.\n\n* Fix openssl install path for Power\n\nInstall openssl to a non-standard directory to avoid interfering with building Velox and Gluten."
    },
    {
      "commit": "8a987f8b81f61b429123431e01d90996c131ad04",
      "tree": "8b6e21e7e37a798496d2c147cfa762b99b87c71a",
      "parents": [
        "0a16d3272c7199fe7b4a28104f33df1cba8bbc4d"
      ],
      "author": {
        "name": "wangguangxin.cn",
        "email": "wangguangxin.cn@bytedance.com",
        "time": "Fri Jun 05 17:26:04 2026 +0800"
      },
      "committer": {
        "name": "WangGuangxin",
        "email": "wangguangxin.cn@gmail.com",
        "time": "Mon Jul 06 22:57:13 2026 +0800"
      },
      "message": "rewrite multi-children Count in window expressions\n"
    },
    {
      "commit": "0a16d3272c7199fe7b4a28104f33df1cba8bbc4d",
      "tree": "d51a5ba729906f74bd8affe71fa7a8e13d957b97",
      "parents": [
        "6853c22b992eaac57f6f9a9313b27c3bcc37d599"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Mon Jul 06 21:15:44 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Mon Jul 06 14:15:44 2026 +0100"
      },
      "message": "[MINOR][CORE] Fail fast in InsertTransitions when a child has no recognizable Convention (#12442)"
    },
    {
      "commit": "6853c22b992eaac57f6f9a9313b27c3bcc37d599",
      "tree": "46d51a02a07c8f67d122023082933af3037689bb",
      "parents": [
        "999a949984d700e27cb9f782f8725bc5fddcaa98"
      ],
      "author": {
        "name": "BRIJ RAJ KISHORE",
        "email": "22271048+brijrajk@users.noreply.github.com",
        "time": "Sun Jul 05 17:35:41 2026 +0530"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Sun Jul 05 20:05:41 2026 +0800"
      },
      "message": "[GLUTEN-10992][VL] Fix MatchError for KeyGroupedPartitioning in native shuffle (#12335)\n\nWhat changes were proposed in this pull request?\nWhen Spark 4.0\u0027s V2 bucketing shuffle (spark.sql.v2.bucketing.shuffle.enabled\u003dtrue) is used in a join where only one side reports partitioning, Spark generates a ShuffleExchangeExec with KeyGroupedPartitioning as its output partitioning.\n\nThe default case _ \u003d\u003e in VeloxSparkPlanExecApi.genColumnarShuffleExchange created a ColumnarShuffleExchangeExec for this node without validation. When the query executed, ExecUtil.genShuffleDependency crashed with a scala.MatchError because KeyGroupedPartitioning was missing from its exhaustive match.\n\nChanges:\n\nVeloxSparkPlanExecApi.genColumnarShuffleExchange: add an explicit case _: KeyGroupedPartitioning \u003d\u003e before the default that adds a fallback tag and returns the vanilla ShuffleExchangeExec. This prevents a ColumnarShuffleExchangeExec from being created for an unsupported partitioning type.\nExecUtil.genShuffleDependency: add an explicit wildcard case other \u003d\u003e that throws GlutenNotSupportException instead of the cryptic scala.MatchError, as a defensive guard for any future unknown partitioning types.\nHow was this patch tested?\nThe existing testGluten(\"SPARK-41471: shuffle one side: only one side reports partitioning\") tests in GlutenKeyGroupedPartitioningSuite (both spark40 and spark41) reproduce the crash exactly — they set V2_BUCKETING_SHUFFLE_ENABLED\u003dtrue with only one bucketed side, which triggers a ShuffleExchangeExec with KeyGroupedPartitioning output and then call checkAnswer. After this fix these tests pass without MatchError.\n\nWas this patch authored or co-authored using generative AI tooling?\nGenerated-by: Claude Code (https://claude.ai/code)\n\nRelated issue: #10992"
    },
    {
      "commit": "999a949984d700e27cb9f782f8725bc5fddcaa98",
      "tree": "903060c051e3800ceeb3f9f5d8059ee5eee3d8b8",
      "parents": [
        "794728da671049c6cc78efff429a51c995776a9a"
      ],
      "author": {
        "name": "inf",
        "email": "ialhazmim@gmail.com",
        "time": "Sat Jul 04 09:29:08 2026 +0000"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Sat Jul 04 10:29:08 2026 +0100"
      },
      "message": "[VL] feat(iceberg): Added dictionary size bytes write config (#12434)\n\nAdded a new config: spark.gluten.sql.columnar.backend.velox.parquet.dictionaryPageSizeBytes to control the write.parquet.dict-size-bytes write confi"
    },
    {
      "commit": "794728da671049c6cc78efff429a51c995776a9a",
      "tree": "3dcb96f10b33ad405bf47be314a7412c1a81cf12",
      "parents": [
        "dfcf4378c6401cec6880eca02ad6f9a1a3648e5c"
      ],
      "author": {
        "name": "YangJie",
        "email": "yangjie01@baidu.com",
        "time": "Sat Jul 04 00:24:29 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 03 17:24:29 2026 +0100"
      },
      "message": "[MINOR][CORE] Rewrite OverAcquire class comments to match current behavior (#12443)"
    },
    {
      "commit": "dfcf4378c6401cec6880eca02ad6f9a1a3648e5c",
      "tree": "d1fa488652f5edb54793c58a75177df6c198168b",
      "parents": [
        "5eee3976ef547a4608440c336a34e92f35189d17"
      ],
      "author": {
        "name": "Ismaël Mejía",
        "email": "iemejia@gmail.com",
        "time": "Fri Jul 03 18:23:09 2026 +0200"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 03 17:23:09 2026 +0100"
      },
      "message": "[GLUTEN][VL] Fix redundant copies in Delta DV plan conversion (#12389)\n\nVeloxPlanConverter.cc (parseDeltaSplitInfo):\n- Fix double dynamic_pointer_cast: the original code called\n  dynamic_pointer_cast\u003cDeltaSplitInfo\u003e twice (once in the ternary\n  condition, once in the true branch). Use a single cast + null check.\n- Avoid unnecessary std::string copy from protobuf: the accessor\n  serialized_deletion_vector() returns a const std::string\u0026 to the\n  internal protobuf field, but \u0027auto\u0027 (by value) triggered a full\n  string copy into a local variable. Changed to \u0027const auto\u0026\u0027 to\n  bind directly to the protobuf internal field, eliminating the\n  intermediate copy. The shared_ptr\u003cstring\u003e construction then copies\n  once from the reference (unavoidable since it needs ownership).\n\nAssisted-by: GitHub Copilot:claude-opus-4.6"
    },
    {
      "commit": "5eee3976ef547a4608440c336a34e92f35189d17",
      "tree": "00d74c6cb77d5b3a637a17019f49b6bf9239f652",
      "parents": [
        "51a9e96516d4d7414a52027d2019acf46e2b6009"
      ],
      "author": {
        "name": "Ismaël Mejía",
        "email": "iemejia@gmail.com",
        "time": "Fri Jul 03 18:22:58 2026 +0200"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 03 17:22:58 2026 +0100"
      },
      "message": "[VL] Optimize Delta DV applyDeletionFilter with iterator-based bulk lookup (#12395)\n\n* [VL] Optimize Delta DV applyDeletionFilter with iterator-based bulk lookup\n\nReplace per-row contains() calls with an iterator-based scan in\nDeltaDeletionVectorReader::applyDeletionFilter().\n\nBefore: O(batch_size) calls to Roaring64Map::contains(), each performing\na std::map lookup (O(log N) on the high-32-bit buckets) plus a Roaring\ncontainer check per row position.\n\nAfter: O(deletions_in_range) iterator advances using move_equalorlarger()\nto seek directly to the first deleted row in the batch range, then\nsequential ++iterator advances (O(1) amortized within a Roaring\ncontainer) until the range end.\n\nBenchmark results (1M row file, batches of 4096, scanning full file):\n\n  Deletion %    Baseline (ms)    Optimized (ms)    Speedup\n  1% (sparse)       9.90             0.050           198x\n  10%               4.06             0.398           10.2x\n  50%               3.82             1.99            1.9x\n  90%               4.37             3.64            1.2x\n\nFor the common case of sparse deletions (1%, typical for Delta\nMERGE/UPDATE), this is a 198x improvement. The optimization is most\neffective when deletions are sparse relative to batch size, which is the\ndominant production pattern.\n\nThe baseline O(batch_size) approach spent constant time (~4ms per 1M\nrows) regardless of deletion density because it always probed every row.\nThe optimized O(deletions) approach scales linearly with actual\ndeletions: 0.05ms at 1%, 0.4ms at 10%, 2ms at 50%.\n\nTests:\n- 21 unit tests pass (12 existing + 9 new corner cases covering: all\n  rows deleted, no rows deleted, first/last row, large 64-bit offsets,\n  single-row batches, dense/sparse mixed patterns, batch before/after\n  all deletions).\n- New BM_ApplyDeletionFilter benchmark added to DeltaBitmapBenchmark.cc\n  with 5 scenarios (sparse/moderate/dense/very-dense/large-batch).\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Fix benchmark counter accuracy for tail rows\n\nAddress review feedback: SetItemsProcessed now reports the actual rows\nprocessed (numBatches * batchSize) rather than totalFileRows, which\nslightly overstated the work when totalFileRows is not evenly divisible\nby batchSize. Added inline comment clarifying padding bits in popcount\nare safe due to memset zeroing in applyDeletionFilter.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Apply clang-format-15 formatting\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Address review: overflow guard on rangeEnd, move MemoryManager init to main()\n\n- Add saturating arithmetic for rangeEnd to prevent uint64_t overflow\n  when baseReadOffset is near UINT64_MAX.\n- Move memory::MemoryManager::testingSetInstance() from benchmark body\n  to a custom main(), ensuring it runs once per process.\n- Add ApplyDeletionFilterOverflowProtection unit test.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Fix narrowing conversion error and remove unused function\n\n- Change createSerializedPayload to accept std::vector\u003cuint64_t\u003e to\n  match RoaringBitmapArray::addSafe(uint64_t) and avoid narrowing\n  conversion when passing values near UINT64_MAX.\n- Update all local deletedRows vectors and loop counters to uint64_t.\n- Remove unused readUint64LittleEndian() that triggered\n  -Wunused-function.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6\n\n* Fix overflow test: use valid bitmap values with high baseReadOffset\n\nRoaringBitmapArray::addSafe rejects values above kMaxRepresentableValue,\nso the test cannot insert row indices near UINT64_MAX. Instead, test the\noverflow guard by calling applyDeletionFilter with a baseReadOffset near\nUINT64_MAX on a bitmap with low-value entries, verifying no crash and\nempty result.\n\nAssisted-by: GitHub Copilot:claude-opus-4.6"
    },
    {
      "commit": "51a9e96516d4d7414a52027d2019acf46e2b6009",
      "tree": "5efb60f441ce898e7652a5ff9a24090e4244fe89",
      "parents": [
        "d666ceb9c016598e3a1ed4dbe12c365f54248693"
      ],
      "author": {
        "name": "李扬",
        "email": "taiyangli@apache.org",
        "time": "Fri Jul 03 22:51:33 2026 +0800"
      },
      "committer": {
        "name": "GitHub",
        "email": "noreply@github.com",
        "time": "Fri Jul 03 15:51:33 2026 +0100"
      },
      "message": "[GLUTEN][CORE] Expose shuffle reader metrics iterator delegate (#12432)\n\nChange-Id: Ia535725b902e511da1693bd9f7fbaa38f84cc3c8\n\nCo-authored-by: Aime \u003caime@bytedance.com\u003e"
    }
  ],
  "next": "d666ceb9c016598e3a1ed4dbe12c365f54248693"
}
