[SPARK-55569][SQL] Collapse non-matching pivot column values in fast-path firstAgg - #59149
Open
shrirangmhalgi wants to merge 1 commit into
Open
shrirangmhalgi wants to merge 1 commit into
shrirangmhalgi wants to merge 1 commit into
Conversation
…path firstAgg In the PivotFirst fast path, the firstAgg groups by (groupByKeys, pivotColumn). When the table has many distinct values outside the explicit pivot IN list, firstAgg creates one group per distinct non-matching value -- all of which PivotFirst ignores in secondAgg. This change wraps the pivot column group-by key in IF(pivotColumn IN (pivotValues), pivotColumn, null), collapsing all non-matching values into a single null group. PivotFirst skips the null group as before, but the group-by keys are preserved in secondAgg, keeping semantics identical to the un-optimized plan. The optimization is skipped when: - The pivot column contains an aggregate expression (SPARK-24722) - A NULL pivot value is present (non-matching rows would merge with legitimate null-key rows) Co-authored-by: Claude Opus 4.8
shrirangmhalgi
commented
Sep 30, 2026
Contributor
Author
There was a problem hiding this comment.
@cloud-fan / @peter-toth could you please take a look? This optimizes the PivotFirst fast path by collapsing non-matching pivot column values into a single null group in firstAgg, reducing the number of wasted groups from O(distinct_values) to O(1) per group-by key. The query results are unchanged.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
In the PivotFirst fast path,
firstAgggroups by(groupByKeys, pivotColumn). When the table has many distinct pivot column values outside the explicitINlist,firstAggcreates one group per distinct non-matching value - all of whichPivotFirstignores insecondAgg.This PR wraps the pivot column group-by key in
IF(pivotColumn IN (pivotValues), pivotColumn, null), collapsing all non-matching values into a single null group.PivotFirstskips the null group as before, and the group-by keys are preserved insecondAgg, keeping query results identical to the un-optimized plan.The optimization is skipped when:
Why are the changes needed?
For a table with 10,000 distinct course values pivoting on 2:
firstAggcreates 10,000 groups per group-by key (99.98% wasted)firstAggcreates 3 groups per group-by key (2 matching + 1 null)This reduces memory (fewer HashAggregate entries), CPU (fewer hash/compare/aggregate operations), and spill risk.
Does this PR introduce any user-facing change?
No. Query results are identical. This only affects the internal execution plan for the PivotFirst fast path.
How was this patch tested?
Five new tests in
DataFramePivotSuite:If(In(...), col, null)appears in the analyzed planRow(year, null, null)Existing golden file tests regenerated (
pivot.sql,udf-pivot.sql,pipe-operators.sql). 49/49DataFramePivotSuitetests pass.Was this patch authored or co-authored using generative AI tooling?
Yes. Co-authored using Claude Opus 4.8.