Conversation
Co-authored-by: Cursor <cursoragent@cursor.com>
Keep the review focused on Spark-owned worker routing by dropping the unrelated formatting churn. Co-authored-by: Cursor <cursoragent@cursor.com>
The class-name hook never shipped, so a compatibility check is unnecessary. Co-authored-by: Cursor <cursoragent@cursor.com>
…class loader The thread context class loader includes user artifacts, so a user jar could supply the Spark-owned factory class name while spark-udf-worker-grpc is absent from the engine classpath. Co-authored-by: Cursor <cursoragent@cursor.com>
…end through a direct gRPC worker Runs a scalar external UDF query through the built-in DIRECT dispatcher, which spawns the gRPC echo worker over a Unix domain socket, and asserts the returned rows. Replaces the test that asserted the gRPC runtime was missing from the sql/core test classpath. Co-authored-by: Cursor <cursoragent@cursor.com>
Member
|
cc @haiyangsun-db FYI |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on #59153. The commits before
00ece1b3are that pull request. Please review00ece1b3only.What changes were proposed in this pull request?
PlanExternalUDFsSuiteruns a scalar external UDF through the realDIRECTpath.ExecuteExternalUDFExecasks SparkEnv for the built-in dispatcher.DirectGrpcDispatcherstartsEchoGrpcWorkerMainand talks to it over a Unix domain socket. The echo worker returns each Arrow batch unchanged, so a bigint column is an identity function.The test input is
Range(0, 6)in two partitions. The collected rows are(0, 0)through(5, 5).sql/coregains a test-scoped dependency onspark-udf-worker-grpcand its test jar, so the reflective load ofDirectDispatcherFactorysucceeds in this suite.corestill does not depend on gRPC.This replaces the suite's missing-runtime assertion.
SparkEnvUDFDispatcherSuitestill covers that error on thecoreclasspath, which does not contain the gRPC module.Why are the changes needed?
#59153 routes
DIRECTto the built-in factory, and the SQL node can already stream Arrow batches. Nothing insql/corechecked that a query returns rows from a spawned worker. This test is that check. Shading and a Python worker stay in SPARK-57214 and the Python worker tickets.Does this PR introduce any user-facing change?
No. The new dependency is test scope. A packaged Spark still reports the missing runtime until SPARK-57214.
How was this patch tested?
sql/testOnly org.apache.spark.sql.execution.externalUDF.PlanExternalUDFsSuiteon macOS. 32 tests passed, including this one. The log shows one kqueue domain socket per task.sql/Test/scalastylereported 0 errors.Was this patch authored or co-authored using generative AI tooling?
Generated-by: Cursor Grok 4.7
Made with Cursor