The library path passed to CometNativeUDF.register must already be valid on every executor.
Comet does not ship the file anywhere; users have to bake it into their image, mount it, or stage
it with their own tooling, and a path that exists only on the driver fails at execution time.
Spark already has a mechanism for this: SparkContext.addFile / spark.files, with
SparkFiles.get(name) resolving the local copy on each executor. register could accept a library
added that way (or add it itself) and send the file name rather than an absolute path, with the
planner resolving it through SparkFiles on the executor side.
Points to settle:
- Cache keying. The loaded-library cache is keyed by path and never unloads. Two versions of a
library added under the same name must not collide, so the key probably needs a content hash.
- Loading from
SparkFiles directories, which differ per application, interacts with the "never
copy over a loaded library" rule in the user guide.
Follow-up to #4459.
The library path passed to
CometNativeUDF.registermust already be valid on every executor.Comet does not ship the file anywhere; users have to bake it into their image, mount it, or stage
it with their own tooling, and a path that exists only on the driver fails at execution time.
Spark already has a mechanism for this:
SparkContext.addFile/spark.files, withSparkFiles.get(name)resolving the local copy on each executor.registercould accept a libraryadded that way (or add it itself) and send the file name rather than an absolute path, with the
planner resolving it through
SparkFileson the executor side.Points to settle:
library added under the same name must not collide, so the key probably needs a content hash.
SparkFilesdirectories, which differ per application, interacts with the "nevercopy over a loaded library" rule in the user guide.
Follow-up to #4459.