CometNativeUdfSuite and the native adapter tests cover every supported Spark type with nulls,
one- and two-argument calls, literal arguments in each position, and each operator a UDF can be
placed in. What they do not cover yet:
- Empty batches. A zero-row batch reaches
invoke_with_args with number_rows = 0. Nothing
checks that the kernel sees zero-length arrays and that a zero-length result passes the row count
check.
- Sliced arrays. An array with a non-zero offset (as produced by a
LIMIT or a filter over a
batch) crosses the C Data Interface with its offset, and the kernel has to honor it. Nothing
exercises a non-zero offset in either direction.
- Dictionary-encoded input. A dictionary-encoded string column from a Parquet scan reaches
the UDF as whatever Comet's scan hands the projection. The type the kernel sees then depends on
whether the dictionary was unpacked first, and return_field implementations that match on
DataType::Utf8 would reject it.
Follow-up to #4459.
CometNativeUdfSuiteand the native adapter tests cover every supported Spark type with nulls,one- and two-argument calls, literal arguments in each position, and each operator a UDF can be
placed in. What they do not cover yet:
invoke_with_argswithnumber_rows = 0. Nothingchecks that the kernel sees zero-length arrays and that a zero-length result passes the row count
check.
LIMITor a filter over abatch) crosses the C Data Interface with its offset, and the kernel has to honor it. Nothing
exercises a non-zero offset in either direction.
the UDF as whatever Comet's scan hands the projection. The type the kernel sees then depends on
whether the dictionary was unpacked first, and
return_fieldimplementations that match onDataType::Utf8would reject it.Follow-up to #4459.