Describe the bug
A cast whose target struct type repeats a field name fails the task when Comet runs it natively. Spark allows the cast as long as each pair of fields casts, so CAST(s AS STRUCT<x: INT, x: INT>) over a STRUCT<p: INT, q: INT> column returns {1, 10}. Comet's native cast_struct_to_struct builds the result with the target fields, and the projection itself works: to_json over the cast matches Spark with {"x":1,"x":10}. Importing the result on the JVM then fails, because Java Arrow keys struct children by name and the two x children collapse into one:
java.lang.IllegalStateException: ArrowArray struct has 2 children (expected 1)
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
at org.apache.arrow.c.ArrayImporter.importChild(ArrayImporter.java:83)
at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:101)
at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
at org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:236)
The same happens when the struct is nested in an array or a map value.
Steps to reproduce
Default configs, reproduced on main at fef94f6cd with Spark 4.1.3:
CREATE TABLE ts (
s STRUCT<p: INT, q: INT>,
arr ARRAY<STRUCT<p: INT, q: INT>>,
m MAP<STRING, STRUCT<p: INT, q: INT>>) USING parquet;
INSERT INTO ts VALUES
(named_struct('p', 1, 'q', 10), array(named_struct('p', 1, 'q', 10)),
map('k', named_struct('p', 1, 'q', 10))),
(named_struct('p', 2, 'q', NULL), array(), map()),
(NULL, NULL, NULL);
SELECT CAST(s AS STRUCT<x: INT, x: INT>) FROM ts; -- IllegalStateException
SELECT CAST(arr AS ARRAY<STRUCT<x: INT, x: INT>>) FROM ts; -- IllegalStateException
SELECT CAST(m AS MAP<STRING, STRUCT<x: INT, x: INT>>) FROM ts; -- IllegalStateException
SELECT CAST(s AS STRUCT<x: INT, y: INT>) FROM ts; -- matches Spark
SELECT to_json(CAST(s AS STRUCT<x: INT, x: INT>)) FROM ts; -- matches Spark
The plan is CometProject over CometNativeScan. Spark returns {1, 10}, {2, null} and null for the first query.
Expected behavior
Either match Spark or fall back. The smallest fix is probably for the struct branch of CometCast.isSupported to return Unsupported when the target struct repeats a field name, the same way CometCreateNamedStruct declines repeated names. The array and map branches already recurse into the struct branch, so one check there also covers the nested cases.
Additional context
This is the same Java Arrow limitation as #6251 (arrays_zip, fix in #6324), #5605 and #1015. Main declines structs with repeated field names wherever they would cross into Java Arrow: the shuffle and row conversion checks (#5866), the codegen dispatcher (#5766) and the cache serializer (#6004). Those checks guard the paths that bring a struct into Comet, though, and a native cast creates one. The broadcast exchange admits these types too (CometSink.convert calls supportedDataType with the default allowDuplicateStructFieldNames = true), so declining in the serde that produces the struct is what keeps it out.
Found while checking whether anything in #5603 was worth keeping.
Describe the bug
A cast whose target struct type repeats a field name fails the task when Comet runs it natively. Spark allows the cast as long as each pair of fields casts, so
CAST(s AS STRUCT<x: INT, x: INT>)over aSTRUCT<p: INT, q: INT>column returns{1, 10}. Comet's nativecast_struct_to_structbuilds the result with the target fields, and the projection itself works:to_jsonover the cast matches Spark with{"x":1,"x":10}. Importing the result on the JVM then fails, because Java Arrow keys struct children by name and the twoxchildren collapse into one:The same happens when the struct is nested in an array or a map value.
Steps to reproduce
Default configs, reproduced on
mainatfef94f6cdwith Spark 4.1.3:The plan is
CometProjectoverCometNativeScan. Spark returns{1, 10},{2, null}andnullfor the first query.Expected behavior
Either match Spark or fall back. The smallest fix is probably for the struct branch of
CometCast.isSupportedto returnUnsupportedwhen the target struct repeats a field name, the same wayCometCreateNamedStructdeclines repeated names. The array and map branches already recurse into the struct branch, so one check there also covers the nested cases.Additional context
This is the same Java Arrow limitation as #6251 (
arrays_zip, fix in #6324), #5605 and #1015. Main declines structs with repeated field names wherever they would cross into Java Arrow: the shuffle and row conversion checks (#5866), the codegen dispatcher (#5766) and the cache serializer (#6004). Those checks guard the paths that bring a struct into Comet, though, and a native cast creates one. The broadcast exchange admits these types too (CometSink.convertcallssupportedDataTypewith the defaultallowDuplicateStructFieldNames = true), so declining in the serde that produces the struct is what keeps it out.Found while checking whether anything in #5603 was worth keeping.