You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Array functions evaluate their second argument on rows where the array is NULL, raising errors that Spark skips #6613
Spark's ArrayContains is a null-intolerant BinaryExpression. When the array is NULL it returns NULL without evaluating the second argument. Comet's native array_contains evaluates both arguments over the whole batch first, because ScalarFunctionExpr evaluates every child before calling the function. So a second argument that throws under ANSI raises on rows Spark never evaluates it for, and a query that Spark completes fails under Comet.
The same happens for array_position, array_remove, arrays_overlap, array_union and map_contains_key, which Spark rewrites to array_contains(map_keys(m), k). element_at is not affected, because CometElementAt wraps the lookup in a CASE WHEN <array> IS NOT NULL guard (#5766). Neither are array_intersect and array_except, which always go through the codegen dispatcher and so run Spark's own code.
Steps to reproduce
SETspark.sql.ansi.enabled=true; -- the default on Spark 4.xCREATETABLEt(ai ARRAY<INT>, m MAP<INT, INT>, s STRING) USING parquet;
INSERT INTO t VALUES (NULL, NULL, 'bad'), (array(1), map(1, 10), '1');
SELECT array_contains(ai, CAST(s ASINT)) FROM t;
SELECT array_position(ai, CAST(s ASINT)) FROM t;
SELECT array_remove(ai, CAST(s ASINT)) FROM t;
SELECT arrays_overlap(ai, array(CAST(s ASINT))) FROM t;
SELECT array_union(ai, array(CAST(s ASINT))) FROM t;
SELECT map_contains_key(m, CAST(s ASINT)) FROM t;
Each query runs as CometProject over CometNativeScan and fails with:
org.apache.spark.SparkNumberFormatException: [CAST_INVALID_INPUT] The value 'bad' of the type "STRING" cannot be cast to "INT" because it is malformed. Correct the value as per the syntax, or change its target type. Use `try_cast` to tolerate malformed input and return NULL instead. SQLSTATE: 22018
Reproduced on main at 9866221 with Spark 4.1.3, the default profile, with one CometSqlFileTestSuite fixture per query.
Expected behavior
Spark returns NULL for the first row, and true, 1, [], true, [1] and true for the second. When the bad value sits on a row whose array is not NULL, Spark raises CAST_INVALID_INPUT too, and Comet already matches it there.
Additional context
branch-1.1 routes all five the same way for an ARRAY<INT>, so 1.1.0 should be affected too. I have not run it there.
For the casts above, only ANSI mode is affected. With ANSI off the cast returns NULL, and try_cast is a workaround. A second argument that fails in any mode fails in both. With a NULL a, array_union(a, slice(b, 0, 1)) returns NULL in Spark, while Comet evaluates the slice anyway and fails. On main at 9dc8c3c that reproduces for an ARRAY<INT> on Spark 4.1.3 and 4.2.0 (Unexpected value for start in function slice), and for an ARRAY<DOUBLE> on 4.2.0, which already runs float elements natively (INVALID_PARAMETER_VALUE.START). It was found in the feat: normalize floats in native array_distinct and array_union #6563 review, which would also run float elements natively on 4.0.5+ and 4.1.4+.
On main, array_contains over an array with a float leaf avoids this, because it goes through the codegen dispatcher, which runs Spark's generated code. feat: run array_contains on float elements natively with Spark's equa… #6599 would move those arrays to a native kernel and bring this back for them. The review there covers it.
changed the title [-]Array lookup functions evaluate their second argument on rows where the array is NULL, raising ANSI errors that Spark skips[/-][+]Array functions evaluate their second argument on rows where the array is NULL, raising errors that Spark skips[/+]on Oct 6, 2026
Describe the bug
Spark's
ArrayContainsis a null-intolerantBinaryExpression. When the array is NULL it returns NULL without evaluating the second argument. Comet's nativearray_containsevaluates both arguments over the whole batch first, becauseScalarFunctionExprevaluates every child before calling the function. So a second argument that throws under ANSI raises on rows Spark never evaluates it for, and a query that Spark completes fails under Comet.The same happens for
array_position,array_remove,arrays_overlap,array_unionandmap_contains_key, which Spark rewrites toarray_contains(map_keys(m), k).element_atis not affected, becauseCometElementAtwraps the lookup in aCASE WHEN <array> IS NOT NULLguard (#5766). Neither arearray_intersectandarray_except, which always go through the codegen dispatcher and so run Spark's own code.Steps to reproduce
Each query runs as
CometProjectoverCometNativeScanand fails with:Reproduced on
mainat 9866221 with Spark 4.1.3, the default profile, with oneCometSqlFileTestSuitefixture per query.Expected behavior
Spark returns NULL for the first row, and
true,1,[],true,[1]andtruefor the second. When the bad value sits on a row whose array is not NULL, Spark raisesCAST_INVALID_INPUTtoo, and Comet already matches it there.Additional context
branch-1.1routes all five the same way for anARRAY<INT>, so 1.1.0 should be affected too. I have not run it there.try_castis a workaround. A second argument that fails in any mode fails in both. With a NULLa,array_union(a, slice(b, 0, 1))returns NULL in Spark, while Comet evaluates thesliceanyway and fails. Onmainat 9dc8c3c that reproduces for anARRAY<INT>on Spark 4.1.3 and 4.2.0 (Unexpected value for start in function slice), and for anARRAY<DOUBLE>on 4.2.0, which already runs float elements natively (INVALID_PARAMETER_VALUE.START). It was found in the feat: normalize floats in native array_distinct and array_union #6563 review, which would also run float elements natively on 4.0.5+ and 4.1.4+.main,array_containsover an array with a float leaf avoids this, because it goes through the codegen dispatcher, which runs Spark's generated code. feat: run array_contains on float elements natively with Spark's equa… #6599 would move those arrays to a native kernel and bring this back for them. The review there covers it.CometElementAt(feat: address remaining issues forCreateArray#5766) andCometArrayJoin.orderInsensitiveshow two existing fixes: aCASE WHENguard that masks the lookup, or keeping any second argument other than a literal or column read on the dispatcher.