Skip to content

Array functions evaluate their second argument on rows where the array is NULL, raising errors that Spark skips #6613

Description

@andygrove

Describe the bug

Spark's ArrayContains is a null-intolerant BinaryExpression. When the array is NULL it returns NULL without evaluating the second argument. Comet's native array_contains evaluates both arguments over the whole batch first, because ScalarFunctionExpr evaluates every child before calling the function. So a second argument that throws under ANSI raises on rows Spark never evaluates it for, and a query that Spark completes fails under Comet.

The same happens for array_position, array_remove, arrays_overlap, array_union and map_contains_key, which Spark rewrites to array_contains(map_keys(m), k). element_at is not affected, because CometElementAt wraps the lookup in a CASE WHEN <array> IS NOT NULL guard (#5766). Neither are array_intersect and array_except, which always go through the codegen dispatcher and so run Spark's own code.

Steps to reproduce

SET spark.sql.ansi.enabled=true; -- the default on Spark 4.x
CREATE TABLE t(ai ARRAY<INT>, m MAP<INT, INT>, s STRING) USING parquet;
INSERT INTO t VALUES (NULL, NULL, 'bad'), (array(1), map(1, 10), '1');

SELECT array_contains(ai, CAST(s AS INT)) FROM t;
SELECT array_position(ai, CAST(s AS INT)) FROM t;
SELECT array_remove(ai, CAST(s AS INT)) FROM t;
SELECT arrays_overlap(ai, array(CAST(s AS INT))) FROM t;
SELECT array_union(ai, array(CAST(s AS INT))) FROM t;
SELECT map_contains_key(m, CAST(s AS INT)) FROM t;

Each query runs as CometProject over CometNativeScan and fails with:

org.apache.spark.SparkNumberFormatException: [CAST_INVALID_INPUT] The value 'bad' of the type "STRING" cannot be cast to "INT" because it is malformed. Correct the value as per the syntax, or change its target type. Use `try_cast` to tolerate malformed input and return NULL instead. SQLSTATE: 22018

Reproduced on main at 9866221 with Spark 4.1.3, the default profile, with one CometSqlFileTestSuite fixture per query.

Expected behavior

Spark returns NULL for the first row, and true, 1, [], true, [1] and true for the second. When the bad value sits on a row whose array is not NULL, Spark raises CAST_INVALID_INPUT too, and Comet already matches it there.

Additional context

  • branch-1.1 routes all five the same way for an ARRAY<INT>, so 1.1.0 should be affected too. I have not run it there.
  • For the casts above, only ANSI mode is affected. With ANSI off the cast returns NULL, and try_cast is a workaround. A second argument that fails in any mode fails in both. With a NULL a, array_union(a, slice(b, 0, 1)) returns NULL in Spark, while Comet evaluates the slice anyway and fails. On main at 9dc8c3c that reproduces for an ARRAY<INT> on Spark 4.1.3 and 4.2.0 (Unexpected value for start in function slice), and for an ARRAY<DOUBLE> on 4.2.0, which already runs float elements natively (INVALID_PARAMETER_VALUE.START). It was found in the feat: normalize floats in native array_distinct and array_union #6563 review, which would also run float elements natively on 4.0.5+ and 4.1.4+.
  • On main, array_contains over an array with a float leaf avoids this, because it goes through the codegen dispatcher, which runs Spark's generated code. feat: run array_contains on float elements natively with Spark's equa… #6599 would move those arrays to a native kernel and bring this back for them. The review there covers it.
  • This is one shape of the family tracked in Track evaluation masks for data-dependent errors beyond unbase64 operator fallback #6006. CometElementAt (feat: address remaining issues for CreateArray #5766) and CometArrayJoin.orderInsensitive show two existing fixes: a CASE WHEN guard that masks the lookup, or keeping any second argument other than a literal or column read on the dispatcher.

Activity

  1. added
    bugSomething isn't working
    priority:mediumFunctional bugs, performance regressions, broken features
    on Oct 4, 2026
  2. changed the title [-]Array lookup functions evaluate their second argument on rows where the array is NULL, raising ANSI errors that Spark skips[/-] [+]Array functions evaluate their second argument on rows where the array is NULL, raising errors that Spark skips[/+] on Oct 6, 2026
  3. self-assigned this
    on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:expressionsExpression evaluationbugSomething isn't workingpriority:mediumFunctional bugs, performance regressions, broken features

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions