Skip to content

[Variant] Consolidate Spark-compatible missing-value validation and errors #5977

Description

@peterxcli

Move the missing-state checks in rewrite_shredding_state into the native unshredding traversal. #5868 currently checks these states before Arrow reconstruction.

Required behavior

For whole-value reads, “missing” means both value and typed_value are absent or SQL NULL. A present value = X'00' is explicit Variant null.

State Result
SQL NULL root SQL NULL; ignore its children
Valid root or list-element state with a missing value MALFORMED_VARIANT
Valid object-field state with a missing value Omit the field
NULL object-field state under a present typed object MALFORMED_VARIANT
Present residual Variant null, no typed value Retain Variant null
NULL typed value with a present residual Read the residual; ignore inactive typed descendants
Active root with NULL metadata MALFORMED_VARIANT
NULL list-element state struct Controlled MALFORMED_VARIANT rejection

These rules follow Spark's full reconstruction, checked on 4.0.4, 4.1.3 and 4.2.0. The NULL-list-struct case is an explicit exception: Spark can throw an NPE; Comet should retain its controlled error. Pushed field extraction has different validation coverage and is outside this change.

Proposed change

Add MissingValuePolicy::{Parquet, Spark} to an options-based Arrow unshredder. This API is proposed. Existing unshred_variant must retain the Parquet missing-value policy.

Retain the original state StructArray and its context (root, object field or list element) alongside each recursive row builder. In Spark mode, check the table above before decoding. Skip null roots and inactive typed containers; visit only the current list row's offsets, including sliced/ListView arrays.

Apply the same checks to the residual-only fast path and value-only/null builders. A schema without typed_value may preserve its bytes, but unmasked NULL value/metadata entries must still fail.

Once the pinned dependency provides the API, select Spark mode in Comet and remove the missing-state loop and allow_missing argument. Keep residual/metadata rewrites still needed elsewhere, and keep Arrow-error conversion to SparkError::MalformedVariant and JVM MALFORMED_VARIANT in Comet.

Completion

Test both policies, every state above, parent masking and sliced lists. Run native scan regressions and unchanged upstream Spark Variant assertions across supported profiles, with the documented NULL-list-struct exception. The checks must execute during reconstruction and the duplicate Comet precheck must be deleted.

Parent: #5477. Related: #5978 (output encoding), #5979 (metadata), #5980 (precedence). Removing these checks alone does not remove the surrounding rewrite pass.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:scanParquet scan / data readingenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions