Move the missing-state checks in rewrite_shredding_state into the native unshredding traversal. #5868 currently checks these states before Arrow reconstruction.
Required behavior
For whole-value reads, “missing” means both value and typed_value are absent or SQL NULL. A present value = X'00' is explicit Variant null.
| State |
Result |
| SQL NULL root |
SQL NULL; ignore its children |
| Valid root or list-element state with a missing value |
MALFORMED_VARIANT |
| Valid object-field state with a missing value |
Omit the field |
| NULL object-field state under a present typed object |
MALFORMED_VARIANT |
| Present residual Variant null, no typed value |
Retain Variant null |
| NULL typed value with a present residual |
Read the residual; ignore inactive typed descendants |
| Active root with NULL metadata |
MALFORMED_VARIANT |
| NULL list-element state struct |
Controlled MALFORMED_VARIANT rejection |
These rules follow Spark's full reconstruction, checked on 4.0.4, 4.1.3 and 4.2.0. The NULL-list-struct case is an explicit exception: Spark can throw an NPE; Comet should retain its controlled error. Pushed field extraction has different validation coverage and is outside this change.
Proposed change
Add MissingValuePolicy::{Parquet, Spark} to an options-based Arrow unshredder. This API is proposed. Existing unshred_variant must retain the Parquet missing-value policy.
Retain the original state StructArray and its context (root, object field or list element) alongside each recursive row builder. In Spark mode, check the table above before decoding. Skip null roots and inactive typed containers; visit only the current list row's offsets, including sliced/ListView arrays.
Apply the same checks to the residual-only fast path and value-only/null builders. A schema without typed_value may preserve its bytes, but unmasked NULL value/metadata entries must still fail.
Once the pinned dependency provides the API, select Spark mode in Comet and remove the missing-state loop and allow_missing argument. Keep residual/metadata rewrites still needed elsewhere, and keep Arrow-error conversion to SparkError::MalformedVariant and JVM MALFORMED_VARIANT in Comet.
Completion
Test both policies, every state above, parent masking and sliced lists. Run native scan regressions and unchanged upstream Spark Variant assertions across supported profiles, with the documented NULL-list-struct exception. The checks must execute during reconstruction and the duplicate Comet precheck must be deleted.
Parent: #5477. Related: #5978 (output encoding), #5979 (metadata), #5980 (precedence). Removing these checks alone does not remove the surrounding rewrite pass.
Move the missing-state checks in
rewrite_shredding_stateinto the native unshredding traversal. #5868 currently checks these states before Arrow reconstruction.Required behavior
For whole-value reads, “missing” means both
valueandtyped_valueare absent or SQL NULL. A presentvalue = X'00'is explicit Variant null.MALFORMED_VARIANTMALFORMED_VARIANTMALFORMED_VARIANTMALFORMED_VARIANTrejectionThese rules follow Spark's full reconstruction, checked on 4.0.4, 4.1.3 and 4.2.0. The NULL-list-struct case is an explicit exception: Spark can throw an NPE; Comet should retain its controlled error. Pushed field extraction has different validation coverage and is outside this change.
Proposed change
Add
MissingValuePolicy::{Parquet, Spark}to an options-based Arrow unshredder. This API is proposed. Existingunshred_variantmust retain the Parquet missing-value policy.Retain the original state StructArray and its context (root, object field or list element) alongside each recursive row builder. In Spark mode, check the table above before decoding. Skip null roots and inactive typed containers; visit only the current list row's offsets, including sliced/ListView arrays.
Apply the same checks to the residual-only fast path and value-only/null builders. A schema without
typed_valuemay preserve its bytes, but unmasked NULL value/metadata entries must still fail.Once the pinned dependency provides the API, select Spark mode in Comet and remove the missing-state loop and
allow_missingargument. Keep residual/metadata rewrites still needed elsewhere, and keep Arrow-error conversion toSparkError::MalformedVariantand JVMMALFORMED_VARIANTin Comet.Completion
Test both policies, every state above, parent masking and sliced lists. Run native scan regressions and unchanged upstream Spark Variant assertions across supported profiles, with the documented NULL-list-struct exception. The checks must execute during reconstruction and the duplicate Comet precheck must be deleted.
Parent: #5477. Related: #5978 (output encoding), #5979 (metadata), #5980 (precedence). Removing these checks alone does not remove the surrounding rewrite pass.