Skip to content

[Variant] Support whole-value VariantType in native Parquet writes #5433

Description

@peterxcli

Enable ordinary native Parquet writes of direct, top-level whole Variant values. The input Field/scan foundation landed in #5868; writer admission remains to be added.

Implementation

Reuse the protobuf-to-Arrow Field path and parquet::arrow::ArrowWriter. Preserve the parent arrow.parquet.variant extension marker, field name/nullability and [value: Binary, metadata: Binary] order.

The existing Arrow writer mapping already supplies the logical annotation. Emit a Parquet VARIANT(1) group with required binary value and metadata children; no dependency upgrade is needed for that mapping.

Completion

Assert a native scan-to-write plan, the footer annotation and Spark VariantType on readback. Verify exact value/metadata bytes for objects, arrays/scalars, Variant null, SQL NULL and multiple Variant columns with adjacent ordinary columns. Preserve Spark 3 behavior.

This issue covers unshredded top-level output. #3983 owns shredded writes; nested and Iceberg writes remain outside this change. Variant producers from #5425 can later feed the same writer.

Parent: #5438.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions