Skip to content

[Variant] Support VariantType in native columnar-to-row conversion #5436

Description

@peterxcli

Add native columnar-to-row conversion for top-level Variant columns from the foundation merged in #5868.

Spark's UnsafeWriter encodes Variant as 4-byte value length + value bytes + metadata bytes. Generic Struct encoding would produce an incompatible UnsafeRow.

Implementation

Preserve Variant Field identity when initializing native C2R; the physical Struct datatype alone is insufficient. Add a dedicated writer consuming canonical [value: Binary, metadata: Binary] storage and emitting Spark's payload with its row offset/size and alignment conventions.

Preserve SQL NULL versus Variant null and reject malformed child/null combinations without panicking. Admit only the implemented top-level shape.

Completion

Round-trip through UnsafeRow.getVariant with exact value/metadata bytes and assert native C2R. Cover objects, arrays/scalars, both null forms, empty batches and columns before/after Variant. Retain nested fallback and Spark 3 behavior.

Parent: #5438. #5434 reuses this row contract for JVM shuffle; #5425 supplies computed Variant inputs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:ffiArrow FFI / JNI boundaryenhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions