Skip to content

[Variant] Support Variant payloads through Comet shuffle and spill paths #5434

Description

@peterxcli

Carry top-level Variant columns as payloads through Comet shuffle and spill. For example, repartition by an ordinary id while retaining id, v. The columnar foundation landed in #5868.

Implementation

For native shuffle, preserve the parent arrow.parquet.variant Field marker, [value, metadata] Binary children and parent nulls through IPC, spill and merge.

For JVM shuffle, decode Spark's dedicated UnsafeRow Variant payload into the canonical Arrow Field. Coordinate the row codec with #5436.

Keep Spark's restriction on Variant partitioning expressions. Admit supported payload shapes without enabling Variant keys or nested shapes.

Completion

Compare results and assert native/JVM Comet shuffle plans for non-Variant hash keys, round-robin, single partition, AQE/coalescing and forced spill. Cover containers/scalars, Variant null, SQL NULL, multiple Variant columns and adjacent ordinary columns. Check Field metadata and layout after spill/readback, and retain Spark-compatible rejection of Variant keys.

Parent: #5438. Native expression producers are tracked by #5425.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:shuffleShuffle (JVM and native)enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions