Skip to content

[Variant] Remove Variant UTF-16 output rewriting #5474

Description

@peterxcli

Remove output-specific UTF-16 Variant ordering once every supported Spark profile includes SPARK-58949 or equivalent canonical UTF-8 write/search and legacy-read support.

The implementation merged in #5868 must currently match Spark 4.0.4, 4.1.3 and experimental 4.2.0. Canonical Parquet/Arrow uses UTF-8 ordering; legacy Spark objects use UTF-16 ordering. Historical files will still require compatible reads after Spark is upgraded.

Change after the Spark prerequisite

Use canonical UTF-8 comparison for reconstructed output and remove only the output-specific UTF-16 traversal/sort helpers. Preserve scalar encoding, dictionary traversal, metadata flags and residual bytes; #5978 owns the complete reconstruction replacement.

Retain input-side handling for legacy residual objects until the reader accepts them directly. Helpers shared with that path cannot be deleted merely because output ordering changes.

Completion

Read canonical and legacy files through native scans, including nested/partially shredded objects and Unicode keys. Exercise Spark lookup above its binary-search threshold. Verify the redundant output rewrite is removed, preserve Spark 3 behavior and report before/after whole-value scan timings.

SPARK-58949 landed on Spark master and branch-4.x; it still needs adoption in Comet's supported profiles. apache/parquet-java#3746 resolved the corresponding library issue but does not update Spark's pinned builder.

Parent: #5477; roadmap: #5438.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:scanParquet scan / data readingenhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions