Remove output-specific UTF-16 Variant ordering once every supported Spark profile includes SPARK-58949 or equivalent canonical UTF-8 write/search and legacy-read support.
The implementation merged in #5868 must currently match Spark 4.0.4, 4.1.3 and experimental 4.2.0. Canonical Parquet/Arrow uses UTF-8 ordering; legacy Spark objects use UTF-16 ordering. Historical files will still require compatible reads after Spark is upgraded.
Change after the Spark prerequisite
Use canonical UTF-8 comparison for reconstructed output and remove only the output-specific UTF-16 traversal/sort helpers. Preserve scalar encoding, dictionary traversal, metadata flags and residual bytes; #5978 owns the complete reconstruction replacement.
Retain input-side handling for legacy residual objects until the reader accepts them directly. Helpers shared with that path cannot be deleted merely because output ordering changes.
Completion
Read canonical and legacy files through native scans, including nested/partially shredded objects and Unicode keys. Exercise Spark lookup above its binary-search threshold. Verify the redundant output rewrite is removed, preserve Spark 3 behavior and report before/after whole-value scan timings.
SPARK-58949 landed on Spark master and branch-4.x; it still needs adoption in Comet's supported profiles. apache/parquet-java#3746 resolved the corresponding library issue but does not update Spark's pinned builder.
Parent: #5477; roadmap: #5438.
Remove output-specific UTF-16 Variant ordering once every supported Spark profile includes SPARK-58949 or equivalent canonical UTF-8 write/search and legacy-read support.
The implementation merged in #5868 must currently match Spark 4.0.4, 4.1.3 and experimental 4.2.0. Canonical Parquet/Arrow uses UTF-8 ordering; legacy Spark objects use UTF-16 ordering. Historical files will still require compatible reads after Spark is upgraded.
Change after the Spark prerequisite
Use canonical UTF-8 comparison for reconstructed output and remove only the output-specific UTF-16 traversal/sort helpers. Preserve scalar encoding, dictionary traversal, metadata flags and residual bytes; #5978 owns the complete reconstruction replacement.
Retain input-side handling for legacy residual objects until the reader accepts them directly. Helpers shared with that path cannot be deleted merely because output ordering changes.
Completion
Read canonical and legacy files through native scans, including nested/partially shredded objects and Unicode keys. Exercise Spark lookup above its binary-search threshold. Verify the redundant output rewrite is removed, preserve Spark 3 behavior and report before/after whole-value scan timings.
SPARK-58949 landed on Spark master and branch-4.x; it still needs adoption in Comet's supported profiles. apache/parquet-java#3746 resolved the corresponding library issue but does not update Spark's pinned builder.
Parent: #5477; roadmap: #5438.