Track removal of compatibility code introduced by #5868. Parent roadmap: #5438. Keep completed prerequisites checked and their remaining Comet integration steps visible.
As of 2026-10-02, Comet main resolves Arrow/Parquet 59.3.0 and DataFusion 55.1.0. Dependency upgrades remain deferred. An upstream merge completes its prerequisite; the workaround stays open until Comet adopts the fix and passes the existing regressions.
Completed milestones
Upstream-dependent cleanup — deferred
For the schema-provider step, derive types from the physical Parquet schema, ignore advisory ARROW:schema, map ENUM to Utf8 and retain raw BINARY. Preserve metadata/index/decryptor state and unread-column pruning. Reassess the encrypted-scan fallback after encryption tests pass.
Spark compatibility and performance
These child issues retain the concrete algorithms and acceptance tests.
Compatibility to retain
Unsigned typed_value is excluded by the canonical specification (apache/arrow#50810); apache/arrow-rs#10417 closed without merging. Keep Spark's unsigned widening, including UInt64 -> Decimal(20,0). Millisecond timestamps, unannotated fixed-length binary and FixedSizeList conversion also remain Spark compatibility work.
Field metadata/FFI export, ordinary Binary children in [value, metadata] order, parent nulls, fallback gates and embedded-NUL handling at the C ABI remain integration requirements.
Complete this tracker when each temporary workaround is removed or has a documented retention reason. Each removal must pass its existing native and Spark regressions.
Track removal of compatibility code introduced by #5868. Parent roadmap: #5438. Keep completed prerequisites checked and their remaining Comet integration steps visible.
As of 2026-10-02, Comet main resolves Arrow/Parquet 59.3.0 and DataFusion 55.1.0. Dependency upgrades remain deferred. An upstream merge completes its prerequisite; the workaround stays open until Comet adopts the fix and passes the existing regressions.
Completed milestones
Upstream-dependent cleanup — deferred
metadata.value/typed_valuenormalization.proptestfuzzing to parquet-variant and implement fixes for findings arrow-rs#10352.canonicalize_spark_empty_key_metadataand the retry, and retain fallible full validation and empty-key regressions.Decimal256(p,s) -> Decimal128(p,s)whenp <= 38; verify positive/negative values in 17- and 32-byte physical fields.ParquetFileSchemaProvider, supply the full per-file schema and remove footer reconstruction.For the schema-provider step, derive types from the physical Parquet schema, ignore advisory
ARROW:schema, map ENUM to Utf8 and retain raw BINARY. Preserve metadata/index/decryptor state and unread-column pruning. Reassess the encrypted-scan fallback after encryption tests pass.Spark compatibility and performance
extend_shredded_metadata.These child issues retain the concrete algorithms and acceptance tests.
Compatibility to retain
Unsigned
typed_valueis excluded by the canonical specification (apache/arrow#50810); apache/arrow-rs#10417 closed without merging. Keep Spark's unsigned widening, includingUInt64 -> Decimal(20,0). Millisecond timestamps, unannotated fixed-length binary and FixedSizeList conversion also remain Spark compatibility work.Field metadata/FFI export, ordinary Binary children in
[value, metadata]order, parent nulls, fallback gates and embedded-NUL handling at the C ABI remain integration requirements.Complete this tracker when each temporary workaround is removed or has a documented retention reason. Each removal must pass its existing native and Spark regressions.