Skip to content

[EPIC] DataFusion 56 upgrade: regressions found by tracking DataFusion main #6410

Description

@andygrove

What is the problem the feature request solves?

Draft PR #6404 builds Comet against DataFusion main instead of a release. That way, upstream changes reach Comet's CI weeks before DataFusion 56.0.0 is released (targeted for Oct/Nov, apache/datafusion#24461). The PR is long-lived and bumped weekly. It is not meant to merge before that release. It currently pins these versions:

Dependency Version
DataFusion 4a5d580 (main, 2026-09-29)
arrow, parquet 60.0.0
object_store 0.14.2
iceberg-rust 0d5eb12, the head of apache/iceberg-rust#3257 (Arrow 60, not merged yet)

This issue tracks the regressions that PR finds. The goal is to fix them upstream, or work around them in Comet, before the release rather than after it.

Describe the potential solution

Regressions

Check an item off once #6404 pins a revision that has the fix.

  • The native Iceberg writer puts NaN bounds in manifests.
    • CometIcebergWriteActionSuite "native acceleration: NaN float/double manifest metrics match the JVM writer" fails (run). The native writer produces [1,0.0,NaN,2,0.25,NaN,3,NaN,NaN] where the JVM writer produces [1,0.0,2.5,2,0.25,1.5,3,null,null].
    • Cause: Parquet 60 now writes NaN min/max statistics for a column chunk whose values are all NaN, where 59 wrote none (get_min_max in parquet/src/column/writer/encoder.rs). iceberg-rust's MinMaxColAggregator::update copies the Parquet statistics into lower and upper bounds without filtering out NaN. So a data file with an all-NaN float or double column gets NaN bounds, where iceberg-java writes none.
    • Impact: iceberg-java readers treat a NaN bound as unreliable, so this should not cause wrong results there. But the manifests no longer match the JVM writer's.
    • The fix belongs in iceberg-rust, ideally as part of its Arrow 60 bump (Bump to Arrow 60 iceberg-rust#3257). Not reported upstream yet.

Behavior changes reviewed and accepted

  • Arrow 60 accepts empty entries in unsorted Variant dictionaries (Add proptest fuzzing to parquet-variant and implement fixes for findings arrow-rs#10352). Comet's empty-key workaround used to reject two malformed dictionaries. Both now pass through unchanged, as they do in Spark.
  • A failed read from a JVM input stream now reports Cannot get next batch from input stream. Error code: N. Producer error: .... This is because Arrow 60 made the C Stream callbacks private, so Comet now uses the stock ArrowArrayStreamReader.
  • An invalid shuffle IPC schema message now returns a decode error instead of panicking (try_fb_to_schema).
  • DataFusion's new floor/ceil return types and map_extract absent-key results don't reach Comet, because Comet registers its own implementations.
  • DataFusion now treats a missing Parquet null count as unknown rather than zero. That is a correctness fix. It can reduce pruning on files written by parquet-rs versions before 53.1.

Upstream follow-ups

Tested revisions

Date DataFusion iceberg-rust Suites Result
2026-09-29 4a5d580 0d5eb12 PR tier (run) 1 failure (the NaN bounds bug)

Additional context

Steps for triaging a red run on #6404:

  1. Re-run it once and compare with main's latest nightly, since some failures are known flakes.
  2. Run the failing test on the merge commit with the previous upstream revision. This tells an upstream change apart from a Comet change on main.
  3. Bisect the upstream commits between the two revisions, using a local [patch] override. DataFusion main gets 100 to 120 commits a week, so this takes about 7 steps.

Spark's SQL tests, Iceberg's Spark tests and the non-default Spark profiles have not run on #6404 yet. They will run once the run-spark-4.1-tests, run-iceberg-tests and run-all-spark-profiles labels are applied.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions