Skip to content

Native Iceberg writer's value and null counts for a float under a NULL struct differ from iceberg-java 1.9+ #6659

Description

@andygrove

Describe the bug

For a float or double field inside a struct, the native Iceberg writer's value_counts and null_value_counts differ from iceberg-java's on Iceberg 1.9 and later whenever the struct is NULL in some rows. iceberg-java counts only the rows where the struct is present. The native path counts every row, and counts the NULL-struct rows as nulls.

With the reproduction below on Spark 4.1 / Iceberg 1.11 (100 rows, s NULL on the 50 odd ids), the table's data files add up to a value_count of 100 and a null_value_count of 50 for s.x natively, against 50 and 0 through iceberg-java. #6562 saw the same numbers.

The native path takes these counts from the Parquet footer. IcebergReflection.buildFloatFieldMetrics builds one FieldMetrics per float/double leaf for ParquetUtil.footerMetrics, with the value and null counts copied from the native DataFile, which carries the footer's counts. Its doc assumes they are "footer-derived on both sides". That held through Iceberg 1.8, whose footerMetrics takes value and null counts from the footer and uses FieldMetrics only for NaN counts and bounds. From 1.9, footerMetrics delegates to ParquetMetrics, whose metricsFromFieldMetrics takes the value, null and NaN counts from the FieldMetrics whenever one exists for the field. The JVM writer's FieldMetrics come from its value writers, and ParquetValueWriters.OptionWriter writes a NULL struct's columns with column.writeNull directly, so the field's own writer never sees those rows.

Fields under a list or map are not affected, because ParquetMetrics keeps no metrics for them.

The native counts never prune more than iceberg-java's. They report at least as many nulls, so IS NULL prunes no more files, and "only nulls", "only NaNs" and "nulls or NaNs" come out the same or less often. So this is a metadata parity difference, visible in the files and data_files metadata tables, rather than a wrong query result.

Steps to reproduce

CREATE TABLE t (id INT, s STRUCT<x: DOUBLE>) USING iceberg;

INSERT INTO t
SELECT id, IF(id % 2 = 0, named_struct('x', CAST(id AS DOUBLE)), NULL)
FROM range(100);

SELECT value_counts, null_value_counts FROM t.data_files;

Run the insert once with spark.comet.write.iceberg.splitOperator.enabled=true and spark.comet.write.iceberg.enabled=true, where the native writer engages (CometIcebergWriteExec in the plan), and once with only the split operator on. The test "native acceleration: NaN counts skip values under NULL structs, lists and maps" in #6658 builds the same shape and compares NaN counts only.

Expected behavior

The native path reports the counts iceberg-java reports for the linked Iceberg version: the footer's counts through 1.8, and from 1.9 only the rows where every enclosing struct is present.

Additional context

Split out of #6562, which covers the NaN counts. Part of #5649.

Activity

  1. andygrove commented on Oct 5, 2026

    @andygrove
    MemberAuthor

    Closing as a duplicate of an accepted divergence the user guide already documents. Under "Manifest metadata visible to later readers", iceberg-writes.md says that manifest value_counts / null_value_counts for float/double columns nested under a nullable struct count the rows whose parent struct is null, while iceberg-java's writer-tracked counts do not. Both counts grow by the same amount, so null ratios and IS NULL / IS NOT NULL pruning are unaffected, which matches the "never prunes more" analysis above.

    That entry dates the difference to Iceberg 1.10, but it starts in 1.9, when ParquetMetrics arrived. #6681 corrects the version.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:Icebergarea:writerNative Parquet writerbugSomething isn't workingpriority:mediumFunctional bugs, performance regressions, broken features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions