Describe the bug
For a float or double field inside a struct, the native Iceberg writer's value_counts and null_value_counts differ from iceberg-java's on Iceberg 1.9 and later whenever the struct is NULL in some rows. iceberg-java counts only the rows where the struct is present. The native path counts every row, and counts the NULL-struct rows as nulls.
With the reproduction below on Spark 4.1 / Iceberg 1.11 (100 rows, s NULL on the 50 odd ids), the table's data files add up to a value_count of 100 and a null_value_count of 50 for s.x natively, against 50 and 0 through iceberg-java. #6562 saw the same numbers.
The native path takes these counts from the Parquet footer. IcebergReflection.buildFloatFieldMetrics builds one FieldMetrics per float/double leaf for ParquetUtil.footerMetrics, with the value and null counts copied from the native DataFile, which carries the footer's counts. Its doc assumes they are "footer-derived on both sides". That held through Iceberg 1.8, whose footerMetrics takes value and null counts from the footer and uses FieldMetrics only for NaN counts and bounds. From 1.9, footerMetrics delegates to ParquetMetrics, whose metricsFromFieldMetrics takes the value, null and NaN counts from the FieldMetrics whenever one exists for the field. The JVM writer's FieldMetrics come from its value writers, and ParquetValueWriters.OptionWriter writes a NULL struct's columns with column.writeNull directly, so the field's own writer never sees those rows.
Fields under a list or map are not affected, because ParquetMetrics keeps no metrics for them.
The native counts never prune more than iceberg-java's. They report at least as many nulls, so IS NULL prunes no more files, and "only nulls", "only NaNs" and "nulls or NaNs" come out the same or less often. So this is a metadata parity difference, visible in the files and data_files metadata tables, rather than a wrong query result.
Steps to reproduce
CREATE TABLE t (id INT, s STRUCT<x: DOUBLE>) USING iceberg;
INSERT INTO t
SELECT id, IF(id % 2 = 0, named_struct('x', CAST(id AS DOUBLE)), NULL)
FROM range(100);
SELECT value_counts, null_value_counts FROM t.data_files;
Run the insert once with spark.comet.write.iceberg.splitOperator.enabled=true and spark.comet.write.iceberg.enabled=true, where the native writer engages (CometIcebergWriteExec in the plan), and once with only the split operator on. The test "native acceleration: NaN counts skip values under NULL structs, lists and maps" in #6658 builds the same shape and compares NaN counts only.
Expected behavior
The native path reports the counts iceberg-java reports for the linked Iceberg version: the footer's counts through 1.8, and from 1.9 only the rows where every enclosing struct is present.
Additional context
Split out of #6562, which covers the NaN counts. Part of #5649.
Describe the bug
For a
floatordoublefield inside a struct, the native Iceberg writer'svalue_countsandnull_value_countsdiffer from iceberg-java's on Iceberg 1.9 and later whenever the struct is NULL in some rows. iceberg-java counts only the rows where the struct is present. The native path counts every row, and counts the NULL-struct rows as nulls.With the reproduction below on Spark 4.1 / Iceberg 1.11 (100 rows,
sNULL on the 50 odd ids), the table's data files add up to avalue_countof 100 and anull_value_countof 50 fors.xnatively, against 50 and 0 through iceberg-java. #6562 saw the same numbers.The native path takes these counts from the Parquet footer.
IcebergReflection.buildFloatFieldMetricsbuilds oneFieldMetricsper float/double leaf forParquetUtil.footerMetrics, with the value and null counts copied from the nativeDataFile, which carries the footer's counts. Its doc assumes they are "footer-derived on both sides". That held through Iceberg 1.8, whosefooterMetricstakes value and null counts from the footer and usesFieldMetricsonly for NaN counts and bounds. From 1.9,footerMetricsdelegates toParquetMetrics, whosemetricsFromFieldMetricstakes the value, null and NaN counts from theFieldMetricswhenever one exists for the field. The JVM writer'sFieldMetricscome from its value writers, andParquetValueWriters.OptionWriterwrites a NULL struct's columns withcolumn.writeNulldirectly, so the field's own writer never sees those rows.Fields under a list or map are not affected, because
ParquetMetricskeeps no metrics for them.The native counts never prune more than iceberg-java's. They report at least as many nulls, so
IS NULLprunes no more files, and "only nulls", "only NaNs" and "nulls or NaNs" come out the same or less often. So this is a metadata parity difference, visible in thefilesanddata_filesmetadata tables, rather than a wrong query result.Steps to reproduce
Run the insert once with
spark.comet.write.iceberg.splitOperator.enabled=trueandspark.comet.write.iceberg.enabled=true, where the native writer engages (CometIcebergWriteExecin the plan), and once with only the split operator on. The test "native acceleration: NaN counts skip values under NULL structs, lists and maps" in #6658 builds the same shape and compares NaN counts only.Expected behavior
The native path reports the counts iceberg-java reports for the linked Iceberg version: the footer's counts through 1.8, and from 1.9 only the rows where every enclosing struct is present.
Additional context
Split out of #6562, which covers the NaN counts. Part of #5649.