Repository navigation
fix: don't let S3A committer, delete and read defaults block native Iceberg writes - #6808
Conversation
…ceberg writes apache#6441 makes a native Iceberg write to S3 fall back when the Hadoop configuration has an fs.s3a.* setting the native writer doesn't support. Spark 4.1 sets fs.s3a.committer.magic.enabled and fs.s3a.committer.name for every application when spark-hadoop-cloud is on the classpath (SPARK-47618), and many deployments set fs.s3a.experimental.input.fadvise, fs.s3a.bulk.delete.page.size or fs.s3a.committer.threads cluster-wide. Any of them makes every native S3 write fall back. None of these settings affects an Iceberg data-file write: the native writer doesn't go through S3A, fadvise is a read hint, and Iceberg commits through table metadata rather than a Hadoop output committer. Ignore them like the Spark-seeded read settings already in IgnoredHadoopS3Keys.
|
LGTM overall. Some minor questions.
|
|
Not a broader list, no. I've only added keys that S3A wouldn't apply to an Iceberg data-file write and that arrive as cluster-wide defaults rather than a per-job choice: Spark 4.1 sets the two committer keys itself whenever And yes to the guides. Both name the ignored keys (vectored-read and |
|
Thanks @manuzhang @snmvaughan |
…ache#6828) The user guide and the contributor guide both name the Hadoop S3A settings that the native Iceberg write gate ignores, and both list only the vectored-read and downgrade.syncable.exceptions settings. apache#6808 adds the S3A committer, bulk.delete.page.size and experimental.input.fadvise settings to that list. Name them in both guides, point the S3A row of the eligibility table at the list, state when a key belongs on it, and note that the list matches the plain key, so a per-bucket spelling still falls back.
Which issue does this PR close?
No issue. This follows up on #6441.
Rationale for this change
#6441 makes a native Iceberg write to S3 fall back to Spark when the Hadoop configuration has an
fs.s3a.*setting the native writer doesn't support. It already ignores Hadoop's built-in defaults and a few S3A read settings that Spark seeds into every session. Some other settings are also common in S3 deployments, and any one of them makes every native S3 write fall back:SparkContextsetsfs.s3a.committer.magic.enabled=trueandfs.s3a.committer.name=magicfor every application whenspark-hadoop-cloudis on the classpath (SPARK-47618).fs.s3a.committer.threads,fs.s3a.experimental.input.fadvise(Hadoop's S3A docs recommendrandomfor columnar formats) orfs.s3a.bulk.delete.page.sizecluster-wide.The fallback reason looks like this:
None of these settings affects an Iceberg data-file write:
fadviseis a read hint.What changes are included in this PR?
The five keys are added to
IgnoredHadoopS3KeysinCometIcebergNativeWrite, alongside the Spark-seeded read settings, with a comment explaining why.How are these changes tested?
There's a new test in
CometIcebergWriteDetectionSuite. It checks that these settings produce no unsupported keys, and that a setting the native writer can't honor (fs.s3a.encryption.algorithm) is still reported next to them. The whole suite passes locally.