docs: correct what the Iceberg split-operator plan moves into AQE - #6407
Merged
Merged
Conversation
Spark's InsertAdaptiveSparkPlan already wraps the input query of a V2CommandExec write in AQE; what sits outside AQE is the write operator itself. The split moves data-file writing inside AQE, apart from the commit, rather than making the write's input visible to AQE.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
No issue; documentation only. Follow-up to review feedback from @mbutrovich on the Comet 1.1.0 blog post (apache/datafusion-site#205), which repeated this wording.
Rationale for this change
Both Iceberg writes guides say the query feeding an Iceberg write can't be re-planned by AQE, and that the split-operator plan makes the write's input visible to AQE. That isn't right. Spark's
InsertAdaptiveSparkPlanhascase c: V2CommandExec => c.withNewChildren(c.children.map(apply)), in both Spark 3.5 and 4.0, so the input query of a V2 write such asAppendDataExecis already wrapped in AQE. What sits outside AQE is the write operator itself.As #4658 describes it, the split moves the data-file writing step inside AQE, so it can be re-planned in response to its upstream operators, and separates it from the metadata writing and commit, which Comet does not accelerate.
What changes are included in this PR?
user-guide/latest/iceberg-writes.md: the Overview now says AQE already re-plans the sub-query, that the write operator is what sits outside AQE, and that the split moves data-file writing inside AQE, apart from the commit.contributor-guide/iceberg-writes.md: the Split-Operator Plan section says the same and namesInsertAdaptiveSparkPlan. The plan diagram now marks which operators run inside and outside AQE instead of labeling the input query "now visible to AQE and Comet".How are these changes tested?
Documentation only.
npx prettier@latest --checkpasses on both files.