docs: refresh stale roadmap entries - #6309
Conversation
Update roadmap sections whose status changed since the last pass: - Iceberg writes landed behind two experimental flags (apache#4658, apache#5361); point at the production-quality epic apache#5649. - Native Iceberg deletion vector reads landed (apache#5853); list the V3 features that still fall back, and drop the stale iceberg_scan.rs scheme-match reference. - The codegen dispatcher covers only scalar expressions, not aggregates; link the expression reference, which lists dispatched expressions. - mapInArrow and mapInPandas have experimental support (apache#4234). - Replace the closed TPC-DS epic apache#858 with apache#2551. - datafusion-spark function wiring (apache#4150) is nearly complete. - Link the hash join spill issue and the upstream DataFusion design. - Describe the delta-spark contrib scan (apache#5365) and the convergence proposal (apache#5411).
sunchao
left a comment
There was a problem hiding this comment.
Summary
- Prior state and problem: The contributor roadmap described several implemented features as pending and linked to outdated tracking work.
- Design approach: Refresh capability descriptions and issue references in
docs/source/contributor-guide/roadmap.md. - Correctness / compatibility analysis: The revised claims agree with the relevant implementation and linked issue/PR states. Checked aggregate dispatch, Iceberg read/write gates, and Python support against Comet shims and relevant Spark sources across 3.4, 3.5, 4.0, 4.1, and experimental 4.2. Python acceleration remains limited to Spark 4.x.
- Key design decisions: Preserve experimental/default-off qualifications and distinguish implemented features from proposed work. The change adds no runtime abstraction or complexity.
- Implementation sketch: One documentation file changes, with 51 additions and 37 deletions. The entire base-relative diff was reviewed.
- Behavioral changes worth calling out: Contributor guidance changes. Execution behavior, Spark compatibility, and runtime overhead are unchanged.
- Suggested improvements: None meeting the P1/P2 bar. No introduced P1/P2 issues found within this review.
Reviewed head 47dea8e350adf711c220f18f31825ce3a77e2e0c against base 76493618769ff85e7a43e6f6c79f8b58d9e0b1ad. The PR is not a draft. The snapshot and live discussion contain no reviews, issue comments, inline comments, or review threads.
Routed skills: review-comet-pr. No sibling skill applies under its documentation-only routing rule.
Exact-head CI: Seven successful checks and fifteen skipped jobs, with no failures or unfinished checks. Runtime builds, Spark/Iceberg suites, and site deployment were skipped.
Validation: git diff --check, reference-definition checks, and local-link checks passed. Linked development statuses were verified. Prettier could not run because prettier/npx are unavailable, and Sphinx is not installed. No documentation build or runtime tests were run. The working tree remains unchanged.
Which issue does this PR close?
No issue. This is a documentation refresh of the contributor roadmap, following the last pass in #5064.
Rationale for this change
Several roadmap sections still describe work as unstarted or blocked when it has since landed, and a few references point at issues that have closed or code that has moved. Contributors use the roadmap to find where work is coordinated, so stale entries send them to the wrong place.
What changes are included in this PR?
Only
docs/source/contributor-guide/roadmap.mdchanges. Each section was checked against main and the linked issues and PRs:variant/geometry/geography/unknowntypes. The HDFS note referred to a scheme match iniceberg_scan.rsthat has since moved toiceberg_common.rs. It now describes the supported storage backends without naming a file.CometBatchKernelCodegenrejectsAggregateFunction), so incompatible aggregates fall back to Spark. The link now points at the expression reference, whose Implementation column lists the dispatched expressions; the compatibility guide it pointed at doesn't have that list.mapInArrowandmapInPandashave had experimental support since feat: Add experimental support for accelerated PyArrow UDFs #4234.datafusion-sparkfunction now has a Comet serde (Support expressions already implemented in datafusion-spark crate #4150); the remaining exception,json_tuple, is tracked in [Feature] Support Spark expression: json_tuple #3160.HashJoinExec([EPIC] Spilling Hash Join — run any join in a bounded memory budget datafusion#24768).delta-kernel-rspath (Converge the two Delta read paths into one plugin (clean architecture + performance) #5411). The dormant plain-table draft feat: support native Comet scan of plain Delta Lake tables #4669 is dropped because feat: add native Delta Lake scan contrib module (page/row-group pruning) #5365 supersedes its approach; I can add it back if we'd rather keep it listed.The window, lambda, native Parquet write, and memory management sections are still accurate and are unchanged.
How are these changes tested?
This is a documentation-only change.
prettier --checkpasses on the file, and every reference-style link it uses is defined, with no unused definitions.