You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.
No type reclassifications. Every bug or enhancement label an author had already applied matched the issue's content. The 20 issues that arrived without a type label were classified from their bodies.
The guide's area table still lacks area:memory, area:Iceberg, area:udf and area:joins, so this pass did not add them. Where an issue already carries one, it is listed as "also carries" below, along with other labels outside the guide's table.
Field id matching misses a container id after INT96 coercion drops container metadata (#6131)
Area labels: area:scan
Rationale: With spark.sql.parquet.fieldId.read.enabled=true, the native scan returns nulls for an id-matched struct that contains a timestamp, where Spark returns the data. That is silent wrong results at step 1. The Spark setting is opt-in, but the INT96 timestamps that trigger it are Spark's default encoding.
CometNativeScanExec evaluates the inner (unrewritten) dynamic-pruning subquery from outputPartitioning: q5 fails with SubqueryAdaptiveBroadcastExec.execute(), q64 intermittently returns 0 rows at SF1000 (#6133)
Native Iceberg write merges -0.0 and 0.0 rows into one float/double identity partition (#6138)
Area labels: area:writer; also carries area:Iceberg, correctness
Rationale: Rows are written under the other signed zero's partition value, so a filter that prunes on partition values can drop them with no error. That is the guide's data-corruption case at step 1, reachable only with the off-by-default native writer (see escalations).
Native Iceberg write silently drops S3 settings it cannot honour (credentials provider, SSE, ACL, tags, remote signing) instead of falling back (#6139)
Area labels: area:writer; also carries area:Iceberg, correctness
Rationale: Objects can be written without the table owner's SSE-KMS or SSE-C encryption, ACL, tags or storage class, and nothing reports it. The guide's priority:critical row covers security vulnerabilities. The native writer is off by default (see escalations).
Comparison operators on arrays and structs with floating-point leaves do not match Spark for signed zero (#6157)
Area labels: area:expressions; also carries correctness
CometSort sorts a multi-column key containing a collated string by raw bytes (#6158)
Area labels: none; also carries correctness
Rationale: row_number() over a multi-column key with a UTF8_LCASE string comes back in byte order, with no fallback. That is silent wrong results at step 1, on Spark 4.0+ collations.
Codegen dispatcher writes a null map key as the key type's default value (#6172)
Area labels: area:expressions
Rationale: On the default configuration, map_keys(transform_values(try_cast(...))) returns [1, 0] where Spark returns [1, NULL], which is silent wrong results at step 1.
AQE reuses one exchange for scans with different dynamic pruning filters and drops rows (#6264)
Storage-partitioned self-join of an Iceberg table returns duplicate rows with partially clustered distribution (#6278)
Area labels: area:scan
Rationale: On the default-on native Iceberg scan, Comet returns 378 rows where Spark returns 126, which is silent wrong results at step 1.
Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results (#6288)
Area labels: area:ffi; also carries correctness
Rationale: On default configs, after LIMIT ... OFFSET or a grouped aggregate with more than one batch of groups, Comet returns wrong boolean values and puts nulls on the wrong rows. The guide lists "boolean arrays with non-zero offset" at the FFI boundary as a critical example. Escalated from the author's priority:high.
Accelerated mapInArrow reads Python output at the declared types without checking them (#6290)
Area labels: none; also carries area:udf, correctness
Rationale: Where Spark raises on an output type mismatch, the accelerated path reads the buffers at the declared width. It returns wrong integers or decimals, or reads past the end of a buffer, which is step 1. Escalated from the author's priority:medium; the feature is experimental and off by default (see escalations).
Nested DATE to numeric casts in structs and maps return the day count or fail (#6316)
Area labels: area:expressions; also carries correctness
Rationale: In legacy mode, the Spark 3.x default, a DATE to INT cast inside a struct or map returns the day count where Spark returns NULL. The plan is fully native and there is no error, which is step 1; the other numeric and boolean targets fail the query with an internal error. Confirms the author's label.
A native plan with no JVM input returns truncated output without an error when the Tokio runtime shuts down (#6294)
Area labels: area:ffi; also carries correctness
Rationale: When an executor shuts down mid-task, the task reports success with truncated output: 1.2M of 4.8M rows in the reproduction. That is silent wrong results at step 1, and confirms the author's label.
priority:high
arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)" (#6251)
JVM columnar shuffle fails with ArrayIndexOutOfBoundsException when spark.shuffle.checksum.enabled=false (#6256)
Area labels: area:shuffle
Rationale: With Spark's shuffle checksums disabled, every task of a sort-based JVM columnar shuffle throws an unhandled ArrayIndexOutOfBoundsException. That matches the guide's "NPE on supported code path" example at step 2, and it affects 1.0.0 and 1.1.0.
spark.comet.batchSize below 8192 makes CometConf fail to initialize on every executor (#6286)
Area labels: none
Rationale: A valid tuning value makes CometConf throw ExceptionInInitializerError, then NoClassDefFoundError for every later task on the executor, so the job aborts. That is an unhandled exception that breaks every query, step 2.
priority:medium
Iceberg scan fails with a hard-wired 10s OpenDAL io_timeout that cannot be configured (#6124)
Area labels: area:scan; also carries area:Iceberg
Rationale: A native Iceberg read fails on a 10-second timeout that iceberg-java does not impose. The failure is visible and there are workarounds: disable the native Iceberg scan, or raise COMET_WORKER_THREADS if the cause is worker starvation. Step 3.
Native Iceberg write gate reads hdfs:/path as file, so the write fails natively instead of falling back (#6140)
Area labels: area:writer; also carries area:Iceberg
Native Iceberg write fails on a V1 spec that mixes a live field with a void field whose source column was dropped (#6141)
Area labels: area:writer; also carries area:Iceberg
Rationale: Every task fails where iceberg-java succeeds, with no fallback. The failure is visible and needs the off-by-default native writer, so step 3.
Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 (#6145)
Area labels: area:writer; also carries area:Iceberg
Rationale: The NULL partition value is silent wrong metadata, but it needs dates past year 262143 on the off-by-default writer. Held at medium the way Iceberg serde: delete-file fields fall back to wrong defaults on reflection failure #5256 was held for a correctness path that cannot be reached in practice (see escalations). The author called it low priority.
Native Iceberg write over-counts NaNs for nested float fields when the input batch is already sliced (#6146)
Area labels: area:writer; also carries area:Iceberg
Rationale: Data files get wrong nan_value_counts for list and map float elements, which no predicate can prune on. So this is broken metadata rather than wrong query results, step 3.
RevertNativeForTransitionHeavyStages strips transitions in the stage below when AQE is off (#6152)
Area labels: none
Rationale: Found by code reading: the rule can leave a row operator over a columnar child. That needs both the off-by-default spark.comet.exec.transitionRevert.enabled and AQE off, so step 3.
CometCast.isSupported returns the first non-Compatible child, so Incompatible can mask Unsupported (#6200)
Area labels: area:expressions
Rationale: For a struct or map cast with an Unsupported child, the support-level check reports Incompatible, and under allowIncompatible that child would run natively. No query reaches it today, because the only Incompatible cast needs a negative-scale decimal that cannot normally be built, so it is held at medium (see escalations).
ANSI integral overflow errors: wrong error class for Byte/Short arithmetic and missing try_ suggestion (#6217)
fair_unified: a JVM consumer freeing its last bytes can fail a parked native acquire with NoSuchElementException (#6224)
Area labels: none; also carries area:memory
Rationale: A race in Spark's ExecutionMemoryPool fails the task with NoSuchElementException instead of refusing the request so the operator can spill. It is shown at component level only, and the task fails visibly, so step 3.
Reduce retained buffer allocation when list_extract selects short nested arrays (#6225)
Area labels: area:expressions
Rationale: perf: optimize list_extract without defaults using Arrow take #5174, which is in the 1.1 branch, makes list_extract retain about 32 times the buffer capacity when it selects short inner arrays, and native shuffle charges that capacity against its reservation. The values are correct and there is no end-to-end measurement yet, so this is a memory regression at step 3.
Spark errors after the first batch of a JVM input stream reach the user as CometNativeException (#6234)
Grouped integer SUM doesn't count its per-group state in the aggregate's memory reservation (#6252)
Area labels: area:aggregation; also carries area:memory
Rationale: size() reports the struct rather than its per-group vectors, so every grouped integer SUM under-reserves by 16 bytes per group and spills late. The results are correct, so step 3.
Native window operators reserve no memory for the batches they buffer (#6253)
Area labels: none; also carries area:memory
Rationale: WindowAggExec runs Spark's default frame for agg(...) OVER (PARTITION BY k). It keeps every input batch of the task's partition with no reservation, where Spark's WindowExec spills, so a large partition can push the executor past its container limit. Not yet measured, so step 3.
A native final aggregate that has spilled can fail the task during its replay (#6254)
Area labels: area:aggregation; also carries area:memory
Rationale: The replay after a spill fails the task with Additional allocation failed instead of spilling again, and it reproduces. The failure is visible and falling back to Spark's aggregate avoids it, so step 3; see the escalations.
Hash-based JVM columnar shuffle reports its output as disk spill instead of bytes written (#6258)
Spark 3.4 and 3.5 release jars require Java 17 since 0.11.0, though the docs list Java 11 (#6283)
Area labels: area:ci; also carries build
Rationale: The published Spark 3.x jars fail to load on Java 11 with UnsupportedClassVersionError. The failure is visible and running on Java 17 avoids it, and the cause is in the release build tooling. Confirms the author's label.
Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset (#6292)
Area labels: none; also carries performance
Rationale: A standalone executor with spark.executor.cores unset, the standalone default, gets a single Tokio worker. In the reporter's measurement, a native scan feeding a sort ran 2.4 times slower with one worker than with four, which is significant performance degradation at step 3. Confirms the author's label; see the escalations for the deadlock.
A JVM consumer's parked page allocation fails when another consumer of the task empties its balance (#6304)
Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE (#6315)
Area labels: area:writer
Rationale: On Spark 3.4 and 3.5, a native INSERT INTO ... SELECT reads back empty until the table is refreshed, which is silent. But the native Parquet write needs both the Testing-category spark.comet.parquet.write.enabled and the DataWritingCommandExecallowIncompatible opt-in, so as with [Bug] Native make_interval overflows valid time components or loses seconds precision #5131, users only see it after accepting incompatible behavior. Confirms the author's label (see escalations).
Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval (#6313)
Area labels: none; also carries performance, regression
native: JVMClasses::with_env: JAVA_VM not initialized aborts a multi-suite JVM (#6096)
Area labels: area:ffi
Rationale: The JNI lifecycle abort has only been seen when 132 Comet suites share one JVM with debug assertions on, and CI splits the suites and passes. That is a test-only failure at step 4 (see escalations).
Three dev/diffs weaken a local-shuffle-read assertion into a contradiction, breaking the Spark-only baseline (#6122)
Area labels: spark sql tests
Rationale: The patched test cannot pass for any number of local reads. It is skipped under Comet and breaks only the Spark-only baseline run, so this is test-only, step 4.
Native memory usage log underestimates the executor overhead for PySpark and SparkR on Kubernetes, and warns on standalone clusters (#6188)
Area labels: none; also carries area:memory
Rationale: The log warns of a container kill that cannot happen, and warns on standalone clusters that have no container. That is a misleading log line, step 4.
In-memory cache tests that use checkSparkAnswer compare the cache with itself (#6203)
Area labels: none; also carries test
Rationale: 23 cache tests cannot catch a wrong stored value or a wrong pruning bound, which is a test-only problem at step 4. Confirms the author's label; it bears on the default-on decision in feat: enable Comet's in-memory cache by default #5634.
With several native plans in one task, the non-zero memory usage warning comes from the wrong plan (#6255)
Area labels: none; also carries area:memory
Rationale: The first plan logs a leak that is not one, and a real leak in a later plan is never reported. Log diagnostics only, step 4.
spark.comet.shuffle.jvm.batchSize=0 hangs a task, and an unknown spark.comet.exec.memoryPool fails every task (#6259)
Area labels: area:shuffle; also carries area:memory
Rationale: Two settings are not validated, and the failures need an invalid value: a zero batch size, or a misspelled pool name that fails every task with a clear message. Step 4.
The native memory usage log understates untracked memory while a pool is overcommitted (#6260)
Area labels: none; also carries area:memory
Rationale: While a pool is overcommitted, the log's untracked figure and the container warning come out low by the overcommit. The author notes that is usually small and short-lived. Log diagnostics, step 4.
An input ArrowArrayStream that native never takes is never released (#6289)
Area labels: area:ffi
Rationale: The stream, and the first input batch it buffered, stay pinned for the executor's life only when a task ends before native takes the stream, for example a failed createPlan or an early cancellation. That is a leak on error paths, and confirms the author's label.
Area labels: none; also carries documentation, area:memory
Rationale: Documentation corrections and the remaining work for the config deprecation.
Kubernetes guide example never enables Comet because it sets no off-heap memory (#6190)
Area labels: none; also carries documentation
Rationale: A documentation correction, which the guide's type table places under enhancement.
ExtractANSIIntervalDays and the other interval field extractors have no serde, so date + <day interval column> and extract of an interval fall back (#6193)
Area labels: area:expressions
Rationale: New expression support; the fallback is correct today.
Charge the native shuffle writer's write buffers to the memory pool (#6196)
Area labels: area:shuffle; also carries area:memory
Support anonymous S3 access through the CometS3CredentialProvider SPI (#6298)
Area labels: area:scan
Rationale: New capability for the credential provider SPI. Today the adapter fails with a clear error for anonymous buckets.
Support Spark 4.2 geometry type + functions (#6317)
Area labels: area:expressions
Rationale: New type and function support for Spark 4.2's GEOMETRY and its initial ST_* functions.
Escalations to consider
Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results (#6288)
Escalated in this pass from the author's priority:high to priority:critical. The guide's list of correctness bugs includes "Data corruption in FFI boundary (e.g., boolean arrays with non-zero offset)". The reproductions return wrong values on default configs through ordinary LIMIT ... OFFSET and grouped-aggregate plans.
Accelerated mapInArrow reads Python output at the declared types without checking them (#6290)
Escalated in this pass from the author's priority:medium to priority:critical on decision-tree step 1. The feature is experimental and off by default (spark.comet.exec.pyarrowUDF.enabled), and the report comes from reading the code rather than a run. The reviewer may prefer to restore priority:medium under the "core path over experimental" principle.
array_append: ANSI item error is swallowed when the array is NULL (#6086)
Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 (#6145)
Held at priority:medium. The year and month NULL partition value is silent wrong metadata, which step 1 would place at critical. But it needs dates beyond year 262143 on the off-by-default writer, so it is held at medium, as Iceberg serde: delete-file fields fall back to wrong defaults on reflection failure #5256 was. The author's own "low priority" reading is also defensible.
CometCast.isSupported returns the first non-Compatible child, so Incompatible can mask Unsupported (#6200)
A native final aggregate that has spilled can fail the task during its replay (#6254)
Matches the guide's trigger "A priority:medium bug is reported by multiple users or affects a common workload → consider escalating to priority:high". Any grouped aggregate whose final stage spills under memory pressure can hit it. The upstream report, aggregate_memory_spill.slt Case G fails intermittently with ResourcesExhausted datafusion#25423, was closed by a test-only change, so the behavior is unchanged in DataFusion 55.1.
Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset (#6292)
Matches the same trigger. Leaving spark.executor.cores unset is the standalone default. On main, the single worker also deadlocked in 4 of 4 runs with 96 MB or 128 MB of off-heap memory. fix: return a native plan's memory before its Spark task ends #6261, still open, fixes the deadlock but not the slowdown.
Hash-based JVM columnar shuffle reports its output as disk spill instead of bytes written (#6258)
native: JVMClasses::with_env: JAVA_VM not initialized aborts a multi-suite JVM (#6096)
Labeled priority:low as a test-harness failure. If the JAVA_VM lifecycle gap can be hit by a long-lived production process, for example one that loads the native library a second time, the result is a JVM abort, which is priority:high at step 2.
Skipped (needs more info)
Both are tracking or record issues rather than bug reports or feature requests. No bug or enhancement label fits, so requires-triage was left in place, and they will reappear in each pass until they are closed.
[EPIC] Bug fixes to consider backporting to branch-1.0 (#6201)
The summary issue from the 2026-08-24 pass. None of its labels were applied at the time, but all 28 issues it covers have since been triaged: each carries a type label and none still carries requires-triage. The three earlier summaries it links are closed. Nothing in it is outstanding, so it can be closed.
Triage pass over the open
requires-triagequeue, per the project Bug Triage Guide.priority:critical15,priority:high3,priority:medium20,priority:low8Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.
Notes on this pass:
priority:mediumtopriority:critical. In both, Comet returns rows where Spark raises an error. This pass applies that rule directly, so Nested duplicate names in a metadata-free Parquet file bypass the native resolver when the file schema equals the requested schema #6136 and array_append: ANSI item error is swallowed when the array is NULL #6086 are critical. Two other corrections are cited in the escalations below: a reviewer raised the metrics defect Native metrics from several plan instances in one task overwrite each other, so a coalesced scan reports only its last partition #5879 to high, and lowered array_distinct and array_union diverge from Spark on -0.0 for Spark versions without SPARK-54918 #5701 from critical to medium.bugorenhancementlabel an author had already applied matched the issue's content. The 20 issues that arrived without a type label were classified from their bodies.priority:high) and Accelerated mapInArrow reads Python output at the declared types without checking them #6290 (frompriority:medium); see the escalations below. Author priorities were confirmed unchanged on In-memory cache tests that use checkSparkAnswer compare the cache with itself #6203, ANSI integral overflow errors: wrong error class for Byte/Short arithmetic and missing try_ suggestion #6217, Spark 3.4 and 3.5 release jars require Java 17 since 0.11.0, though the docs list Java 11 #6283, An input ArrowArrayStream that native never takes is never released #6289, Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset #6292, A native plan with no JVM input returns truncated output without an error when the Tokio runtime shuts down #6294, Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE #6315 and Nested DATE to numeric casts in structs and maps return the day count or fail #6316.mapInArrow(Accelerated mapInArrow reads Python output at the declared types without checking them #6290) are experimental and off by default. Silent wrong results there are labeled critical, as revertToSpark erases CometIcebergWriteExec / CometNativeWriteExec because originalPlan is the node's own child #5719 was, with an escalation note. The exception is Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 #6145, which needs dates past year 262143 and is held at medium. The native Parquet writer (Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE #6315) also needs an explicitallowIncompatibleopt-in, so it follows [Bug] Native make_interval overflows valid time components or loses seconds precision #5131 and stays at medium.priority:low. As in earlier passes, it was left in place, and the reviewer may want to remove it.area:memory,area:Iceberg,area:udfandarea:joins, so this pass did not add them. Where an issue already carries one, it is listed as "also carries" below, along with other labels outside the guide's table.spark 4as an area indicator, but the repository only hasspark 4.0/spark 4.1/spark 4.2/spark 3.x. So nothing was applied to the Spark 4-only issues Support scalar PySpark Arrow UDFs in Comet native execution via PyO3 #6129, CometSort sorts a multi-column key containing a collated string by raw bytes #6158,CometLiteralaccepts a collated string literal and serializes it as a plain string #6232, MoveCometCollationSuitetospark-4.xso every 4.x profile runs it #6233 and Support Spark 4.2 geometry type + functions #6317.regressionis not in the guide, so this pass neither added nor removed it. Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval #6313 carries it, but the repository describes the label as "A bug that did not affect the most recent Comet release", while the issue dates the behavior to perf: executePlan uses a channel to park executor task thread instead of yield_now() [iceberg] #3553 in 0.14.0. The reviewer may want to remove it there. Reduce retained buffer allocation when list_extract selects short nested arrays #6225 does fit that description: perf: optimize list_extract without defaults using Arrow take #5174 introduced it after 1.0.0, and perf: optimize list_extract without defaults using Arrow take #5174 is in the 1.1 branch.Bugs
priority:critical
area:expressionsDIVIDE_BY_ZERO, andgetSupportLevelreportsCompatible, so the error is lost with no fallback. That is step 1, consistent with Codegen dispatcher: whole-tree NullIntolerant short-circuit suppresses ANSI errors, plus TIME type gaps between canHandle and the runtime dispatcher #5218 and with the reviewer corrections that raised Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed #5801 and Field id gating differs from Spark: root-only check and no dependence on fieldId.read.enabled #5936 to critical. Reachable on Spark 3.4 and 3.5 only.area:scanspark.sql.parquet.fieldId.read.enabled=true, the native scan returns nulls for an id-matched struct that contains a timestamp, where Spark returns the data. That is silent wrong results at step 1. The Spark setting is opt-in, but the INT96 timestamps that trigger it are Spark's default encoding.area:scanarea:scanfoundDuplicateFieldInCaseInsensitiveModeError. The 2026-09-14 pass put the duplicate-field-id sibling Duplicate field ids inside a struct are not validated when the file schema equals the requested schema and no predicate is pushed #5801 at medium and a reviewer raised it to critical, so this follows the corrected call.area:writer; also carriesarea:Iceberg,correctnessarea:writer; also carriesarea:Iceberg,correctnesspriority:criticalrow covers security vulnerabilities. The native writer is off by default (see escalations).area:expressions; also carriescorrectness<=>,<and>=over nested-0.0and0.0run fully native and return the opposite answers from Spark, which is step 1. This is the same class as Nested floating-point IN membership does not match Spark for signed zero #6019, which the last pass put at critical.correctnessrow_number()over a multi-column key with aUTF8_LCASEstring comes back in byte order, with no fallback. That is silent wrong results at step 1, on Spark 4.0+ collations.area:expressionsmap_keys(transform_values(try_cast(...)))returns[1, 0]where Spark returns[1, NULL], which is silent wrong results at step 1.area:scanReusedExchangehands one branch the other branch's rows. The reproduction returns 7 or 21 rows where Spark returns 28, with no error, which is step 1. This is the likely cause of the q64 loss in CometNativeScanExec evaluates the inner (unrewritten) dynamic-pruning subquery from outputPartitioning: q5 fails with SubqueryAdaptiveBroadcastExec.execute(), q64 intermittently returns 0 rows at SF1000 #6133.area:scanarea:ffi; also carriescorrectnessLIMIT ... OFFSETor a grouped aggregate with more than one batch of groups, Comet returns wrong boolean values and puts nulls on the wrong rows. The guide lists "boolean arrays with non-zero offset" at the FFI boundary as a critical example. Escalated from the author'spriority:high.area:udf,correctnesspriority:medium; the feature is experimental and off by default (see escalations).area:expressions; also carriescorrectnessDATEtoINTcast inside a struct or map returns the day count where Spark returns NULL. The plan is fully native and there is no error, which is step 1; the other numeric and boolean targets fail the query with an internal error. Confirms the author's label.area:ffi; also carriescorrectnesspriority:high
area:expressions,area:ffiIllegalStateExceptionin the Arrow import instead of falling back. Two same-named inputs arise naturally after a join. That is an unhandled exception on a supported path, which the guide ratespriority:high, as with Hashing a CalendarInterval value fails with "Unsupported data type in hasher: Interval(MonthDayNano)" #5059.area:shuffleArrayIndexOutOfBoundsException. That matches the guide's "NPE on supported code path" example at step 2, and it affects 1.0.0 and 1.1.0.CometConfthrowExceptionInInitializerError, thenNoClassDefFoundErrorfor every later task on the executor, so the job aborts. That is an unhandled exception that breaks every query, step 2.priority:medium
area:scan; also carriesarea:IcebergCOMET_WORKER_THREADSif the cause is worker starvation. Step 3.area:writer; also carriesarea:IcebergUnsupported storage scheme: hdfsinstead of falling back. The failure is visible and needs the off-by-default native writer, so step 3. The default-on scan analog, Iceberg native scan claims schemes it cannot execute; three scheme lists disagree #5541, ispriority:high.area:writer; also carriesarea:Icebergarea:writer; also carriesarea:Icebergarea:writer; also carriesarea:Icebergnan_value_countsfor list and map float elements, which no predicate can prune on. So this is broken metadata rather than wrong query results, step 3.spark.comet.exec.transitionRevert.enabledand AQE off, so step 3.CometCast.isSupportedreturns the first non-Compatible child, soIncompatiblecan maskUnsupported(#6200)area:expressionsUnsupportedchild, the support-level check reportsIncompatible, and underallowIncompatiblethat child would run natively. No query reaches it today, because the onlyIncompatiblecast needs a negative-scale decimal that cannot normally be built, so it is held at medium (see escalations).area:expressionsarea:memoryExecutionMemoryPoolfails the task withNoSuchElementExceptioninstead of refusing the request so the operator can spill. It is shown at component level only, and the task fails visibly, so step 3.area:expressionslist_extractretain about 32 times the buffer capacity when it selects short inner arrays, and native shuffle charges that capacity against its reservation. The values are correct and there is no end-to-end measurement yet, so this is a memory regression at step 3.area:ffiarea:aggregation; also carriesarea:memorysize()reports the struct rather than its per-group vectors, so every grouped integerSUMunder-reserves by 16 bytes per group and spills late. The results are correct, so step 3.area:memoryWindowAggExecruns Spark's default frame foragg(...) OVER (PARTITION BY k). It keeps every input batch of the task's partition with no reservation, where Spark'sWindowExecspills, so a large partition can push the executor past its container limit. Not yet measured, so step 3.area:aggregation; also carriesarea:memoryAdditional allocation failedinstead of spilling again, and it reproduces. The failure is visible and falling back to Spark's aggregate avoids it, so step 3; see the escalations.area:shuffleMapStatusstays correct. That is a defect in existing metrics reporting, classified like Task input metrics are unreliable when a native block mixes a native scan with a JVM input #5336 and Report native child-operator spill metrics in Spark task metrics for unified shuffle plans #5382 (see escalations).area:ci; also carriesbuildUnsupportedClassVersionError. The failure is visible and running on Java 17 avoids it, and the cause is in the release build tooling. Confirms the author's label.performancespark.executor.coresunset, the standalone default, gets a single Tokio worker. In the reporter's measurement, a native scan feeding a sort ran 2.4 times slower with one worker than with four, which is significant performance degradation at step 3. Confirms the author's label; see the escalations for the deadlock.area:shuffleallocatePageinCometUnifiedShuffleMemoryAllocatorfails withNoSuchElementException. It is shown at component level only, and the task fails visibly, so step 3.area:writerINSERT INTO ... SELECTreads back empty until the table is refreshed, which is silent. But the native Parquet write needs both the Testing-categoryspark.comet.parquet.write.enabledand theDataWritingCommandExecallowIncompatibleopt-in, so as with [Bug] Native make_interval overflows valid time components or loses seconds precision #5131, users only see it after accepting incompatible behavior. Confirms the author's label (see escalations).performance,regressionspark.comet.metrics.updateIntervaland publish metrics on every batch. A one-file scan took 343 ms of process CPU against 209 ms with the interval honored. Performance degradation since perf: executePlan uses a channel to park executor task thread instead of yield_now() [iceberg] #3553 (0.14.0), step 3.priority:low
JVMClasses::with_env: JAVA_VM not initializedaborts a multi-suite JVM (#6096)area:ffispark sql testsarea:memorytestarea:memoryarea:shuffle; also carriesarea:memoryarea:memoryarea:fficreatePlanor an early cancellation. That is a leak on error paths, and confirms the author's label.Enhancements
AtLeastNNonNullsnatively (#6093)area:expressionsarea:ciarea:ciarea:scanarea:writer; also carriesarea:Iceberg,performancearea:writer; also carriesarea:Iceberg,area:memoryarea:shuffle; also carriesdocumentationarea:scanarea:scan; also carriesarea:Iceberg,area:memoryArrowEvalPythonExec, which falls back correctly today.area:memoryarea:ffi; also carriesarea:udfarea:udfarea:udfregisterAll(#6177)area:udfregisterAll.fair_unifieddescription and preparespark.comet.exec.memoryPool.fractionfor removal (#6187)documentation,area:memorydocumentationenhancement.ExtractANSIIntervalDaysand the other interval field extractors have no serde, sodate + <day interval column>andextractof an interval fall back (#6193)area:expressionsarea:shuffle; also carriesarea:memoryarea:ci; also carriesbuild,performancelibcomet, then possibly enable it.performance,area:memoryarea:scan; also carriesarea:IcebergCometLiteralaccepts a collated string literal and serializes it as a plain string (#6232)area:expressionsCometCollationSuitetospark-4.xso every 4.x profile runs it (#6233)area:writerarea:memoryarea:ffi; also carriesdocumentationarea:ffi; also carriesperformancearea:ffi; also carriesperformancearea:scanarea:scanarea:expressionsGEOMETRYand its initialST_*functions.Escalations to consider
priority:hightopriority:critical. The guide's list of correctness bugs includes "Data corruption in FFI boundary (e.g., boolean arrays with non-zero offset)". The reproductions return wrong values on default configs through ordinaryLIMIT ... OFFSETand grouped-aggregate plans.priority:mediumtopriority:criticalon decision-tree step 1. The feature is experimental and off by default (spark.comet.exec.pyarrowUDF.enabled), and the report comes from reading the code rather than a run. The reviewer may prefer to restorepriority:mediumunder the "core path over experimental" principle.ArrayAppend.evalshort-circuits exactly as Comet does; only Spark's generated code raises. If the reviewer reads that as Spark disagreeing with itself,priority:mediumfits, just as a reviewer lowered array_distinct and array_union diverge from Spark on -0.0 for Spark versions without SPARK-54918 #5701 from critical to medium.SparkUnsupportedOperationExceptionon the default scan selection, which ispriority:highat step 2.spark.comet.iceberg.write.enabled=true, a Testing-category setting that defaults to false. revertToSpark erases CometIcebergWriteExec / CometNativeWriteExec because originalPlan is the node's own child #5719 was labeled critical on the same basis and has kept it, but the reviewer may preferpriority:highunder "core path over experimental". Either way, it should be settled before the native writer is turned on by default (Enable the split-operator plan and native Iceberg writes by default #5644).priority:highbecause it was framed as a storage identity problem.priority:medium. TheyearandmonthNULL partition value is silent wrong metadata, which step 1 would place at critical. But it needs dates beyond year 262143 on the off-by-default writer, so it is held at medium, as Iceberg serde: delete-file fields fall back to wrong defaults on reflection failure #5256 was. The author's own "low priority" reading is also defensible.CometCast.isSupportedreturns the first non-Compatible child, soIncompatiblecan maskUnsupported(#6200)priority:mediumbecause no query reaches it today. If fix: fall back to Spark for a TRY cast of a map key that can fail #6179 or a newIncompatiblecast makes the maskedUnsupportedchild reachable, that cast would run natively underallowIncompatible, with a crash (as in TRY_CAST on narrowing map keys fails where Spark returns a map with a null key #5995) or a wrong result. Re-evaluate at that point.priority:mediumbug is reported by multiple users or affects a common workload → consider escalating topriority:high". Any grouped aggregate whose final stage spills under memory pressure can hit it. The upstream report,aggregate_memory_spill.sltCase G fails intermittently with ResourcesExhausted datafusion#25423, was closed by a test-only change, so the behavior is unchanged in DataFusion 55.1.spark.executor.coresunset is the standalone default. Onmain, the single worker also deadlocked in 4 of 4 runs with 96 MB or 128 MB of off-heap memory. fix: return a native plan's memory before its Spark task ends #6261, still open, fixes the deadlock but not the slowdown.priority:mediumalongside Task input metrics are unreliable when a native block mixes a native scan with a JVM input #5336 and Report native child-operator spill metrics in Spark task metrics for unified shuffle plans #5382. However, on 2026-09-14 a reviewer raised a similar metrics under-report, Native metrics from several plan instances in one task overwrite each other, so a coalesced scan reports only its last partition #5879, topriority:high.priority:mediumbecause the native Parquet write sits behind anallowIncompatibleopt-in, following [Bug] Native make_interval overflows valid time components or loses seconds precision #5131. The stale read is silent, so step 1 would place it at critical once the native Parquet writer runs without that opt-in.JVMClasses::with_env: JAVA_VM not initializedaborts a multi-suite JVM (#6096)priority:lowas a test-harness failure. If theJAVA_VMlifecycle gap can be hit by a long-lived production process, for example one that loads the native library a second time, the result is a JVM abort, which ispriority:highat step 2.Skipped (needs more info)
Both are tracking or record issues rather than bug reports or feature requests. No
bugorenhancementlabel fits, sorequires-triagewas left in place, and they will reappear in each pass until they are closed.requires-triage. The three earlier summaries it links are closed. Nothing in it is outstanding, so it can be closed.