Skip to content

Bug triage results: 2026-09-28 #6321

Description

@andygrove

Triage pass over the open requires-triage queue, per the project Bug Triage Guide.

  • Date: 2026-09-28
  • Total issues processed: 82 (80 triaged, 2 skipped, 0 failed)
  • Type counts: 46 bugs, 34 enhancements
  • Priority counts applied: priority:critical 15, priority:high 3, priority:medium 20, priority:low 8
  • Guide: docs/source/contributor-guide/bug_triage.md

Labels have already been applied. A reviewer should spot-check the calls below and close this issue when satisfied. Corrections should be made directly on the affected issue.

Notes on this pass:

Bugs

priority:critical

  • array_append: ANSI item error is swallowed when the array is NULL (#6086)
  • Field id matching misses a container id after INT96 coercion drops container metadata (#6131)
    • Area labels: area:scan
    • Rationale: With spark.sql.parquet.fieldId.read.enabled=true, the native scan returns nulls for an id-matched struct that contains a timestamp, where Spark returns the data. That is silent wrong results at step 1. The Spark setting is opt-in, but the INT96 timestamps that trigger it are Spark's default encoding.
  • CometNativeScanExec evaluates the inner (unrewritten) dynamic-pruning subquery from outputPartitioning: q5 fails with SubqueryAdaptiveBroadcastExec.execute(), q64 intermittently returns 0 rows at SF1000 (#6133)
  • Nested duplicate names in a metadata-free Parquet file bypass the native resolver when the file schema equals the requested schema (#6136)
  • Native Iceberg write merges -0.0 and 0.0 rows into one float/double identity partition (#6138)
    • Area labels: area:writer; also carries area:Iceberg, correctness
    • Rationale: Rows are written under the other signed zero's partition value, so a filter that prunes on partition values can drop them with no error. That is the guide's data-corruption case at step 1, reachable only with the off-by-default native writer (see escalations).
  • Native Iceberg write silently drops S3 settings it cannot honour (credentials provider, SSE, ACL, tags, remote signing) instead of falling back (#6139)
    • Area labels: area:writer; also carries area:Iceberg, correctness
    • Rationale: Objects can be written without the table owner's SSE-KMS or SSE-C encryption, ACL, tags or storage class, and nothing reports it. The guide's priority:critical row covers security vulnerabilities. The native writer is off by default (see escalations).
  • Comparison operators on arrays and structs with floating-point leaves do not match Spark for signed zero (#6157)
  • CometSort sorts a multi-column key containing a collated string by raw bytes (#6158)
    • Area labels: none; also carries correctness
    • Rationale: row_number() over a multi-column key with a UTF8_LCASE string comes back in byte order, with no fallback. That is silent wrong results at step 1, on Spark 4.0+ collations.
  • Codegen dispatcher writes a null map key as the key type's default value (#6172)
    • Area labels: area:expressions
    • Rationale: On the default configuration, map_keys(transform_values(try_cast(...))) returns [1, 0] where Spark returns [1, NULL], which is silent wrong results at step 1.
  • AQE reuses one exchange for scans with different dynamic pruning filters and drops rows (#6264)
  • Storage-partitioned self-join of an Iceberg table returns duplicate rows with partially clustered distribution (#6278)
    • Area labels: area:scan
    • Rationale: On the default-on native Iceberg scan, Comet returns 378 rows where Spark returns 126, which is silent wrong results at step 1.
  • Sliced booleans nested in structs, and sliced inputs to the JVM UDF bridge, reach the JVM misaligned and return wrong results (#6288)
    • Area labels: area:ffi; also carries correctness
    • Rationale: On default configs, after LIMIT ... OFFSET or a grouped aggregate with more than one batch of groups, Comet returns wrong boolean values and puts nulls on the wrong rows. The guide lists "boolean arrays with non-zero offset" at the FFI boundary as a critical example. Escalated from the author's priority:high.
  • Accelerated mapInArrow reads Python output at the declared types without checking them (#6290)
    • Area labels: none; also carries area:udf, correctness
    • Rationale: Where Spark raises on an output type mismatch, the accelerated path reads the buffers at the declared width. It returns wrong integers or decimals, or reads past the end of a buffer, which is step 1. Escalated from the author's priority:medium; the feature is experimental and off by default (see escalations).
  • Nested DATE to numeric casts in structs and maps return the day count or fail (#6316)
    • Area labels: area:expressions; also carries correctness
    • Rationale: In legacy mode, the Spark 3.x default, a DATE to INT cast inside a struct or map returns the day count where Spark returns NULL. The plan is fully native and there is no error, which is step 1; the other numeric and boolean targets fail the query with an internal error. Confirms the author's label.
  • A native plan with no JVM input returns truncated output without an error when the Tokio runtime shuts down (#6294)
    • Area labels: area:ffi; also carries correctness
    • Rationale: When an executor shuts down mid-task, the task reports success with truncated output: 1.2M of 4.8M rows in the reproduction. That is silent wrong results at step 1, and confirms the author's label.

priority:high

  • arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)" (#6251)
  • JVM columnar shuffle fails with ArrayIndexOutOfBoundsException when spark.shuffle.checksum.enabled=false (#6256)
    • Area labels: area:shuffle
    • Rationale: With Spark's shuffle checksums disabled, every task of a sort-based JVM columnar shuffle throws an unhandled ArrayIndexOutOfBoundsException. That matches the guide's "NPE on supported code path" example at step 2, and it affects 1.0.0 and 1.1.0.
  • spark.comet.batchSize below 8192 makes CometConf fail to initialize on every executor (#6286)
    • Area labels: none
    • Rationale: A valid tuning value makes CometConf throw ExceptionInInitializerError, then NoClassDefFoundError for every later task on the executor, so the job aborts. That is an unhandled exception that breaks every query, step 2.

priority:medium

  • Iceberg scan fails with a hard-wired 10s OpenDAL io_timeout that cannot be configured (#6124)
    • Area labels: area:scan; also carries area:Iceberg
    • Rationale: A native Iceberg read fails on a 10-second timeout that iceberg-java does not impose. The failure is visible and there are workarounds: disable the native Iceberg scan, or raise COMET_WORKER_THREADS if the cause is worker starvation. Step 3.
  • Native Iceberg write gate reads hdfs:/path as file, so the write fails natively instead of falling back (#6140)
  • Native Iceberg write fails on a V1 spec that mixes a live field with a void field whose source column was dropped (#6141)
    • Area labels: area:writer; also carries area:Iceberg
    • Rationale: Every task fails where iceberg-java succeeds, with no fallback. The failure is visible and needs the off-by-default native writer, so step 3.
  • Native Iceberg write panics or writes a NULL partition value for dates and timestamps beyond year 262143 (#6145)
  • Native Iceberg write over-counts NaNs for nested float fields when the input batch is already sliced (#6146)
    • Area labels: area:writer; also carries area:Iceberg
    • Rationale: Data files get wrong nan_value_counts for list and map float elements, which no predicate can prune on. So this is broken metadata rather than wrong query results, step 3.
  • RevertNativeForTransitionHeavyStages strips transitions in the stage below when AQE is off (#6152)
    • Area labels: none
    • Rationale: Found by code reading: the rule can leave a row operator over a columnar child. That needs both the off-by-default spark.comet.exec.transitionRevert.enabled and AQE off, so step 3.
  • CometCast.isSupported returns the first non-Compatible child, so Incompatible can mask Unsupported (#6200)
    • Area labels: area:expressions
    • Rationale: For a struct or map cast with an Unsupported child, the support-level check reports Incompatible, and under allowIncompatible that child would run natively. No query reaches it today, because the only Incompatible cast needs a negative-scale decimal that cannot normally be built, so it is held at medium (see escalations).
  • ANSI integral overflow errors: wrong error class for Byte/Short arithmetic and missing try_ suggestion (#6217)
  • fair_unified: a JVM consumer freeing its last bytes can fail a parked native acquire with NoSuchElementException (#6224)
    • Area labels: none; also carries area:memory
    • Rationale: A race in Spark's ExecutionMemoryPool fails the task with NoSuchElementException instead of refusing the request so the operator can spill. It is shown at component level only, and the task fails visibly, so step 3.
  • Reduce retained buffer allocation when list_extract selects short nested arrays (#6225)
    • Area labels: area:expressions
    • Rationale: perf: optimize list_extract without defaults using Arrow take #5174, which is in the 1.1 branch, makes list_extract retain about 32 times the buffer capacity when it selects short inner arrays, and native shuffle charges that capacity against its reservation. The values are correct and there is no end-to-end measurement yet, so this is a memory regression at step 3.
  • Spark errors after the first batch of a JVM input stream reach the user as CometNativeException (#6234)
  • Grouped integer SUM doesn't count its per-group state in the aggregate's memory reservation (#6252)
    • Area labels: area:aggregation; also carries area:memory
    • Rationale: size() reports the struct rather than its per-group vectors, so every grouped integer SUM under-reserves by 16 bytes per group and spills late. The results are correct, so step 3.
  • Native window operators reserve no memory for the batches they buffer (#6253)
    • Area labels: none; also carries area:memory
    • Rationale: WindowAggExec runs Spark's default frame for agg(...) OVER (PARTITION BY k). It keeps every input batch of the task's partition with no reservation, where Spark's WindowExec spills, so a large partition can push the executor past its container limit. Not yet measured, so step 3.
  • A native final aggregate that has spilled can fail the task during its replay (#6254)
    • Area labels: area:aggregation; also carries area:memory
    • Rationale: The replay after a spill fails the task with Additional allocation failed instead of spilling again, and it reproduces. The failure is visible and falling back to Spark's aggregate avoids it, so step 3; see the escalations.
  • Hash-based JVM columnar shuffle reports its output as disk spill instead of bytes written (#6258)
  • Spark 3.4 and 3.5 release jars require Java 17 since 0.11.0, though the docs list Java 11 (#6283)
    • Area labels: area:ci; also carries build
    • Rationale: The published Spark 3.x jars fail to load on Java 11 with UnsupportedClassVersionError. The failure is visible and running on Java 17 avoids it, and the cause is in the release build tooling. Confirms the author's label.
  • Comet starts a single Tokio worker on standalone executors when spark.executor.cores is unset (#6292)
    • Area labels: none; also carries performance
    • Rationale: A standalone executor with spark.executor.cores unset, the standalone default, gets a single Tokio worker. In the reporter's measurement, a native scan feeding a sort ran 2.4 times slower with one worker than with four, which is significant performance degradation at step 3. Confirms the author's label; see the escalations for the deadlock.
  • A JVM consumer's parked page allocation fails when another consumer of the task empties its balance (#6304)
  • Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE (#6315)
    • Area labels: area:writer
    • Rationale: On Spark 3.4 and 3.5, a native INSERT INTO ... SELECT reads back empty until the table is refreshed, which is silent. But the native Parquet write needs both the Testing-category spark.comet.parquet.write.enabled and the DataWritingCommandExec allowIncompatible opt-in, so as with [Bug] Native make_interval overflows valid time components or loses seconds precision #5131, users only see it after accepting incompatible behavior. Confirms the author's label (see escalations).
  • Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval (#6313)

priority:low

  • native: JVMClasses::with_env: JAVA_VM not initialized aborts a multi-suite JVM (#6096)
    • Area labels: area:ffi
    • Rationale: The JNI lifecycle abort has only been seen when 132 Comet suites share one JVM with debug assertions on, and CI splits the suites and passes. That is a test-only failure at step 4 (see escalations).
  • Three dev/diffs weaken a local-shuffle-read assertion into a contradiction, breaking the Spark-only baseline (#6122)
    • Area labels: spark sql tests
    • Rationale: The patched test cannot pass for any number of local reads. It is skipped under Comet and breaks only the Spark-only baseline run, so this is test-only, step 4.
  • Native memory usage log underestimates the executor overhead for PySpark and SparkR on Kubernetes, and warns on standalone clusters (#6188)
    • Area labels: none; also carries area:memory
    • Rationale: The log warns of a container kill that cannot happen, and warns on standalone clusters that have no container. That is a misleading log line, step 4.
  • In-memory cache tests that use checkSparkAnswer compare the cache with itself (#6203)
    • Area labels: none; also carries test
    • Rationale: 23 cache tests cannot catch a wrong stored value or a wrong pruning bound, which is a test-only problem at step 4. Confirms the author's label; it bears on the default-on decision in feat: enable Comet's in-memory cache by default #5634.
  • With several native plans in one task, the non-zero memory usage warning comes from the wrong plan (#6255)
    • Area labels: none; also carries area:memory
    • Rationale: The first plan logs a leak that is not one, and a real leak in a later plan is never reported. Log diagnostics only, step 4.
  • spark.comet.shuffle.jvm.batchSize=0 hangs a task, and an unknown spark.comet.exec.memoryPool fails every task (#6259)
    • Area labels: area:shuffle; also carries area:memory
    • Rationale: Two settings are not validated, and the failures need an invalid value: a zero batch size, or a misspelled pool name that fails every task with a clear message. Step 4.
  • The native memory usage log understates untracked memory while a pool is overcommitted (#6260)
    • Area labels: none; also carries area:memory
    • Rationale: While a pool is overcommitted, the log's untracked figure and the container warning come out low by the overcommit. The author notes that is usually small and short-lived. Log diagnostics, step 4.
  • An input ArrowArrayStream that native never takes is never released (#6289)
    • Area labels: area:ffi
    • Rationale: The stream, and the first input batch it buffered, stay pinned for the executor's life only when a task ends before native takes the stream, for example a failed createPlan or an early cancellation. That is a leak on error paths, and confirms the author's label.

Enhancements

  • Support Spark's AtLeastNNonNulls natively (#6093)
    • Area labels: area:expressions
    • Rationale: New native expression support. Routing through the codegen dispatcher works but was measured slower than falling back.
  • ci: key the TPC dataset caches on pinned generators (#6102)
    • Area labels: area:ci
    • Rationale: CI cache efficiency; no job fails today.
  • ci: shard the Iceberg extensions test task (#6103)
    • Area labels: area:ci
    • Rationale: Shortens the longest job in the merge queue; no job fails today.
  • Iceberg S3-family storage rebuilds the opendal Operator and signer on every file open (#6109)
    • Area labels: area:scan
    • Rationale: A performance optimization that needs an upstream iceberg-rust change. Reads are correct today.
  • Native Iceberg writer keeps a dictionary page for high-cardinality columns where iceberg-java writes none (#6114)
    • Area labels: area:writer; also carries area:Iceberg, performance
    • Rationale: The issue states that results are correct either way. This is a file layout and read-performance improvement.
  • Charge native write buffers to the shared off-heap memory pool (#6115)
  • docs: add a user-facing Celeborn integration guide (#6118)
    • Area labels: area:shuffle; also carries documentation
    • Rationale: Documentation addition.
  • Write blog post for 1.1.0 release (#6120)
    • Area labels: none
    • Rationale: Release communication; no code change.
  • perf: reduce Parquet runtime filter schema guard overhead (#6123)
  • Investigate charging native Iceberg scan memory to the task memory pool (#6126)
  • Support scalar PySpark Arrow UDFs in Comet native execution via PyO3 (#6129)
    • Area labels: none
    • Rationale: A new opt-in native operator for Spark 4.1+ ArrowEvalPythonExec, which falls back correctly today.
  • Report resident memory (RssAnon) in the executor memory usage log and warn on it against the container size (#6167)
    • Area labels: none; also carries area:memory
    • Rationale: Observability addition to the memory usage log.
  • Follow-ups from the 1.1.0 user guide review (#6169)
  • Native UDFs: test coverage for empty batches and varied array encodings (#6174)
    • Area labels: area:ffi; also carries area:udf
    • Rationale: Test coverage for C Data Interface edge cases; no failure is reported.
  • Native UDFs: config to disable or restrict library loading (#6175)
    • Area labels: none; also carries area:udf
    • Rationale: A new config for operators to disable or restrict native UDF library loading.
  • Native UDFs: distribute the library to executors (#6176)
    • Area labels: none; also carries area:udf
    • Rationale: New capability to ship the native UDF library to executors.
  • Native UDFs: lift the 4-argument cap and add registerAll (#6177)
    • Area labels: none; also carries area:udf
    • Rationale: Extends the registration API: more than four arguments, and registerAll.
  • Follow-ups to chore: deprecate spark.comet.exec.memoryPool.fraction #6163: correct the fair_unified description and prepare spark.comet.exec.memoryPool.fraction for removal (#6187)
    • Area labels: none; also carries documentation, area:memory
    • Rationale: Documentation corrections and the remaining work for the config deprecation.
  • Kubernetes guide example never enables Comet because it sets no off-heap memory (#6190)
    • Area labels: none; also carries documentation
    • Rationale: A documentation correction, which the guide's type table places under enhancement.
  • ExtractANSIIntervalDays and the other interval field extractors have no serde, so date + <day interval column> and extract of an interval fall back (#6193)
    • Area labels: area:expressions
    • Rationale: New expression support; the fallback is correct today.
  • Charge the native shuffle writer's write buffers to the memory pool (#6196)
  • Release builds of libcomet skip LTO because the crate is also an rlib (#6210)
    • Area labels: area:ci; also carries build, performance
    • Rationale: A build configuration change: measure LTO for libcomet, then possibly enable it.
  • perf: account native allocations without a thread-local lookup (#6213)
    • Area labels: none; also carries performance, area:memory
    • Rationale: Performance optimization of the allocation accounting wrapper.
  • Follow-ups for the IRSA web-identity credential provider (#6231)
  • CometLiteral accepts a collated string literal and serializes it as a plain string (#6232)
  • Move CometCollationSuite to spark-4.x so every 4.x profile runs it (#6233)
    • Area labels: none
    • Rationale: Moves the collation suite so that the Spark 4.2 profile runs it too.
  • Support Native Iceberg MOR write (#6240)
    • Area labels: area:writer
    • Rationale: New native write capability.
  • CometTaskMemoryManager logs a warning and a memory dump every time a native reservation is refused (#6257)
    • Area labels: none; also carries area:memory
    • Rationale: Changes the log level of a routine event; no behavior change.
  • Remove AlignedArrowStreamReader and fix the Native to JVM section of ffi.md (#6291)
    • Area labels: area:ffi; also carries documentation
    • Rationale: Dead-code removal and a documentation rewrite, with no behavior change intended.
  • Blocking JVM calls from native plans hold Tokio workers, delaying other plans and I/O (#6293)
  • Cancelling a task doesn't stop its native plan until the plan produces its next batch (#6295)
    • Area labels: area:ffi; also carries performance
    • Rationale: Adds native cancellation so that a killed task frees its slot sooner. Results are unaffected, and the author filed it as an enhancement.
  • S3 credential refreshes are not coalesced, and bucket region detection has no timeout (#6296)
  • Support anonymous S3 access through the CometS3CredentialProvider SPI (#6298)
    • Area labels: area:scan
    • Rationale: New capability for the credential provider SPI. Today the adapter fails with a clear error for anonymous buckets.
  • Support Spark 4.2 geometry type + functions (#6317)
    • Area labels: area:expressions
    • Rationale: New type and function support for Spark 4.2's GEOMETRY and its initial ST_* functions.

Escalations to consider

Skipped (needs more info)

Both are tracking or record issues rather than bug reports or feature requests. No bug or enhancement label fits, so requires-triage was left in place, and they will reappear in each pass until they are closed.

  • [EPIC] Bug fixes to consider backporting to branch-1.0 (#6201)
  • Bug triage results: 2026-08-24 (#5454)
    • The summary issue from the 2026-08-24 pass. None of its labels were applied at the time, but all 28 issues it covers have since been triaged: each carries a type label and none still carries requires-triage. The three earlier summaries it links are closed. Nothing in it is outstanding, so it can be closed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions