You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Enable spark.comet.exec.localTableScan.enabled when running Spark SQL tests #4347
A lot of Spark SQL tests don't put their source data in a format that Comet accelerates reading from. Comet generally wants data in Parquet or Iceberg. A number of Spark SQL suites (e.g., UDFSuite) don't write to Parquet so their scans are just LocalTableScanExec, and Comet is not exercising UDF compatibility because the scan underneath isn't native. We can either do what #2723 suggests and enable converting to columnar, or try enabling LocalTableScanExec now that we have #2735. @andygrove tried the former in #2714, but the logs are gone now so I'm not sure how ugly it was. Starting with LocalTableScanExec might be a smaller blast radius?
Tracking
Enabling spark.comet.exec.localTableScan.enabled by default is being tried in #4393 (draft) and #6780 (draft, run against Spark 4.1). Remaining work, by area:
Test expectations that assume Spark nodes (skip list or per-test config)
@mbutrovich heads-up: I appended a Tracking section below your description (your text is unchanged) and filed sub-issues #6795-#6803 from a fresh run of the default flip on current main (#6780, Spark 4.1 only so far). They cover the remaining plan-shape test failures and the gaps in your B3a/B5/B7 buckets that still reproduce. Happy to move or reshape the list if you would prefer it differently.
Related to #2723.
A lot of Spark SQL tests don't put their source data in a format that Comet accelerates reading from. Comet generally wants data in Parquet or Iceberg. A number of Spark SQL suites (e.g., UDFSuite) don't write to Parquet so their scans are just
LocalTableScanExec, and Comet is not exercising UDF compatibility because the scan underneath isn't native. We can either do what #2723 suggests and enable converting to columnar, or try enablingLocalTableScanExecnow that we have #2735. @andygrove tried the former in #2714, but the logs are gone now so I'm not sure how ugly it was. Starting withLocalTableScanExecmight be a smaller blast radius?Tracking
Enabling
spark.comet.exec.localTableScan.enabledby default is being tried in #4393 (draft) and #6780 (draft, run against Spark 4.1). Remaining work, by area:Test expectations that assume Spark nodes (skip list or per-test config)
LocalTableScanExecis a Comet operatorKnown Comet gaps newly exposed by the change
collect_set/collect_listelement order differs from Sparkasinh,cosh,tan,cot,atanh,cbrt,exp,pow,atan2)SparkExceptionwhere Spark raises a specific error class (to_time,to_binary,array_insert,mapconstruction)percentilereturns0.0instead of-0.0from_csvreturns wrong results in permissive modes and fails with a variant schemaGROUP BYon a struct containing a map fails withCometNativeExceptionfield name cannot be null)Iceberg
CometIcebergWriteDetectionSuiteexpects the split Iceberg write plan but gets a fully native write overCometLocalTableScan(related to feat: enable native Iceberg writes by default #6664)Done: #4789, #4786, #4787, #4788, plus the
NullTypeandTimeTypehandling from #4393.