Skip to content

Enable spark.comet.exec.localTableScan.enabled when running Spark SQL tests #4347

Description

@mbutrovich

Related to #2723.

A lot of Spark SQL tests don't put their source data in a format that Comet accelerates reading from. Comet generally wants data in Parquet or Iceberg. A number of Spark SQL suites (e.g., UDFSuite) don't write to Parquet so their scans are just LocalTableScanExec, and Comet is not exercising UDF compatibility because the scan underneath isn't native. We can either do what #2723 suggests and enable converting to columnar, or try enabling LocalTableScanExec now that we have #2735. @andygrove tried the former in #2714, but the logs are gone now so I'm not sure how ugly it was. Starting with LocalTableScanExec might be a smaller blast radius?


Tracking

Enabling spark.comet.exec.localTableScan.enabled by default is being tried in #4393 (draft) and #6780 (draft, run against Spark 4.1). Remaining work, by area:

Test expectations that assume Spark nodes (skip list or per-test config)

Known Comet gaps newly exposed by the change

Iceberg

Done: #4789, #4786, #4787, #4788, plus the NullType and TimeType handling from #4393.

Activity

  1. added this to the 1.0.0 milestone on May 18, 2026
  2. added
    priority:lowMinor issues, test failures, tooling, cosmetic
    and removed on May 18, 2026
  3. modified the milestones: 1.0.0, 1.1.0 on Aug 17, 2026
  4. modified the milestones: 1.1.0, 1.2.0 on Oct 8, 2026
  5. andygrove commented on Oct 8, 2026

    @andygrove
    Member

    @mbutrovich heads-up: I appended a Tracking section below your description (your text is unchanged) and filed sub-issues #6795-#6803 from a fresh run of the default flip on current main (#6780, Spark 4.1 only so far). They cover the remaining plan-shape test failures and the gaps in your B3a/B5/B7 buckets that still reproduce. Happy to move or reshape the list if you would prefer it differently.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:lowMinor issues, test failures, tooling, cosmeticspark sql testsSpark SQL test failures

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions