feat: Support lazy per-file Parquet Arrow schema derivation - #25343
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #25343 +/- ##
==========================================
+ Coverage 81.93% 81.95% +0.01%
==========================================
Files 1136 1136
Lines 429152 429482 +330
Branches 429152 429482 +330
==========================================
+ Hits 351633 351964 +331
+ Misses 56475 56471 -4
- Partials 21044 21047 +3 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
sunchao
left a comment
There was a problem hiding this comment.
Thanks @peterxcli! LGTM. The hook runs at the right point for per-file schema derivation, preserves the original metadata and decryption state, and respects explicit file-schema precedence. Provider-bearing sources correctly require an extension codec for DataFusion protobuf serialization.
Reviewed at a14db51. Local validation passed all 264 Parquet datasource unit tests with encryption enabled, the protobuf serialization rejection test, and an independent 20-scan DataSourceExec oracle covering nested LIST/MAP ENUM fields, invalid-UTF8 binary data, differing file layouts, view-type coercions, projections, partition values, and virtual row numbers.
The full workspace suite and end-to-end Comet projected Variant workflow were not run locally.
|
@sunchao thanks! |
Which issue does this PR close?
Closes #25251.
Rationale for this change
Readers such as Comet need each file's physical Parquet schema to choose Arrow types for nested fields, but that schema is only available after loading the footer.
What changes are included in this PR?
Add an optional
ParquetFileSchemaProvidertoParquetSource, invoked after footer loading and before Arrow schema inference. It supplies a complete file schema through Arrow's existingwith_schemaAPI while retaining the original metadata. ExplicitPartitionedFile.arrow_schematakes precedence.What is the testing strategy for this PR?
Regression tests cover differing nested layouts, ENUM/BINARY handling, cached and encrypted scans, schema precedence and validation, and serialization. Extended workspace tests and lint checks pass.
Are there any user-facing changes?
Adds
ParquetSource::with_schema_provider. Sources using a provider require a custom codec for plan serialization. Default scans are unchanged.