What is the problem the feature request solves?
With #6106 every task on an executor shares one FileIO per storage configuration. What that sharing buys depends on what the backend's Storage holds. At the pinned iceberg-rust rev (665c64e), OpenDalStorage::S3, Gcs and Oss hold only the parsed config and, for S3, the access loader. create_operator in iceberg-storage-opendal builds a new opendal Operator on every new_input, new_output, exists and metadata, and opendal-service-s3's build() constructs a new Signer for each operator. The signer is where reqsign keeps the access cache, so with a custom S3 access provider every file open still pays a JNI provide_credential round trip, and every operator re-parses and re-validates its configuration. HTTP connection pooling is unaffected either way: opendal's default transport is process-wide.
For the S3 family, then, #6106 reuses the factory, the config and the bridge, but not the client or signer. The HDFS backend in #5898 already solves this on its side: OpenDalStorage::HdfsNative carries an operator cache keyed by NameNode, so the hdfs-native client is built once per Storage and, with #6106, once per executor.
Describe the potential solution
Apply the HDFS backend's pattern to the S3 family in iceberg-storage-opendal: an operator cache inside OpenDalStorage::S3 (and Gcs, Oss) keyed by the operator's root, so create_operator returns an existing Operator for a bucket it has already built. Since Storage lives as long as the cached FileIO, the signer and its access cache would then be shared by every task on the executor, and the per-open JNI fetch collapses to one per access expiry.
This is an upstream change in apache/iceberg-rust, followed by a pin bump here. Comet cannot do it locally: create_operator is pub(crate) and the Storage trait exposes no operators.
Additional context
Split out of #6105 at review time so it stays open after #6106 merges. The operator cache is only effective once a FileIO outlives a task, which is what #6106 provides.
What is the problem the feature request solves?
With #6106 every task on an executor shares one
FileIOper storage configuration. What that sharing buys depends on what the backend'sStorageholds. At the pinned iceberg-rust rev (665c64e),OpenDalStorage::S3,GcsandOsshold only the parsed config and, for S3, the access loader.create_operatoriniceberg-storage-opendalbuilds a new opendalOperatoron everynew_input,new_output,existsandmetadata, andopendal-service-s3'sbuild()constructs a newSignerfor each operator. The signer is where reqsign keeps the access cache, so with a custom S3 access provider every file open still pays a JNIprovide_credentialround trip, and every operator re-parses and re-validates its configuration. HTTP connection pooling is unaffected either way: opendal's default transport is process-wide.For the S3 family, then, #6106 reuses the factory, the config and the bridge, but not the client or signer. The HDFS backend in #5898 already solves this on its side:
OpenDalStorage::HdfsNativecarries an operator cache keyed by NameNode, so the hdfs-native client is built once perStorageand, with #6106, once per executor.Describe the potential solution
Apply the HDFS backend's pattern to the S3 family in
iceberg-storage-opendal: an operator cache insideOpenDalStorage::S3(andGcs,Oss) keyed by the operator's root, socreate_operatorreturns an existingOperatorfor a bucket it has already built. SinceStoragelives as long as the cachedFileIO, the signer and its access cache would then be shared by every task on the executor, and the per-open JNI fetch collapses to one per access expiry.This is an upstream change in
apache/iceberg-rust, followed by a pin bump here. Comet cannot do it locally:create_operatorispub(crate)and theStoragetrait exposes no operators.Additional context
Split out of #6105 at review time so it stays open after #6106 merges. The operator cache is only effective once a
FileIOoutlives a task, which is what #6106 provides.