Conversation
Fork of HDFGroup#468 which adds a test, and fixes the unit test registration. testall.py registers dset_dn_test in unit_tests, which otherwise would never run it, and reflows that tuple to stay under the 99 character flake8 limit.
d-gski
added a commit
to d-gski/hsds
that referenced
this pull request
Sep 29, 2026
Since 1.0, the chunk shape comes from h5json's getChunkDims, which returns the full dataset shape for any layout that isn't H5D_CHUNKED*. For a contiguous reference that makes the whole dataset one chunk: reading a single element makes the DN range-get, decode and cache the entire dataset from the linked file. This regressed in b6016e0 ("use hdf5-json util classes", HDFGroup#450). Through 0.9.x, POST_Dataset gave every contiguous reference a virtual chunk shape via getContiguousLayout, and testContiguousRefDataset required it (HDFGroup#393/HDFGroup#396 fixed its short last chunk on an NREL `meta` dataset). b6016e0 moved layouts under creationProperties, removed that computation, and switched to h5json's getChunkDims. h5json's rule holds for data HSDS stores itself, since generateLayout only picks H5D_CONTIGUOUS below chunk_min, but not for a reference, whose size comes from the linked file. The read path still assumes virtual chunks (per-chunk offsets in getChunkLocations, the DN's short-chunk padding). The reference-layout bug fixed by HDFGroup#468/HDFGroup#469 hid the regression for domains linked by 0.9.x. Seen on NREL's NSRDB TMY domains (nrel-pds-hsds). Each `meta` dataset is an H5D_CONTIGUOUS_REF (MSG: 2,693,287 x 134 B = 361 MB). One-row reads fetched the whole 361 MB (6-34 s from S3). The two such datasets no longer fit a 512m chunk cache and evicted each other. Concurrent requests during a fetch each started another full copy. Add getChunkDims to hsds.util.dsetUtil, used everywhere in place of h5json's. It defers to h5json except for H5D_CONTIGUOUS_REF. There it derives virtual chunks with getContiguousLayout from the dataset's shape, item size and the min/max_chunk_size config. That is the same split 0.9.x made: 21042 rows for the MSG `meta` above, as stored with the domain. Nothing is stored per chunk; each is a range get into the file. So the shape only has to agree between the SN (chunk index -> byte range) and the DN (decode + cache). Deriving it from shared inputs gives that on both, so GET_Dataset keeps reporting the reference layout without dims, as HDFGroup#468/HDFGroup#469 intend.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fork of #468 which adds a test, and fixes the unit test registration.
testall.py registers dset_dn_test in unit_tests, which otherwise would
never run it, and reflows that tuple to stay under the 99 character
flake8 limit.