Skip to content

fix: preserve and test reference layouts in GET_Dataset response - #469

Merged
mattjala merged 2 commits into
HDFGroup:masterfrom
mattjala:fix/preserve-reference-layout-with-dn-tests
Sep 23, 2026
Merged

mattjala merged 2 commits into
HDFGroup:masterfrom
mattjala:fix/preserve-reference-layout-with-dn-tests

Conversation

@mattjala

Copy link
Copy Markdown
Collaborator

Fork of #468 which adds a test, and fixes the unit test registration.

testall.py registers dset_dn_test in unit_tests, which otherwise would
never run it, and reflows that tuple to stay under the 99 character
flake8 limit.

joaopaulosr95 and others added 2 commits September 18, 2026 12:49
Fork of HDFGroup#468 which adds a test, and fixes the unit test registration.

testall.py registers dset_dn_test in unit_tests, which otherwise would
never run it, and reflows that tuple to stay under the 99 character
flake8 limit.
@github-project-automation github-project-automation Bot moved this to To be triaged in HSDS - TRIAGE & TRACK Sep 22, 2026
@mattjala
mattjala merged commit 9046150 into HDFGroup:master Sep 23, 2026
27 checks passed
@github-project-automation github-project-automation Bot moved this from To be triaged to Done in HSDS - TRIAGE & TRACK Sep 23, 2026
d-gski added a commit to d-gski/hsds that referenced this pull request Sep 29, 2026
Since 1.0, the chunk shape comes from h5json's getChunkDims, which returns the
full dataset shape for any layout that isn't H5D_CHUNKED*. For a contiguous
reference that makes the whole dataset one chunk: reading a single element
makes the DN range-get, decode and cache the entire dataset from the linked
file.

This regressed in b6016e0 ("use hdf5-json util classes", HDFGroup#450). Through 0.9.x,
POST_Dataset gave every contiguous reference a virtual chunk shape via
getContiguousLayout, and testContiguousRefDataset required it (HDFGroup#393/HDFGroup#396 fixed
its short last chunk on an NREL `meta` dataset). b6016e0 moved layouts under
creationProperties, removed that computation, and switched to h5json's
getChunkDims. h5json's rule holds for data HSDS stores itself, since
generateLayout only picks H5D_CONTIGUOUS below chunk_min, but not for a
reference, whose size comes from the linked file. The read path still assumes
virtual chunks (per-chunk offsets in getChunkLocations, the DN's short-chunk
padding). The reference-layout bug fixed by HDFGroup#468/HDFGroup#469 hid the regression for
domains linked by 0.9.x.

Seen on NREL's NSRDB TMY domains (nrel-pds-hsds). Each `meta` dataset is an
H5D_CONTIGUOUS_REF (MSG: 2,693,287 x 134 B = 361 MB). One-row reads fetched
the whole 361 MB (6-34 s from S3). The two such datasets no longer fit a 512m
chunk cache and evicted each other. Concurrent requests during a fetch each
started another full copy.

Add getChunkDims to hsds.util.dsetUtil, used everywhere in place of h5json's.
It defers to h5json except for H5D_CONTIGUOUS_REF. There it derives virtual
chunks with getContiguousLayout from the dataset's shape, item size and the
min/max_chunk_size config. That is the same split 0.9.x made: 21042 rows for
the MSG `meta` above, as stored with the domain.

Nothing is stored per chunk; each is a range get into the file. So the shape
only has to agree between the SN (chunk index -> byte range) and the DN
(decode + cache). Deriving it from shared inputs gives that on both, so
GET_Dataset keeps reporting the reference layout without dims, as HDFGroup#468/HDFGroup#469
intend.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants