Phase 7.1+7.3: Cache-line placement of bytes_until_sample_ + lazy provider overhead validation - #19
Merged
Conversation
…rhead
- Phase 7.1: hoist bytes_until_sample into a dedicated alignas(64/128)
SamplerHotState struct (128 bytes on Apple Silicon, 64 elsewhere) so the
per-thread fast-path counter sits on its own cache line and cannot
false-share with the colder Sampler tail (PRNG state, last_sample_,
initialized_) or with concurrent dealloc slot-clear traffic. Counter is
the first member of the cache-aligned region (offset 0). Adds a
SNMALLOC_LIKELY annotation on the hot subtract+compare.
- Phase 7.3: new func test profile_overhead asserting
a) sizeof(Config::PagemapEntry) is unchanged vs. an explicit
StandardConfigClientMeta<NoClientMetaDataProvider> — proves the
lazy provider type is compiled in but contributes zero bytes when
profiling is off.
b) bytes_until_sample lives at offset 0 of the cache-aligned hot
state (offsetof check).
c) Runtime gate: 1M alloc/free pairs of size 32 under
Sampler::set_sampling_rate(0) (off) and Sampler::set_sampling_rate
(2^40) (on, never fires) — assert ns/alloc ratio < 1.05, i.e. no
branch-misprediction storm in the dealloc null-slot fast-path.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Phase 7.1 — Hoist the per-thread Poisson sampler hot counter into its own cache line.
SNMALLOC_CACHE_LINE_SIZEmacro (128 on Apple Silicon via__APPLE__ + __aarch64__, 64 elsewhere)Sampler::SamplerHotStatenested struct,alignas(SNMALLOC_CACHE_LINE_SIZE), withint64_t bytes_until_sampleat offset 0Sampler→ TLS sampler's hot counter occupies its own cache line, no false-sharing with concurrent dealloc clears on adjacent fieldsSNMALLOC_LIKELYalready present on the fast-path branch (annotated for clarity)Phase 7.3 — Compile-time + runtime validation that lazy ClientMetaDataProvider adds zero overhead when profiling inactive.
src/test/func/profile_overhead/profile_overhead.cc(auto-discovered by CMake subdirlist)static_asserts:sizeof(LazyArrayClientMetaDataProvider<T>::StorageType) == sizeof(void*)sizeof(NoClientMetaDataProvider::StorageType) == sizeof(Empty)sizeof(snmalloc::Config::PagemapEntry) == sizeof(StandardConfigClientMeta<NoClientMetaDataProvider>::PagemapEntry)— proves lazy provider compiled in but NOT instantiated into default config's metadataSampler::kBytesUntilSampleOffset == 0Test plan