Skip to content

Phase 7.1+7.3: Cache-line placement of bytes_until_sample_ + lazy provider overhead validation - #19

Merged
jayakasadev merged 1 commit into
mainfrom
feature/phase-7-1-7-3-perf-tweaks
Jun 11, 2026
Merged

jayakasadev merged 1 commit into
mainfrom
feature/phase-7-1-7-3-perf-tweaks

Conversation

@jayakasadev

Copy link
Copy Markdown
Owner

Summary

Phase 7.1 — Hoist the per-thread Poisson sampler hot counter into its own cache line.

  • SNMALLOC_CACHE_LINE_SIZE macro (128 on Apple Silicon via __APPLE__ + __aarch64__, 64 elsewhere)
  • New Sampler::SamplerHotState nested struct, alignas(SNMALLOC_CACHE_LINE_SIZE), with int64_t bytes_until_sample at offset 0
  • First non-static member of Sampler → TLS sampler's hot counter occupies its own cache line, no false-sharing with concurrent dealloc clears on adjacent fields
  • SNMALLOC_LIKELY already present on the fast-path branch (annotated for clarity)

Phase 7.3 — Compile-time + runtime validation that lazy ClientMetaDataProvider adds zero overhead when profiling inactive.

  • New test src/test/func/profile_overhead/profile_overhead.cc (auto-discovered by CMake subdirlist)
  • static_asserts:
    • sizeof(LazyArrayClientMetaDataProvider<T>::StorageType) == sizeof(void*)
    • sizeof(NoClientMetaDataProvider::StorageType) == sizeof(Empty)
    • sizeof(snmalloc::Config::PagemapEntry) == sizeof(StandardConfigClientMeta<NoClientMetaDataProvider>::PagemapEntry) — proves lazy provider compiled in but NOT instantiated into default config's metadata
    • Sampler::kBytesUntilSampleOffset == 0
  • Runtime: 1M × 32B alloc microbench, profile-off vs profile-on-but-inactive; asserts < 5% overhead ratio

Test plan

  • func-profile_overhead-fast/check (SNMALLOC_PROFILE=ON)
  • func-profile_overhead-fast (SNMALLOC_PROFILE=OFF, layout-only smoke)
  • All 14 func-profile_* tests
  • func-memory, func-statistics, func-first_operation, func-teardown, func-bits, func-memory_usage, func-external_pointer, func-large_alloc
  • cargo test --release (snmalloc-rs, no features + --features profiling)

…rhead

- Phase 7.1: hoist bytes_until_sample into a dedicated alignas(64/128)
  SamplerHotState struct (128 bytes on Apple Silicon, 64 elsewhere) so the
  per-thread fast-path counter sits on its own cache line and cannot
  false-share with the colder Sampler tail (PRNG state, last_sample_,
  initialized_) or with concurrent dealloc slot-clear traffic.  Counter is
  the first member of the cache-aligned region (offset 0).  Adds a
  SNMALLOC_LIKELY annotation on the hot subtract+compare.
- Phase 7.3: new func test profile_overhead asserting
    a) sizeof(Config::PagemapEntry) is unchanged vs. an explicit
       StandardConfigClientMeta<NoClientMetaDataProvider> — proves the
       lazy provider type is compiled in but contributes zero bytes when
       profiling is off.
    b) bytes_until_sample lives at offset 0 of the cache-aligned hot
       state (offsetof check).
    c) Runtime gate: 1M alloc/free pairs of size 32 under
       Sampler::set_sampling_rate(0) (off) and Sampler::set_sampling_rate
       (2^40) (on, never fires) — assert ns/alloc ratio < 1.05, i.e. no
       branch-misprediction storm in the dealloc null-slot fast-path.
@jayakasadev
jayakasadev merged commit a1583f0 into main Jun 11, 2026
156 of 202 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant