Skip to content

[Review] MoE normal cores, bounded HBM DMA and fair single-core comparison - #117

Draft
Happymic wants to merge 5 commits into
mainfrom
review/moe-dual-normal-20260905
Draft

Happymic wants to merge 5 commits into
mainfrom
review/moe-dual-normal-20260905

Conversation

@Happymic

@Happymic Happymic commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose and result

Implements the first-stage MoE experiment: two compute cores with independent normal SRAM buffers and accumulators, sharing one HBM system. The comparison holds total multiplier count and SRAM budgets constant.

Numerical execution passes, but the best tested large/small configuration remains slower than the optimized single core. This is a draft for review, not ready to merge.

The architecture refinement proposal specifies the remaining task-partitioning, SRAM-port, accumulation, DMA and dispatch contracts. It is proposed, not implemented; the timings below describe the current model.

Companion: Compiler #80. The Compiler submodule is pinned to that PR's commit.

Implementation

  • Consume fixed-route Compiler manifests and schedule expert jobs through gather → gate/up → SwiGLU → down → weighted route combine.
  • Model independent per-core storage and bounded weight-prefetch slots. Compute waits for element/scale transfers, copying and actual decoding to finish.
  • Use one shared Ramulator HBM2 backend, with per-channel DMA submission, in-flight sector merging, bounded concurrency and SRAM-capacity checks.
  • Account for decoding, copying, operand supply and accumulation dependencies. Native HBM transactions are calibrated to 32 bytes; HBM channels, bandwidth, timing and the underlying DRAM command scheduler remain consistent across configurations.

Dimensions and resource budgets

M = N = BLEN, K = MLEN. Tiles execute over multiple cycles; multiplier count is BLEN × MLEN.

Configuration Single-core baseline Large core Small core
(M, K, N) (8, 512, 8) (4, 768, 4) (2, 512, 2)
Multipliers 4096 3072 1024
Weight SRAM / prefetch slots 64 KiB / 4 48 KiB / 4 16 KiB / 4
Activation SRAM / accumulator 4 MiB / 1 MiB 2 MiB / 512 KiB 2 MiB / 512 KiB
BF16 activation elements per cycle 1024 768 256

HBM weights are stored by output row. The local E4M3/E8M0 format uses one scale per eight elements. Full matrix dimensions D/F are distinct from the core's MLEN tile width.

Measurement and existing DSE results

Latency starts with activations and routes ready and ends when the final BF16 output is ready. Ramulator times weight-memory requests; explicit timing models cover compute and on-chip transfers. Router execution, input/output HBM transfers and full-model execution are outside this measurement boundary.

The completed DSE used archived routes and synthetic nonzero weights/activations, with D=2048, F=512/1408 and 8/32-token windows:

Evaluation set Large/small latency relative to optimized single core
Four tuning inputs 8.88% higher, geometric mean
Four additional validation windows 8.53% higher, geometric mean

The activation supply split is fixed. Under the current service model, the small core needs four cycles per tile instead of two with sufficient operand supply. The kernel also waits for full result readiness between updates to the same accumulator and does not interleave independent output tiles. These are specific port/scheduling assumptions to review; the measured slowdown cannot be attributed solely to the HBM controller.

The PR retains eight paired timings from that experiment. Historical reports, DSE/ablation drivers and large generated datasets are excluded.

Validation on this review branch

  • Compiler export, Rust library/MoE and Python result-validation tests pass, together with formatting and strict Clippy checks.
  • A nonzero D=31, F=47 case covers tails and a shared expert. Each configuration ran twice; both cores received work, and BF16 output matched the independent oracle bit for bit.
  • The full-dimension Qwen B8 case ran twice per configuration. Numerical, resource, HBM-accounting and repeatability checks passed, reproducing 0.524575 ms single-core / 0.564819 ms large/small. This rerun does not repeat the full DSE.

See the review guide and run commands. The entry point is transactional_emulator/testbench/moe_normal_review/smoke.py.

Scope is the standalone Compiler-manifest → Rust experiment. Transpose buffers, Attention and complete dual-core instruction lowering are not included. Review should focus on scheduling dependencies, SRAM/port constraints and DMA timing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant