Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose and result
Implements the first-stage MoE experiment: two compute cores with independent normal SRAM buffers and accumulators, sharing one HBM system. The comparison holds total multiplier count and SRAM budgets constant.
Numerical execution passes, but the best tested large/small configuration remains slower than the optimized single core. This is a draft for review, not ready to merge.
The architecture refinement proposal specifies the remaining task-partitioning, SRAM-port, accumulation, DMA and dispatch contracts. It is proposed, not implemented; the timings below describe the current model.
Companion: Compiler #80. The Compiler submodule is pinned to that PR's commit.
Implementation
Dimensions and resource budgets
M = N = BLEN,K = MLEN. Tiles execute over multiple cycles; multiplier count isBLEN × MLEN.(M, K, N)(8, 512, 8)(4, 768, 4)(2, 512, 2)HBM weights are stored by output row. The local E4M3/E8M0 format uses one scale per eight elements. Full matrix dimensions
D/Fare distinct from the core'sMLENtile width.Measurement and existing DSE results
Latency starts with activations and routes ready and ends when the final BF16 output is ready. Ramulator times weight-memory requests; explicit timing models cover compute and on-chip transfers. Router execution, input/output HBM transfers and full-model execution are outside this measurement boundary.
The completed DSE used archived routes and synthetic nonzero weights/activations, with
D=2048,F=512/1408and 8/32-token windows:The activation supply split is fixed. Under the current service model, the small core needs four cycles per tile instead of two with sufficient operand supply. The kernel also waits for full result readiness between updates to the same accumulator and does not interleave independent output tiles. These are specific port/scheduling assumptions to review; the measured slowdown cannot be attributed solely to the HBM controller.
The PR retains eight paired timings from that experiment. Historical reports, DSE/ablation drivers and large generated datasets are excluded.
Validation on this review branch
D=31, F=47case covers tails and a shared expert. Each configuration ran twice; both cores received work, and BF16 output matched the independent oracle bit for bit.See the review guide and run commands. The entry point is
transactional_emulator/testbench/moe_normal_review/smoke.py.Scope is the standalone Compiler-manifest → Rust experiment. Transpose buffers, Attention and complete dual-core instruction lowering are not included. Review should focus on scheduling dependencies, SRAM/port constraints and DMA timing.