matrix-multiply
Notes of semi-random semi-useless experiments
- Benchmarking Trillion-Parameter Open Models Without Fooling Yourself
Kimi K3 against DeepSeek V4 Pro and Flash on the same harness and hardware. The scores are the less interesting half; the ways they were nearly wrong are the rest.
- 896 Experts, None Dead: What MoE Load Actually Looks Like at Token Level
Per-expert routing histograms from a 1.5 TB MoE: the hottest expert runs 22x uniform load, no expert is dead, and one-expert-per-chip would strand most of your silicon.
- The Free Lunch That Wasn't (Part 2: Ten Seeds Later)
Part 1 found +5.5 points of free HumanEval from wider expert routing. Instrumenting the router and rerunning with ten seeds per arm turned that gain into noise — and one benchmark into a significant regression.
- Does Kimi K3 Want More Experts? (Part 1: The Tempting Result)
Kimi K3 activates 16 of 896 experts per token. You can change that at inference time. Three runs on eight B300s produced a +5.5 point HumanEval gain — and every reason not to believe it yet.