r/LocalLLaMA • u/Prestigious-Taste-63 • 14h ago
I Built A Thing I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
https://huggingface.co/YOON1v/Apex-2First of all, thank you for reading.
I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons.
Apex-2
- Architecture: Decoder-only MoE, every layer is MoE (no dense layers)
- Size: 3.87B total parameters, 1.45B active per token
- 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4
- Context: 4096
- Tokenizer: Qwen3 (151k)
- Hugging Face: https://huggingface.co/YOON1v/Apex-2
(loads with transformers / vLLM via Qwen3MoeForCausalLM mapping)
Training
- Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo)
- SFT: ~2.5B tokens (code-heavy + math + instruction)
- DPO: tried it, scores dropped, so I dropped the checkpoint
Key numbers (SFT, greedy, chat template)
Benchmark
HumanEval 43.9
HumanEval+ 41.5
MBPP 56.3
MBPP+ 48.9
GSM8K (0-shot CoT) 32.4
MATH-500 21.0
IFEval (prompt strict) 44.7
MMLU (5-shot) 28.6
interesting comparison
With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T).
Knowledge (MMLU) and math still lag far behind, as expected with the data gap.
What didn’t work
DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval.
I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up.
Limitations (honest)
- English-centric (almost no multilingual ability)
- Weak knowledge → frequent hallucinations
- LiveCodeBench medium/hard is near zero
- 4k context only
Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.
7
u/FullOf_Bad_Ideas 12h ago
That's really cool. I have a similar project, a 4B A1.15B MoE trained from scratch only on Polish data. You have good results. How much compute did it take to get there, what was the throughput like? Since you reuse Qwen 3 tokenizer, you can do off-policy or on-policy distillation from other Qwen 3 models that use this tokenizer. That's what I plan to do in my project with models that share APT4 tokenizer that I'm using.
"max_position_embeddings": 4096, "rope_theta": 10000.0,
That's a wrong theta for 4096 context, read https://arxiv.org/abs/2405.14591v1
10k is good enough only for about 1500 tokens.
6
u/wizard_of_menlo_park 13h ago edited 13h ago
Great work! Training is hard.
What was the total cost of hardware to train the model? How long did it take? Did you train this from scratch ?
2
u/Healthy_Dependent854 13h ago
Maybe you could make Reinforced learning with verifiable rewards instead DPO? I heard from a paper that made it pretty cheap https://arxiv.org/pdf/2605.06241
DISCLAIMER: I never tried the paper myself so I do not know if it actually works or if its just benchmaxed or something. I am just giving something that might be helpfull to try
3
u/FullOf_Bad_Ideas 12h ago
RL is very ineffective and expensive, especially for small models.
1
u/Healthy_Dependent854 12h ago
I know thats a paper that at least claims that its possible in about 20 questions -but as I said its just telling what I thought it might be helpfull I am explicity not shure if it really works and do not want to make claims I cant proof.
1
1
u/sn2006gy 11h ago
Love seeing this. Curious if you built a dense model to compare against? Then I wonder if you could take the delta and use what you learn to optimize how the MoE could be improved.
1
u/toothpastespiders 5h ago
I'd be most curious about what you've learned along the way and what you wish you'd known in advance. Along with any oddities you've found working through this as a MoE rather than dense.
0
0
15
u/MehTheHedgehog 13h ago
It encouraging seeing other people trying to build the SML from scratch, not another finetune of Qwen. Starred to check out in the future.
Do you plan to further post train it with Self-distillation as this is what interest me the most.