Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
Olmo-core 3 scaled an MoE expert pool from 8 to 128 with four experts active per token, growing total capacity from 4.6B to 47B while losing under 5% training throughput, and has been benchmarked past one trillion total parameters. The key engineering change is moving from FSDP weight gather/reshard to DDP with experts resident on GPUs and token routing to experts, which makes open MoE training infrastructure more practical if you’re trying to add parameter capacity without blowing up per-token compute.
Mistral Large quota or rate limit — check usage and plan. Original headline: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs