Agents & InferenceHugging Face

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Olmo-core 3 scaled an MoE expert pool from 8 to 128 with four experts active per token, growing total capacity from 4.6B to 47B while losing under 5% training throughput, and has been benchmarked past one trillion total parameters. The key engineering change is moving from FSDP weight gather/reshard to DDP with experts resident on GPUs and token routing to experts, which makes open MoE training infrastructure more practical if you’re trying to add parameter capacity without blowing up per-token compute.