Thursday, Oct 1, 2026
1 Allen AI Releases Open Training Stack for 1 Trillion Parameter MoE Models 🤖 AI Oct 1, 11:20 AM EDT 4/3
Olmo-core 3 serves as the structural foundation for the next generation of AI models developed by Allen AI. The newly released training infrastructure for mixture-of-experts (MoE) architectures enables these systems to scale up to 1 trillion parameters, providing an open-source framework for creating massive model architectures.
The system's primary technical shift replaces FSDP weight gather and resharding with DDP and GPU-resident experts. This update allows the expert pool to increase from 8 to 128, growing from 4.6B to 47B total parameters with approximately 3.2B active. By utilizing rowwise expert parallelism and grouped GEMM, the stack increases performance to 2.7x tokens/s/GPU compared to the previous version while maintaining throughput loss below 5%.