← Back to live feed

Thursday, Oct 1, 2026

1
Allen AI Releases Open Training Stack for 1 Trillion Parameter MoE Models

Olmo-core 3 serves as the structural foundation for the next generation of AI models developed by Allen AI. The newly released training infrastructure for mixture-of-experts (MoE) architectures enables these systems to scale up to 1 trillion parameters, providing an open-source framework for creating massive model architectures.

The system's primary technical shift replaces FSDP weight gather and resharding with DDP and GPU-resident experts. This update allows the expert pool to increase from 8 to 128, growing from 4.6B to 47B total parameters with approximately 3.2B active. By utilizing rowwise expert parallelism and grouped GEMM, the stack increases performance to 2.7x tokens/s/GPU compared to the previous version while maintaining throughput loss below 5%.

See all 4 tweets →