โ† Back to live feed

Wednesday, Sep 23, 2026

1
Scale AI and CAIS Release HLE Diamond Benchmark With 60.6% Top ScorePVT:SCAI
topics ๐Ÿค– AI tags AIAI ModelsAI Research PVT:SCAI keywords ScaleScale AICAIS

Evaluation of the most advanced artificial intelligence models now has a more reliable tool following the debut of a curated testing set based on the Humanity's Last Exam. The HLE-Diamond subset was developed by Scale AI and the Center for AI Safety (CAIS) after one year of cleaning and refinement with assistance from research communities. The highest scoring model tested achieved a mark of 60.6% overall.

The project focused on improving reliability for measuring frontier models by incorporating extensive community feedback. Scale AI expects the HLE-Diamond results to provide a useful signal for measuring model progress over the next 6 to 12 months.

See all 2 tweets โ†’