Friday, Sep 18, 2026
1 GPT 6 Astra Outperforms Rivals in First WeirdML v3 Agentic Benchmark π€ AI Sep 18, 6:20 AM EDT 11/6
The newest version of a machine learning evaluation suite favors OpenAI's latest model, particularly in token efficiency. WeirdML v3 consists of 11 complex hand-made tasks requiring models to explore unfamiliar data and develop machine learning pipelines with limited feedback. One such challenge, the Ship Detect task, uses 5,000 simulated images to test whether a model can build a detection pipeline from scratch using hand-written algorithms.
The TOD Pipeline task requires models to filter raw radio telescope data to extract sky signals from noise and atmospheric emission. Developed by researcher @htihle with support from the Norwegian Defence Research Establishment, EpochAI, and METR, the benchmark produced results for open source models that user @teortaxestex called brutal. A previous version of the project was verified by a benchmark audit initiative, and new models will be added to the v3 tests over time.