← Back to live feed

Friday, Sep 18, 2026

1
GPT 6 Astra Outperforms Rivals in First WeirdML v3 Agentic Benchmark
topics πŸ€– AI tags AIAI ModelsAI Research keywords OpenEpochAIMETR

The newest version of a machine learning evaluation suite favors OpenAI's latest model, particularly in token efficiency. WeirdML v3 consists of 11 complex hand-made tasks requiring models to explore unfamiliar data and develop machine learning pipelines with limited feedback. One such challenge, the Ship Detect task, uses 5,000 simulated images to test whether a model can build a detection pipeline from scratch using hand-written algorithms.

The TOD Pipeline task requires models to filter raw radio telescope data to extract sky signals from noise and atmospheric emission. Developed by researcher @htihle with support from the Norwegian Defence Research Establishment, EpochAI, and METR, the benchmark produced results for open source models that user @teortaxestex called brutal. A previous version of the project was verified by a benchmark audit initiative, and new models will be added to the v3 tests over time.

Image via @htihle on X
You're reading an older version of the story.
GPT 6 Astra Leads WeirdML v3 Agentic Benchmark of 11 Tasks
13 tweets β€’ 8 sources
Earlier version from Friday, Sep 18
GPT 6 Astra Tops Initial Token Efficiency on 11 Task WeirdML v3 Benchmark
5 tweets β€’ 3 sources
See all 11 tweets β†’