← Back to live feed

Friday, Sep 18, 2026

1
GPT 6 Astra Tops Early Tests of 11 Task WeirdML v3 BenchmarkPVT:OPAI
topics 🤖 AI tags AIAI ModelsAI Research PVT:OPAI keywords Harvard

OpenAI's newest AI model outpaced competitors in token efficiency during the first round of testing for a new agentic evaluation suite. The WeirdML v3 benchmark features 11 hand-made tasks that require models to create machine learning and data analysis pipelines from unfamiliar data and limited feedback. These tests evaluate how AI agents handle unspecified goals and restricted information to produce results.

Open source models lagged frontier systems in early metrics, with one evaluator describing the results for non-frontier models as "brutal." The benchmark, credited as work from Harvard, assesses a model's ability to build pipelines without explicit instructions. Its launch follows a review of the preceding WeirdML v2 version, which was verified as functional despite limitations in its text interface.

Image via @htihle on X
You're reading an older version of the story.
GPT 6 Astra Leads WeirdML v3 Agentic Benchmark of 11 Tasks
13 tweets • 8 sources
Earlier version from Friday, Sep 18
GPT 6 Astra Outperforms Rivals in First WeirdML v3 Agentic Benchmark
11 tweets • 6 sources
See all 13 tweets →