← Back to live feed

Friday, Sep 25, 2026

1
TypeSafe AI JEV Judge Cuts Agent Eval Costs 277x vs GPT-6
topics πŸ€– AI tags AIAI ModelsAI Research keywords

A multi-stage filtering process for model verification maintained 99% of the accuracy of top-tier LLMs on 510 preference pairs while operating at a reduced expense. This setup utilizes JEV, a decision-only tool from TypeSafe AI that costs $0.044 per 1,000 judgments, compared to $12.182 for GPT-6. The system provides verdicts and label probabilities without reasoning text and reaches a median latency of 0.152 seconds, versus 1.885 seconds for GPT-6, making it roughly 277 times cheaper.

JEV performs within 3 points of GPT-6 on RewardBench and HaluEval, but the gap widens to 9 or 20 points on JudgeBench, where JEV scored 78.6% against 93.1%. While a cascade of JEV and GPT-6 Astra can maintain high accuracy by escalating low-confidence calls, separate experiments show that other LLM verdicts repeat JEV's most confident errors 96.0% of the time. This correlation indicates that confident mistakes may not be caught through escalation, compared to 50% if errors were independent.

Image via @deliprao on X
You're reading an older version of the story.
LLMs Repeat 96% of JEV Confident Errors and Undermine Evaluation Cascades
6 tweets β€’ 3 sources
Continues from Thursday, Sep 24
TypeSafe AI Cuts Eval Cost to 57% with 99% of GPT 6 Accuracy
4 tweets β€’ 3 sources
See all 4 tweets β†’