← Back to live feed

Thursday, Sep 24, 2026

1
TypeSafe AI Cuts Eval Cost to 57% with 99% of GPT 6 Accuracy
topics πŸ€– AI tags AIAI ModelsAI Research keywords

A research paper introducing a model cascade system called Jev-as-a-Judge demonstrates a method for routing artificial intelligence evaluation tasks to cheaper models based on confidence thresholds. Using JEV, a decision-only judge from TypeSafe AI, the system handles high-confidence verdicts and escalates only uncertain calls to GPT-6 Astra. This approach preserved 99% of the frontier model's accuracy across 510 preference pairs while reducing total fees to roughly 57% of the cost of using GPT-6 alone.

JEV costs $0.044 per 1,000 judgments with a median latency of 0.152 seconds, making it approximately 277 times cheaper than GPT-6, which costs $12.182 per 1,000 judgments with a latency of 1.885 seconds. On RewardBench and HaluEval, JEV stayed within 3 points of GPT-6, scoring 92.2% and 87.5% respectively. However, a gap of 9 to 20 points appeared on tasks requiring complex derivation checks, such as JudgeBench, where JEV scored 78.6% against 93.1% for GPT-6.

Continued in
LLMs Repeat 96% of JEV Confident Errors and Undermine Evaluation Cascades
6 tweets β€’ 3 sources
See all 4 tweets β†’