Thursday, Sep 24, 2026
1 TypeSafe AI Cuts Eval Cost to 57% with 99% of GPT 6 Accuracy π€ AI Sep 24, 11:40 AM EDT 4/3
A research paper introducing a model cascade system called Jev-as-a-Judge demonstrates a method for routing artificial intelligence evaluation tasks to cheaper models based on confidence thresholds. Using JEV, a decision-only judge from TypeSafe AI, the system handles high-confidence verdicts and escalates only uncertain calls to GPT-6 Astra. This approach preserved 99% of the frontier model's accuracy across 510 preference pairs while reducing total fees to roughly 57% of the cost of using GPT-6 alone.
JEV costs $0.044 per 1,000 judgments with a median latency of 0.152 seconds, making it approximately 277 times cheaper than GPT-6, which costs $12.182 per 1,000 judgments with a latency of 1.885 seconds. On RewardBench and HaluEval, JEV stayed within 3 points of GPT-6, scoring 92.2% and 87.5% respectively. However, a gap of 9 to 20 points appeared on tasks requiring complex derivation checks, such as JudgeBench, where JEV scored 78.6% against 93.1% for GPT-6.