CASE / stanislav-sorokinenSun Sep 27
MEDIA / 1video
SStanislav Sorokin@stas_sorokin_
代码与开发者工具JEV
Jev sorted 154 bank support messages into 77 intents for $0.011.
Jev sorted 154 bank support messages into 77 intents for $0.011.
Claude Opus 5.5 did the same 154 for $2.11 and got 83.8% right. Jev got 79.9%.
What came back:
→ one public test, Banking77, 2 messages per intent
→ 0.26 s a message vs 4.2 s
→ accuracy, latency, cost per 1,000, raw results per message
The twist: on the 122 messages where Jev said it was sure, it was right 88.5%. Higher than Opus overall.
Send only the 32 unsure ones to Opus: 83.1% for $0.51.
New open Jev alternatives like @TheVixhal's Gero-4b ship with speed numbers. This puts accuracy next to them, one command per model. @typesafeai
Real API calls, one run, built with Claude Opus 5.5.
Which model should run this test next?
Repost this now, because the next team picking a classifier should see accuracy, not only milliseconds.
Follow for more open source JEV builds. Code in the reply.
1,503浏览
3点赞
1收藏
0转发