CASE / steven-bui-ai-agent-r-denMon Sep 28

MEDIA / 1image
SSteven Bui | AI / Agent / R&D@bmtriet
Код и инструменты разработчикаJEV
Multilingual Internal Routing Benchmark
Multilingual Internal Routing Benchmark
Testing off-the-shelf models, zero-shot — no fine-tuning, just download and run.
60 real-world routing cases across VI / EN / ZH.
Current results:
TypeSafe Jev — 53/60 (88.3%)
GLiNER 2.5 Multi — 49/60 (81.7%)
Tev1-0.8B — 42/60 (70%)
Jev is still leading, with two strong candidates right behind.
Goal: find the smallest, fastest model that is reliable enough for a specific enterprise task.
18просмотров
0лайков
0сохранений
0репостов