Felo API PlatformFelo API Platform
К списку примеров
CASE / elvisenSun Sep 27
MEDIA / 1image
Eelvis@omarsar0
Исследования и экспериментыJEV

Cool paper on catching alignment failures with Jev.

Cool paper on catching alignment failures with Jev. The ideas is to ask Jev one generic yes/no question about a model's response, and use its probability as a score. With no extra training, that score separates failures from good responses well, with a median AUROC of 0.886. It is also cheap. On 19 benchmarks, a Jev pass cost $0.30, while the LLM judges those benchmarks use cost $18.96. Researchers built RLCDAlignBench from 44 existing benchmarks across ten failure types, including sycophancy, jailbreaks, deception, prompt injection and reward hacking. Jev, TypeSafe AI's calibrated decision model, answers many typed questions about one input in a single call, each with a probability. On StrongREJECT, it agrees with human labels as well as the GPT-4o-mini scorer does, and ranks responses better (AUROC 0.971 vs 0.929). Where Jev confidently disagreed with benchmark labels, it found label errors in three benchmarks. They find that the question wording matters little. The thresholds do not transfer between benchmarks, and fitting one on 10 labelled items raises median F1 from 0.706 to 0.793. Paper: https://t.co/vUZV4k786o Chat with Paper: https://t.co/3BLEIhdpYy
11,350просмотров
134лайков
111сохранений
12репостов

Оригинал и ссылки