Felo API PlatformFelo API Platform
返回案例列表
CASE / elvisenSun Sep 27
MEDIA / 1image
Eelvis@omarsar0
研究与实验JEV

Cool paper on catching alignment failures with Jev.

Cool paper on catching alignment failures with Jev. The ideas is to ask Jev one generic yes/no question about a model's response, and use its probability as a score. With no extra training, that score separates failures from good responses well, with a median AUROC of 0.886. It is also cheap. On 19 benchmarks, a Jev pass cost $0.30, while the LLM judges those benchmarks use cost $18.96. Researchers built RLCDAlignBench from 44 existing benchmarks across ten failure types, including sycophancy, jailbreaks, deception, prompt injection and reward hacking. Jev, TypeSafe AI's calibrated decision model, answers many typed questions about one input in a single call, each with a probability. On StrongREJECT, it agrees with human labels as well as the GPT-4o-mini scorer does, and ranks responses better (AUROC 0.971 vs 0.929). Where Jev confidently disagreed with benchmark labels, it found label errors in three benchmarks. They find that the question wording matters little. The thresholds do not transfer between benchmarks, and fitting one on 10 labelled items raises median F1 from 0.706 to 0.793. Paper: https://t.co/vUZV4k786o Chat with Paper: https://t.co/3BLEIhdpYy
11,350浏览
134点赞
111收藏
12转发

原创与引用