Felo API PlatformFelo API Platform
事例一覧に戻る
CASE / fingerlingenSun Sep 27
MEDIA / 1image
FFingerling@lovelylogicss
エージェント自動化JEV

LANCET Nano v0.4.0 is out. Demo runs in your browser: https://t.co/NXOnWloADg

LANCET Nano v0.4.0 is out. Demo runs in your browser: https://t.co/NXOnWloADg I also changed how I compare it to other command guards. A plain "risky or not" score didn't reflect how an agent actually uses a guard. What you want is: - safe commands just run - if it's unsure, ask - if it's clearly dangerous, block it So, the new Triage Score counts a risky command as caught whether it gets blocked or sent to me to confirm. Nothing for letting it through. And if a guard interrupts more than 10% of normal, safe commands, its score gets cut down, because in practice I'd just turn it off. I ran every guard at its default settings on 3 benchmarks (5,053 commands) and weighted them by size. v0.4.0 came out at 73.2, v0.3.0 at 64.0, and @typesafeai Jev at 63.5. Jev actually stops more risky commands (86% vs 75%), but it interrupts about 18% of safe ones to do it. v0.4.0 interrupts about 6%. The rule-based tools catch nearly everything, but they also stop most normal work. It's small (110M), runs locally in about 14 ms per command on a CPU, and never sends or runs your commands.
41表示
0いいね
0保存
0リポスト

オリジナルとリンク