CASE / fingerlingenSun Sep 27

MEDIA / 1image
FFingerling@lovelylogicss
エージェント自動化JEV
LANCET Nano v0.4.0 is out. Demo runs in your browser: https://t.co/NXOnWloADg
LANCET Nano v0.4.0 is out. Demo runs in your browser: https://t.co/NXOnWloADg
I also changed how I compare it to other command guards. A plain "risky or not" score didn't reflect how an agent actually uses a guard. What you want is:
- safe commands just run
- if it's unsure, ask
- if it's clearly dangerous, block it
So, the new Triage Score counts a risky command as caught whether it gets blocked or sent to me to confirm. Nothing for letting it through. And if a guard interrupts more than 10% of normal, safe commands, its score gets cut down, because in practice I'd just turn it off.
I ran every guard at its default settings on 3 benchmarks (5,053 commands) and weighted them by size. v0.4.0 came out at 73.2, v0.3.0 at 64.0, and @typesafeai Jev at 63.5.
Jev actually stops more risky commands (86% vs 75%), but it interrupts about 18% of safe ones to do it. v0.4.0 interrupts about 6%. The rule-based tools catch nearly everything, but they also stop most normal work.
It's small (110M), runs locally in about 14 ms per command on a CPU, and never sends or runs your commands.
41表示
0いいね
0保存
0リポスト