Felo API PlatformFelo API Platform
К списку примеров
CASE / miracleenSun Sep 27
MEDIA / 1image
MMIRACLE!@0xMiiracle
Автоматизация агентовJEV

I built Choiceproof because a decision can look consistent until one small detail changes.

I built Choiceproof because a decision can look consistent until one small detail changes. Imagine an AI agent deciding whether to approve, escalate, or decline a refund. It gives an answer for a customer who requested a return four days after delivery. What happens if the request was made 45 days after delivery instead? What if you add an irrelevant sentence, reorder the evidence, or place an instruction inside the customer’s message? Choiceproof lets you test that question rather than guess. You define the decision and 2–5 answers it is allowed to choose. @openservai SERV Reasoning makes the original decision, then makes it again after a controlled change. Choiceproof shows both inputs, the exact edit, and both bounded answers side by side. A changed answer is a finding, not automatically a mistake. Sometimes the new fact should change the decision. You choose the expected answer and label what should have happened. That turns the comparison into a saved case instead of a screenshot you forget about. JEV has a separate, smaller role: it helps analyze the challenge and flag what deserves a closer look. It does not make the refund decision or decide whether SERV was right. SERV decides, a human labels the expectation, and Choiceproof keeps the evidence. The useful part comes when you revise the decision later. Choiceproof reruns saved cases against the new version and shows regressions, including evaluation cases kept out of the revision prompt. The goal is a repeatable check for teams building agents, not a claim that one test proves an AI system is safe. You can try the public workbench here: https://t.co/MoLvXPZ6nl I’d love to know what bounded decision you would try to break first. @openservai
109просмотров
9лайков
1сохранений
3репостов

Оригинал и ссылки