CASE / sanjay-kalra-ai-transformation-sherpaenMon Sep 28

MEDIA / 1image
SSanjay Kalra, AI Transformation Sherpaโข๏ธ@sanjaykalra
์ฝ๋ยท๊ฐ๋ฐ์ ๋๊ตฌJEV
๐ก #Jev from TypeSafe AI launched on September 15, and it took over my feed this weekend. I think it's the most
๐ก #Jev from TypeSafe AI launched on September 15, and it took over my feed this weekend. I think it's the most useful model release of the quarter for CIOs, as long as you use it for the narrow job it was built for.
Here's how to read the diagram below.
Left side: Jev accepts only three question types. Choice picks from up to 255 options. Score places something on 2 to 10 levels. Noul returns a yes/no probability.
Middle: a "System One" model. It reads the typed question and returns the type, and it never generates a token.
Right side: an option id, a level, or a probability. You get nothing to parse, and the answer can't land outside your schema.
Green boxes: 70 to 500 ms, $0.042 per million input tokens, output free. TypeSafe claims roughly 190x faster and 440x cheaper than frontier LLMs. Those are vendor best cases, so I treat them as a ceiling.
Red box: it can't generate text, do arithmetic, compare dates, or explain itself. It can also return the wrong valid value, so measuring accuracy is still your job.
Most steps in the agentic workflows our clients bring to @BayOneSolutions are small decisions like routing a ticket, flagging an invoice, or checking whether a claim is complete. Teams run each one through a frontier model and then write parsers and retry logic to clean up the output. Jev is built for that layer.
๐ A helpful guide for #CIOs looking at Jev:
1๏ธโฃ Inventory your decisions first. Tag every production LLM call whose output is a category, a rating, or a yes/no. Those are your candidates.
2๏ธโฃ Build a cascade. Jev routes at the front, code handles what it can, and a frontier model or a human takes the hard minority.
3๏ธโฃ Keep math and dates in code. Compute "over $10K" or "past renewal" first and pass the result in as state.
4๏ธโฃ Test calibration. If Jev says 0.8, about 80% of those cases should be right. Check against a holdout built from your own history.
5๏ธโฃ Pin versions. Use jev-1.13.0 in production, since jev-latest will move. For regulated decisions, log the input state and the schema version, because the model can't explain itself to an auditor.
6๏ธโฃ Guard the playbook. The model is rentable at four cents per million tokens. Your schemas and outcome labels encode how your business decides, so keep them in your own control plane.
7๏ธโฃ Size the vendor risk. It's waitlisted early access from a company with a fresh seed round. Pilot on a workflow that isn't critical.
8๏ธโฃ Expect Jevons Paradox. Sub-cent decisions mean far more decisions, so put the governance in place early.
Rajesh Samai and I will be at The @AIConference at Pier 48 in San Francisco, Sept 29 to Oct 1. Jev isn't on the agenda, so the useful conversations about it will happen in the hallways. If you're working out where a decision model fits in your stack, message me and bring one workflow. We'll map it in 20 minutes.
123์กฐํ
1์ข์์
0์ ์ฅ
0๋ฆฌํฌ์คํธ