CASE / johnenFri Sep 25

MEDIA / 1image
JJohn@empowwerr
智能体自动化JEV
Ever had an LLM tell you it was 99% confident, only to find out the answer was complete BS?
Ever had an LLM tell you it was 99% confident, only to find out the answer was complete BS?
You even asked it to double-check. It gave you a longer explanation, and you spent another hour following advice that was wrong to begin with.
@typesafeai calls out overconfidence as a problem with large language models: they can express more certainty than their answers deserve.
Part of this can come from training. Researchers found that some systems used to score answers rewarded higher stated confidence even when the answer quality did not justify it. The model was being rewarded for sounding sure.
That does not make every confidence score useless. It means the score needs checking, just like the answer.
Suppose you check 100 answers that a model marked 99% confident. Only 80 are correct.
That percentage is giving you a misleading picture of the risk. Over enough comparable examples, roughly 99 out of every 100 answers labelled 99% confident should be right. Matching confidence to actual results is called calibration. Even properly calibrated confidence still allows mistakes.
Now change the example slightly.
The model gives the same answers, but marks every one 80% confident. Eighty are correct.
Its confidence now matches its accuracy perfectly on that test.
But which twenty answers should you check?
The score cannot help you. Every answer got the same number.
Knowing how often a model is wrong is different from knowing when it is likely to be wrong.
This distinction showed up in a study of 80 language models. Good calibration did not necessarily mean a model was good at assigning lower confidence to its mistakes.
That matters when someone builds a workflow that automatically accepts answers above a confidence threshold.
A rule like accepting everything above 95% only helps if those scores actually separate dependable answers from risky ones. Otherwise, you have automated trust in a number without establishing what it means.
Before using that rule, test it on examples of the work you actually need done, with independently checked answers. Measure how much work it accepts and how often those accepted answers are wrong. You need both numbers: refusing almost everything can make a system look reliable without making it useful.
For everyday use, I would change what I ask when an answer matters.
Asking whether the model is sure gives it another opportunity to express confidence. Ask it to perform a check that could expose a mistake.
For a factual claim, have it retrieve the original source and identify the passage that supports the answer. Check that the passage actually says what the model claims.
For code, have it run a test that would fail if the proposed fix were wrong, then show the result.
And let the answer remain unresolved when the evidence is missing. There is no benefit in turning an unanswered question into a confident guess just to finish the conversation.
The next time an LLM gives you 99% confidence, ask what it checked to reach that conclusion. Then judge the check.
#LLM #Jev
172浏览
8点赞
0收藏
2转发