Stop asking the model to sound sure

Checked 22 Sep 2026 · By Luke Czak

ArticleOpinionFree to read

Asking for confident language gets you confident language, not confident answers. What actually needs to sit next to a claim is evidence, not tone.

Early on I used to tell agents, in more or less these words, to stop hedging and just tell me straight. I wanted a report that stated whether the tests passed or not, not three qualifying clauses wrapped around an answer I still had to decode myself. It seemed like a reasonable thing to ask for. Hedged language is genuinely harder to act on than a flat statement, and I did not want to spend my own attention untangling a sentence just to find out whether something had actually been checked or merely assumed.

What that instruction actually produced was a report written in a consistently confident tone regardless of whether the underlying claim had been verified or guessed at. A sentence stating that a migration had run cleanly read exactly the same whether the agent had run it and watched it succeed, or had inferred that it probably would from the shape of the change and never actually executed it. Confidence, it turned out, was a stylistic choice sitting entirely apart from the truth of the claim underneath it, and asking for more of that style did nothing to make the underlying claims more accurate.

It made the inaccurate ones considerably harder to catch, because they now read identically to the accurate ones. Before, a hedge at least told me where to look twice. A false claim delivered with total confidence gives me nowhere to look at all, because nothing in its surface tells me it needs checking, and I only found out it was wrong once something built on top of it failed further down the line, at which point tracing the failure back to a report that had sounded perfectly certain took far longer than it should have.

None of this is an argument for the opposite failure. A report that hedges everything equally, wrapping every statement in the same cloud of maybe and possibly, is just as useless as one that is falsely confident about everything, because neither version tells me which specific claim I can act on without checking it myself first. Uniform hedging and uniform confidence are the same underlying failure wearing different clothes, and both apply one tone across an entire report regardless of what actually happened underneath any individual line of it.

What I actually wanted, it turns out, was never a tone at all. It was a way to tell a checked claim apart from an inferred one, and tone cannot carry that distinction because tone is chosen independently of it, by habit or by instruction, rather than by what was actually done. So I stopped asking for confidence and started asking for evidence attached to every claim that mattered: the command that was run, the output it produced, the specific thing that was observed rather than assumed to be true. A claim with a command and its output sitting next to it does not need a confident tone to be worth trusting. A claim without one does not become trustworthy no matter how sure it sounds.

The uncomfortable part of this change is that it is more work than a tone instruction ever was, because it means actually reading the evidence rather than pattern-matching on how certain a sentence sounds, which was always the faster habit and the one I keep having to catch myself falling back into. I would rather do the slower version than keep training myself to trust a voice that has no relationship to whether the thing underneath it is true. Asking a model to sound sure was asking it to solve a presentation problem. The problem I actually had was never presentation.

Comments (0)

Sign in to comment.