Accuracy and confidence describe different parts of an answer

Imagine two colleagues who usually solve the same kinds of problems correctly. One sounds completely sure whenever an answer feels familiar, while the other can distinguish a remembered fact from an educated guess. Their accuracy may be similar, yet their confidence gives different guidance about when another check would help. Accuracy describes whether an answer is right; calibration describes the relationship between expressed confidence and results over time. That relationship matters when you decide whether to submit a calculation, check a source or allow extra time for a delivery.

A confidence judgement of 80% means expecting about eight correct answers in every ten comparable judgements made at that confidence. It leaves room for some errors. Neither an isolated wrong answer nor an isolated success can settle whether your confidence is well calibrated, because a probability that allows uncertainty can be followed by either kind of outcome. In their 1982 review, Sarah Lichtenstein, Baruch Fischhoff and Lawrence Phillips described substantial mismatches between confidence and correctness across judgement tasks. They also discussed differences between tasks and the effects of difficulty, which makes a universal verdict about someone's confidence unhelpful. Keep separate records for areas where the evidence available to you differs.

A calibration plot makes the mismatch visible

The diagonal below represents confidence that agrees with the proportion of correct answers. An overconfident pattern falls below that diagonal because the answers succeed less often than their stated confidence suggests. These invented values illustrate the shape found in overconfidence research; they describe no participant sample and estimate no population average. When making a plot from your own record, compare similar decisions and retain the number of observations behind each point, because a handful of answers can move a percentage sharply and an early curve deserves patient observation before you change your habits around it.

Confidence compared with correctness in two schematic patterns Confidence runs from zero to one hundred per cent. At confidence values of zero, twenty, forty, sixty, eighty and one hundred per cent, the calibrated line has zero, twenty, forty, sixty, eighty and one hundred correct answers per hundred hypothetical answers. The illustrative overconfident curve has zero, ten, thirty, forty, fifty and sixty correct answers per hundred hypothetical answers. These are invented values with no study sample. 0 20 40 60 80 100 0 20 40 60 80 100 Stated confidence (%) Correct answers per 100 hypothetical answers Calibrated line Overconfident pattern
Schematic, not data: the diagonal matches confidence to correctness, while the illustrative overconfident curve falls below it across hypothetical answers.
Invented plot values, with no measured sample
Stated confidence (%)Calibrated correct answers per 100 hypothetical answersOverconfident correct answers per 100 hypothetical answers
000
202010
404030
606040
808050
10010060

Estimation ranges reveal uncertainty that a single number hides

Suppose you are estimating how long an unfamiliar repair will take. A precise duration may conceal uncertainty about the fault, the tools available and interruptions. Giving a lower and upper limit forces you to consider those possibilities. A 90% range is intended to contain the true value in about nine of every ten comparable estimates over time. Its width should reflect what you know. A narrow range around an appealing guess can sound informative while leaving out plausible outcomes you never considered.

Alpert and Raiffa's 1982 chapter reported overprecision when people assessed uncertain quantities, because the ranges they gave caught the true value less often than they intended. Soll and Klayman (2004) later found the same pattern in interval estimates: the correct answer fell inside people's intervals much less often than the stated confidence implied. They traced the main cause to intervals that were systematically too narrow for the accuracy of the information behind them. The degree of overconfidence varied greatly with the way the limits were elicited and with the subject area, so no single coverage rate describes everyone.

For an everyday estimate, consider the lower and upper limits separately before settling on a central guess. Think about what could make the result surprisingly low, then give the same attention to what could make it high. This suggestion follows from the risk of overlooking possibilities, rather than guaranteeing a particular improvement. When the result arrives, record both whether it fell inside your range and whether it missed above or below. Repeated misses in one direction give you a concrete assumption to inspect next time.

Forecasting research connects confidence with checkable outcomes

Philip Tetlock's Expert Political Judgment (2005) examined expert predictions against subsequent events, making expressed belief something that could be evaluated. Mellers and colleagues (2014), including Tetlock, reported on a geopolitical forecasting tournament in which probability training and teamwork improved forecast accuracy under their analysis. The training covered ideas such as comparison classes and averaging several estimates, and every forecast could later be checked against what happened. The work gives a reason to take recorded predictions seriously as a learning practice, although it cannot show that every feedback routine would produce the same benefit.

A later reanalysis by Hauenstein and colleagues (2025) questioned how firmly the tournament had shown that teaming and training improved underlying forecasting ability. Using a statistical model of latent ability, they found that once key extraneous variables were controlled, the effects of teaming and training were substantially reduced, eliminated or, in some cases, reversed. Their result depends on their own modelling choices, so the evidence about those interventions remains contested rather than settled in either direction. The practical lesson for your own record is to compare like with like, because an easier set of questions can look like improved judgement when nothing about your judgement has changed.

The Brier score gives confidence a consequence

The Brier score, introduced by Glenn Brier (1950), evaluates probability forecasts against outcomes. With an event that either happens or does not, it squares the gap between the stated probability and the outcome, then averages those squared gaps across forecasts. Lower scores indicate closer agreement overall. Squaring makes a large error count more heavily than a small one, so confidently backing an outcome that fails carries a substantial penalty. A cautious probability also scores worse than a confident one when the event does occur, so the measure rewards confidence only when it is earned.

That score measures overall forecast quality, which includes more than calibration alone. Someone who gives the same middling probability to every event might be calibrated across the whole collection while offering little help in distinguishing likely outcomes from unlikely ones. Reading a calibration plot alongside the decisions themselves makes this limitation easier to see. If your record suggests excessive certainty about delivery dates, inspect how you estimate the work as well as the confidence you attach to it.

A daily record makes feedback easier to use

Choose recurring judgements whose outcomes you can check, such as whether a planned task will finish before a stated deadline. Write down the event, your probability and the evidence available before the outcome, including any reason the task resembles previous ones. Specify what counts as finishing the task. Later, compare the outcome with the original entry rather than with a remembered version that may have become more cautious. Mark unresolved predictions as unresolved, because quietly discarding them can make the remaining record misleading.

Funga Wega provides a smaller setting for this habit: before locking in each answer, you set a confidence dial from 20% to 95%, and an estimation question asks for a range you are 90% sure contains the truth. The Insights screen compares confidence with results and shows whether those ranges catch the truth about nine times in ten. An error made with high confidence is an invitation to inspect the reasoning behind it. Use that information to choose a similar real task, record your next judgement before checking and examine whether the same assumption still leads you astray.

Frequently asked questions about calibration

Does being calibrated mean being right every time?

Calibration permits errors at a rate that matches expressed uncertainty. Judgements made with 80% confidence should be correct about eight times in ten comparable cases over time, so some mistakes are consistent with that confidence. Improvement still involves learning the subject, because appropriately acknowledged uncertainty cannot supply missing knowledge.

Should I lower my confidence after every mistake?

An isolated mistake gives limited information about a repeated pattern. Check whether you misunderstood the problem, overlooked evidence or faced an unusual outcome, then compare that explanation with other entries in the same area. A change based on several comparable observations has a clearer basis than a reaction to embarrassment.

Can a very wide range count as a successful estimate?

A wide range can include the truth while giving little guidance for a decision. Check its coverage and its usefulness together, asking whether the limits would help someone plan time or resources. As you learn more about the quantity, test whether narrower limits can retain the intended coverage across later estimates.

Sources