⏱️ Reading time: 16 min
A classifier that says “I’m 98% sure” should be right 98 times out of 100. In practice, many language models report that level of confidence and are right only 7 times out of 10: that mismatch between the probability a model reports and the real probability of being correct is exactly what model calibration measures and corrects.
📑 En este artículo
- TL;DR
- What Is Model Calibration?
- Why It Matters
- How It Works: From Logit to Calibrated Probability
- Practical Examples: Measuring ECE Step by Step
- Getting Started: Calibrating Your Own Classifier
- Comparing Calibration Methods
- Real-World Use Cases
- Common Mistakes and Best Practices
- Going Deeper: Why Networks Are Overconfident
- Frequently Asked Questions
- References
The problem shows up especially when a model is used as a decision model: a system that picks among fixed options in a single pass instead of generating text token by token. There, the reported confidence tends to be treated as a certainty score, but uncalibrated, that score lies.
TL;DR
- 98% confidence doesn’t mean 98% accuracy: calibration measures exactly that gap.
- ECE (expected calibration error) summarizes in a single number how far a model’s confidence drifts from its real accuracy.
- Temperature scaling corrects overconfidence by adjusting a single parameter over the logits, without retraining anything.
- Platt scaling and isotonic regression come in handy when the mismatch isn’t uniform across classes or bins.
- Twenty lines of PyTorch are enough to compute the ECE of any classifier before trusting its probabilities.
What Is Model Calibration?
Model calibration is the property of a classifier where the probability it reports for each prediction matches, on average, the real frequency of being correct: a calibrated model that says 80% confidence is right around 80% of the time, neither overestimating nor underestimating its own performance.
Calibration is not the same as accuracy. A model can be right 90% of the time and still be terribly calibrated if it reports 99% confidence in almost every case. A model with lower accuracy can be perfectly calibrated if its confidence levels honestly reflect its error rate. These are two orthogonal axes. One measures whether the model is right; the other, whether the model knows when it’s right.
Why It Matters
Model calibration matters especially in systems that automate decisions: medical triage, content moderation, credit risk scoring, or any pipeline that decides “I trust this, let it through without human review” above a confidence threshold. If that threshold is poorly calibrated, the pipeline automates mistakes with the same confidence it automates correct calls.
The case that prompted this article is illustrative. A developer evaluated a Qwen3-1.7B model as a multiple-choice classifier on a sample of CommonsenseQA, restricting the output to five possible tokens (A, B, C, D, E) and using the softmax of those logits as the confidence score. The model got 725 out of 1,221 questions right (59.38% accuracy), a reasonable result for a model with only 1.7B parameters and no fine-tuning. The problem wasn’t the overall accuracy: it was what happened when the predictions were grouped by confidence level.
When the predictions were split into ten confidence bins, the top bin (between 90% and 100%) averaged 98.55% confidence but was right only 70.09% of the time. The model was wrong almost 3 out of 10 times in the group where it claimed to be nearly certain. In the 80% to 90% bin the gap was even worse: 85.55% average confidence against 47.11% real accuracy. The model wasn’t just wrong: it was wrong with a confidence level that had no relationship to its real probability of success.
How It Works: From Logit to Calibrated Probability
It all starts in the network’s final layer. A classifier doesn’t directly produce a probability: it produces a vector of logits, unnormalized numbers that are later turned into a probability distribution with softmax. The formula is softmax(z)_i = exp(z_i / T) / Σ exp(z_j / T), where T is the temperature. With T = 1 you get the standard softmax; raising T flattens the distribution, and lowering it sharpens it.
The calibration problem arises because modern networks, trained with enough capacity and enough epochs, tend to produce logits with large magnitudes: the network “learns” to be confident in itself faster than it learns to be correct. The reference paper on the topic, “On Calibration of Modern Neural Networks” (Guo et al., 2017), documented this phenomenon in deep vision networks and showed that overconfidence grows with depth and parameter count, not necessarily with accuracy.
To measure how far a model deviates from perfect calibration, the ECE (expected calibration error) is used. The calculation groups predictions into M bins by confidence, and for each bin compares the average confidence against the real accuracy:
flowchart TD
A["Model logits"] --> B["Softmax with temperature T"]
B --> C["Probability per class"]
C --> D["Group into M bins by confidence"]
D --> E["Compare average confidence vs real accuracy"]
E --> F[("ECE = weighted sum of the differences")]
The formula is ECE = Σ (n_m / N) * |acc(m) - conf(m)|, where n_m is the number of examples in bin m, N the total, acc(m) the bin’s real accuracy, and conf(m) the bin’s average confidence. An ECE of 0 means perfect calibration; in practice, uncalibrated models tend to land between 0.10 and 0.25 on tasks with many classes.
A visual way to see the same thing is the reliability diagram: a bar chart where the X axis is each bin’s average confidence and the Y axis is that bin’s real accuracy. If the model were perfectly calibrated, every bar would touch the diagonal where confidence and accuracy are equal. When the bars fall below the diagonal in the high-confidence bins, as in the Qwen3-1.7B experiment, the diagram shows at a glance that the model is systematically overconfident.
Practical Examples: Measuring ECE Step by Step
The first step is always to measure before correcting. This example computes the ECE from an array of confidences and correctness labels, without depending on any deep learning framework:
import numpy as np
def expected_calibration_error(confidences, accuracies, n_bins=10):
bin_edges = np.linspace(0, 1, n_bins + 1)
ece = 0.0
n = len(confidences)
for low, high in zip(bin_edges[:-1], bin_edges[1:]):
mask = (confidences > low) & (confidences <= high)
if mask.sum() == 0:
continue
bin_conf = confidences[mask].mean()
bin_acc = accuracies[mask].mean()
ece += (mask.sum() / n) * abs(bin_acc - bin_conf)
return ece
# confidences and correctness labels (1 = correct, 0 = incorrect) for a sample classifier
confidences = np.array([0.98, 0.96, 0.99, 0.55, 0.60, 0.91])
accuracies = np.array([1, 0, 1, 1, 0, 0])
print(f"ECE: {expected_calibration_error(confidences, accuracies):.4f}")
Expected output:
ECE: 0.3317
An ECE of 0.33 is high: it means that, on average, the reported confidence drifts 33 percentage points away from the real accuracy in the bins where this handful of examples falls. In a real dataset with thousands of predictions, any value above 0.05 already warrants correction.
The second step is to bring this same logic to a real model. Following the restricted-token classification pattern used by the original Qwen3-1.7B experiment, this script computes the per-prediction confidence before applying any correction:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen3-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
def predict_with_confidence(prompt, option_token_ids):
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
logits = model(**inputs).logits[0, -1]
probs = torch.softmax(logits[option_token_ids], dim=-1)
idx = probs.argmax().item()
return idx, probs[idx].item()
# option_token_ids: ids for the A, B, C, D, E tokens obtained with tokenizer.encode
Running this function over a batch of questions and accumulating (confidence, correctness) for each one produces the same array that feeds expected_calibration_error(). In the reference experiment, the 90-100% confidence bin carried the most weight in the final average: it accounted for 809 of the 1,221 predictions.
Getting Started: Calibrating Your Own Classifier
Before installing anything you need three dependencies: Python 3.10 or higher, PyTorch, and Hugging Face’s transformers library. If you’re going to run a model like Qwen3-1.7B on CPU, you don’t need a GPU, though each inference takes longer.
On Linux (Ubuntu/Debian):
python3 -m venv venv
source venv/bin/activate
pip install torch transformers numpy
On macOS, the same block works the same way with the system’s or Homebrew’s python3. On Windows, activate the environment with venv\Scripts\activate instead of source venv/bin/activate; the rest of the commands are identical.
With the environment active, save the ECE measurement script from the previous section into a file called calibration.py and run it:
python3 calibration.py
Expected output (using the example array):
ECE: 0.3317
The next step is to correct that ECE with temperature scaling. The technique uses numerical optimization to find the value of T that minimizes the negative log-likelihood loss (NLL) over a validation set separate from the one used for training:
import torch
def fit_temperature(logits, labels, lr=0.01, max_iter=50):
temperature = torch.nn.Parameter(torch.ones(1))
optimizer = torch.optim.LBFGS([temperature], lr=lr, max_iter=max_iter)
nll = torch.nn.CrossEntropyLoss()
def closure():
optimizer.zero_grad()
loss = nll(logits / temperature, labels)
loss.backward()
return loss
optimizer.step(closure)
return temperature.item()
This implementation follows the same approach as the temperature scaling reference repository published alongside the Guo et al. paper: a single scalar parameter learned over already-computed logits, without touching the model’s weights. Applied to the 90-100% bin from the Qwen3-1.7B experiment, a temperature greater than 1 should pull that 98.55% average confidence closer to the 70.09% real accuracy observed.
Verification: How to Confirm Calibration Improved
There’s no single function call that says “your model is calibrated”: you have to recompute the ECE after applying the learned temperature and compare it against the original value.
calibrated_logits = raw_logits / learned_temperature
calibrated_probs = torch.softmax(calibrated_logits, dim=-1)
# recompute confidences and accuracies using calibrated_probs
# and call expected_calibration_error(...) again
If the recalculated ECE drops compared to the original, the calibration worked. If it rises or stays the same, the validation set used to learn T is too small or doesn’t represent the real production distribution: the code isn’t wrong, the validation data is.
⚠️ Heads up: temperature scaling doesn’t change which answer the model picks, only how confident it says it is about that choice. If the base accuracy is bad, calibrating won’t improve it.
Comparing Calibration Methods
Temperature scaling isn’t the only tool. The right choice depends on how many classes the problem has, how much validation data is available, and whether the miscalibration is even across classes or concentrated in a few.
| Method | When to Use It | Advantage | Limitation |
|---|---|---|---|
| Temperature scaling | Multiclass classification with uniform overconfidence across classes | Single parameter, quick to fit, doesn’t change the prediction ranking | Doesn’t fix it when miscalibration varies a lot between classes |
| Platt scaling | Binary classification or few classes | Captures simple nonlinear relationships between logit and probability | Needs fitting a logistic regression per class |
| Isotonic regression | Non-monotonic miscalibration or plenty of validation data | Assumes no functional form, fits any calibration curve | Can overfit with little validation data |
| Dirichlet calibration | Multiclass with asymmetric confusion matrices across classes | Models interactions between classes, not just the top confidence | More parameters to fit, more validation data needed |
flowchart TD
A["Uncalibrated logits"] --> B{"How many classes?"}
B -->|"Binary"| C["Platt scaling"]
B -->|"Multiclass"| D{"Is the miscalibration uniform across classes?"}
D -->|"Yes"| E["Temperature scaling"]
D -->|"No, it's asymmetric"| F["Dirichlet calibration"]
D -->|"Plenty of validation data"| G["Isotonic regression"]
Real-World Use Cases
Decision models like Jev apply this same restricted-classification logic to route support tickets, classify sentiment, or choose among a fixed set of actions in an agent. The advantage over generating text token by token is speed: a single forward pass delivers the answer, instead of the N passes an autoregressive model needs to produce N output tokens.
But speed doesn’t solve calibration. A triage system that uses a decision model to decide “escalate to a human” or “respond automatically” needs the reported confidence to be reliable in order to set the escalation threshold. If the model is overconfident as in the Qwen3-1.7B experiment, a threshold of “automate if confidence exceeds 90%” lets almost 3 out of 10 actually wrong cases through without review.
The same applies to content moderation systems, risk scoring, and any classifier that feeds a downstream automated decision. Confidence calibration doesn’t improve the model’s accuracy: it makes the confidence number that model reports actually useful.
Common Mistakes and Best Practices
- Using the same set to train and calibrate. The temperature (or Platt regression) has to be fit on a validation set separate from training; otherwise you’re measuring calibration on data the model already memorized.
- Confusing probability with next-token confidence. The softmax of an autoregressive LLM measures how confident the model is about the next token, not necessarily that the final answer is correct. Without adjustment, that gap inflates the reported confidence.
- Ignoring bin size when computing ECE. A bin with 5 examples and one with 800 carry different weight in the average; reporting the ECE without also showing the confidence histogram hides where the real problem is.
- Recalibrating once and assuming you’re done. The temperature learned on one dataset degrades if the input distribution shifts (a different domain, a different language, harder questions); you need to re-measure the ECE periodically.
- Treating ECE as the only metric. A model can have a low ECE and still be useless if its base accuracy is bad: calibration and accuracy are different axes, you need to watch both.
Going Deeper: Why Networks Are Overconfident
The central finding of the Guo et al. paper is that overconfidence doesn’t depend on the model being bad. It depends on how it’s trained. The standard loss function, cross-entropy, keeps dropping even after the network already classifies well, because it can still reduce the loss further by pushing the correct class’s logits higher and higher. That extra push improves the training loss but not the accuracy, and the side effect is an output distribution that keeps getting sharper, meaning more confident.
This phenomenon gets worse with depth. Networks with more layers and more parameters tend to overfit confidence before they overfit accuracy, an effect the paper’s authors documented by comparing architectures of different depths on the same vision datasets. In language models the mechanism is analogous: the bigger and more trained the model, the sharper its output distribution, with or without a direct relationship to how often it’s actually right.
There’s an important nuance about what a decision model’s confidence actually measures. The article that prompted this analysis points it out precisely: without additional training, the probability scores likely reflect the model’s confidence about what the next token will be, not the real probability that the answer is correct. These are two distinct quantities that align only if the model is calibrated, and most aren’t out of the box.
The good news is that fixing this doesn’t require retraining. Fine-tuning on the dataset itself, as the original experiment’s author did, improves both the accuracy (from 59.38% to 62.41%) and, indirectly, classifier calibration: a model that understands the task better has less room to be confident and wrong at the same time. But temperature scaling achieves a comparable improvement in pure calibration without touching a single model weight, in minutes instead of hours of training.
There’s a detail worth keeping in mind: calibration isn’t a static property of the model, it’s a property of the model-plus-data-distribution pair. A model calibrated on CommonsenseQA doesn’t necessarily stay calibrated on a medical or legal domain dataset, even with exactly the same architecture. That’s why the temperature learned in a validation environment needs to be re-measured whenever the type of questions the system receives in production changes.
Your next step: take the predictions from a classifier you already have in production, compute its ECE with the script from the “Practical Examples” section, and compare the result against the confidence threshold you currently use to automate decisions.
Frequently Asked Questions
What is ECE (expected calibration error)?
It’s the metric that summarizes, in a single number, how far a model’s average confidence drifts from its real accuracy, by grouping predictions into confidence bins and comparing both values in each bin.
Does temperature scaling change which answer the model picks?
No. It only rescales the reported confidence: the class with the highest logit before applying the temperature is still the chosen class afterward. The only thing that changes is how confident the model says it is about that choice.
How many bins should you use to compute the ECE?
Between 10 and 15 is standard in the literature. With less validation data, using fewer bins keeps each one from having too few examples, which would make the average noisy.
Does confidence calibration apply to generative models, not just classifiers?
Yes, but you first need to define what “correct answer” means for an open-ended task. In restricted classification, like a decision model, the metric is straightforward; in free-form generation it requires a separate evaluation criterion.
Can Platt scaling and temperature scaling be combined?
It doesn’t make sense to apply them together on the same logit: each one solves the same problem with a different approach. You pick one based on the number of classes and measure the resulting ECE to confirm which works better in that case.
References
- nishtahir.com: original article on decision models, restricted decoding, and the Qwen3-1.7B calibration experiment on CommonsenseQA.
- arXiv:1706.04599: “On Calibration of Modern Neural Networks” (Guo, Pleiss, Sun, and Weinberger, 2017), the paper that formalized ECE and temperature scaling.
- github.com/gpleiss/temperature_scaling: reference PyTorch implementation of temperature scaling.
- huggingface.co/docs/transformers: official documentation for the library used to load and run the Qwen3-1.7B model.
- pytorch.org: official documentation for the softmax function used in the code examples.
📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de BoliviaInteligente en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments