Calibrating LLM estimates

2024-08-29 — 2025-02-24

quality 4.9

Wherein Transformers Are Examined as Generalized Inference Machines and Is Argued That In‑context Learning Is Modeled as Bayesian Conditioning, With Twisted Sequential Monte Carlo Methods Proposed for Sampling.

approximation
Bayes
causality
generative
language
machine learning
meta learning
Monte Carlo
neural nets
NLP
optimization
probabilistic algorithms
probability
statistics
stringology
time series
Figure 1

Dumping ground for notes on probabilistic calibration of NNs. Proper scoring meets Bayesian foundation models, maybe for forecasting.

1 LLM fine-tuning for calibrated estimates

Thought I was clever for thinking this up but turned out to be done (Jang et al. 2024; Levy 2026).

2 Calibrating implicit certainty

TBD

3 RLCD and Jev

Jev is a language model specialized for probabilistic classification outputs instead of textual outputs. WE give it a state (a string or a JSON blob) and a set of typed questions, and it returns a probability distribution over answers for each question. There are three question types: a yes/no probability (“noul”), a selection from a list of (up to) 255 labelled options (choice), and a position on an ordered rubric (score). Recall a chat LLM factorizes text as \(p(x_1,\dots,x_n)=\prod_t p(x_t\mid x_{<t})\) and emits one token at a time sampled from the successive next-token distribution. These can also classify things, probabilistically, but it is speaking their native language when they do so. To get a classification out of one of those, we ask for an answer, sample some tokens, parse the resulting string, and hope it is one of the labels we offered. If we want a probability, we either read off the logits of the label tokens if we have an open model (which is already a bit weird, since there are usually multiple tokens per label) or ask the model to say a number out loud. Models are notoriously overconfident in the second case. Sebastian Raschka posits that Jev is a pretrained transformer with a classification head bolted on the output layer.1 That would explain how the number and wording of options can change per query and yet the output remain consistent with fixed architecture. Because the output space is set before the model runs, it cannot return an answer outside the schema which soothes my OCD, although of course it can still be incorrect.

The training method, Reinforcement Learning for Calibrated Decisions (RLCD), notionally targets calibration rather than human preference (humans not being very calibrated). A forecaster is calibrated if, among all the occasions it says “70%”, the event happens about 70% of the time — formally, \(\Pr(Y=1\mid\hat p=p)=p\) for every \(p\). RLHF rewards answers that human raters like, and humans do love a spuriously confident snake-oil salesman, don’t we now? RLVR rewards answers that pass a verifier, which award 1 point for “correct” and none for “incorrect”; we might imagine this gives insufficient credit for being well-calibrated or hedging appropriately. The natural reward incentivizing calibration and resolution is a proper scoring rule such as the log score \(\log\hat p(y)\) or the Brier score, whose expected value is maximized by reporting an actual belief. So presumably RLCD does that? TypeSafe has not published however, and AFAICT a supervised classifier trained on cross-entropy is already minimizing a proper scoring rule, so it is not clear to me what would be special about their case? It might matter where there is no single label to train against; their evals score Jev against the averaged probabilities of GPT-6 Astra and Fable 5.1, for example (so they clearly think sufficiently powerful models are not terribly badly calibrated).

The result seems to be frugal. Normal reasoning models might emit thousands of thinking tokens to answer a single question. Jev, in contrast, reads the state once and answers every question in the same forward pass. A chat model pays per output token, and a reasoning model may emit thousands of them before it is done. Jev reads the state once and answers every question in the same forward pass, so it charges only for input and returns an answer fast. OTOH, such frugality does not admit such rumination. There is no chain of thought, so a question that needs several steps of reasoning has to be chunked into “instinctual” bits. This motivates the naming of the model family as System One, after Kahneman’s name for the fast, intuitive mode of thinking.

4 References

Hsieh, Fu, and Chen. 2024. “Reasoning and Tools for Forecasting.” In.
Jang, Cho, Lee, et al. 2025. “Reliable Decision‑Making via Calibration‑Oriented Retrieval‑Augmented Generation.” In Advances in Neural Information Processing Systems.
Jang, Lee, Lee, et al. 2024. “Calibrated Decision-Making Through Large Language Model-Assisted Retrieval.”
Levy. 2026. “Reinforcement Learning for LLM-Based Event Forecasting.”

Footnotes

  1. Prior art in this case would be BERT’s [CLS] token.↩︎