Criterica Group — The institutional data science platform for regulated outcomes. A Splitifi company.
Model Governance

Calibrated Probability Is the Institutional Standard

What separates a measured probability from a confident guess, how calibration is actually tested, and what an institutional counterparty should demand before pricing capital against someone else’s number.

September 2026

A probability is calibrated if, among every instance where a model states the same probability, the outcome occurs at that rate over a sufficiently large sample. State a given probability across a large enough population of matters that share that score, and the fraction that resolve that way should match the stated number. That is the entire definition, and it is a stricter standard than it sounds, because it requires the number to mean something outside the single case it was assigned to. A model can be highly accurate in the sense of ranking matters correctly — better outcomes score higher, worse outcomes score lower — while still being badly calibrated, systematically overstating or understating the actual frequency at every point on the scale. Ranking and calibration are different properties, and an institution pricing capital against a probability needs the second one, not the first.

A point estimate is not a probability and should never be priced as one. A single expected-value figure — an average settlement amount, a median time to resolution — collapses an entire distribution into one number and discards the two things a capital allocator actually needs: the width of the distribution and the shape of its tail. Two matters can share an identical expected value and represent entirely different risks, one clustered tightly around the mean, the other split between a fast favorable resolution and a long adverse one with almost nothing in between. Pricing capital off the expected value alone treats these as the same position. They are not, and the difference is exactly the information a calibrated probability distribution, not a point estimate, is built to carry.

Calibration and resolution — in the forecasting sense, not the legal sense — are also not the same property, and treating a confident-sounding model as automatically well calibrated is a related error. A model that assigns every matter the same base-rate probability is trivially well calibrated in aggregate, because its stated number always matches the overall frequency, but it is useless, because it carries no case-specific information. A useful model has to be both calibrated and resolved: its probabilities have to spread meaningfully across a real population, and each point on that range has to hold up against realized frequency. Reporting only that a model is calibrated, without reporting how widely its scores actually spread, describes half of what a buyer needs to know.

A language-model inference is a different failure mode, and a more dangerous one, because it produces a number formatted identically to a calibrated probability. Ask a general-purpose language model for the likelihood a case settles, and it will answer with a specific figure, delivered with the same confident tone it uses for any other question. That figure was generated by predicting a plausible next token given text describing similar disputes in its training corpus. It was not fit against a labeled history of actual case outcomes, it was not checked for calibration against a holdout set, and there is no mechanism by which it improves when shown that its estimate was wrong, because there is no feedback loop connecting the answer to a ground truth. The number is a stylistic artifact of how confidently such figures tend to appear in the text the model was trained on. It should be read as a description of rhetorical convention, not as a measurement.

The distinction matters because the two numbers are visually indistinguishable and functionally opposite. A calibrated model's stated probability and a language model's stated probability look identical in a report — the same two or three digits followed by a percent sign. One was measured against a defined population of resolved matters and can be audited by re-running it against new resolutions as they arrive. The other cannot be audited in any meaningful sense, because there is no population it was validated against and no mechanism for tracking whether it drifts. An institutional buyer who cannot distinguish these two sources in a vendor's output is exposed to a category error that no amount of downstream risk management corrects.

Calibration has to be measured, not asserted, and the measurement has a specific shape: a reliability curve that plots stated probability against realized frequency across bins of the score, alongside the sample size behind each bin. A model that claims to be calibrated but cannot produce this curve, or produces it with bins too sparse to mean anything, is making a claim it cannot support. The Brier score, the standard scoring rule for probabilistic forecasts, decomposes cleanly into a calibration term and a resolution term, and a vendor citing an aggregate score without showing that decomposition is choosing not to show which of the two properties it actually has.

The measurement also has to be out-of-sample and out-of-time, or it measures nothing. A model checked for calibration against the data it was trained on will look calibrated by construction; fitting a curve to minimize error on a dataset and then checking whether that curve fits the same dataset is circular. The only test that carries information is one where the model was frozen at a cutoff date, scored matters filed or pending after that date, and the actual resolutions were compared to the frozen model's stated probabilities once they came in. This is the only structure that reflects how the model will actually be used: forming an opinion about a matter whose outcome is not yet known.

The word 'backtested' has been stretched to cover almost any claim a vendor wants it to support, and institutional buyers should treat it as meaningless until it is defined. A backtest that reruns a frozen model against historical data the model was trained on is not a test of forecasting ability; it is a check that the model remembers its own training set. A genuine backtest fixes the model, scores matters it has never seen, drawn from a period after the freeze date, and waits for those matters to resolve before comparing. Any claim using the word 'backtested' should come with the freeze date, the scored population's date range, and confirmation that the scored population was excluded from training — otherwise the word is doing marketing work rather than describing a method.

Aggregate calibration is a necessary claim and an insufficient one. A model can be well calibrated in aggregate across an entire corpus while being badly miscalibrated within a specific circuit, case type, or procedural posture, if errors in one segment happen to offset errors in another. An institution whose actual exposure concentrates in a narrow set of those segments is exposed to exactly the miscalibration the aggregate number is hiding. The correct question is not “is this model calibrated,” it is “is this model calibrated within the population my book actually sits in,” and a counterparty that cannot answer that question at the segment level has only answered the easier, less useful version of it.

Calibration also decays. The population a model was validated against is a snapshot, and courts, statutes, and litigant behavior do not hold still. A calibration claim without a stated as-of date and a stated re-validation cadence is a historical fact being sold as a current one. The institutional standard is a defined refresh cycle, disclosed rather than implied, so a buyer knows not just that the model was calibrated once, but when it was last checked and what changed since.

What a counterparty should demand, concretely, is five things: the reliability curve itself, not a single summary statistic; the sample size and date range behind that curve; confirmation, in writing, that the validation was out-of-sample and out-of-time rather than checked against training data; calibration broken out by the segments that match the buyer's own exposure, not only in aggregate; and a stated cadence for re-checking calibration as new resolutions accumulate. A vendor that can produce all five has a probability. A vendor that can produce a single percentage and a confident tone has an opinion, and institutional capital should decline to price the difference away.

← All InsightsRequest Access →