A reliability curve is the specific chart that turns a calibration claim from an assertion into evidence: it plots a model's stated probability against the actual, realized frequency of the outcome across bins of the score, so that a viewer can see directly whether matters scored at, for instance, a stated seventy percent probability actually resolved favorably at close to that rate across the population behind that bin. A reliability curve is only as trustworthy as the sample size behind each bin, and a curve presented without disclosing that sample size, particularly at the tails of the score range where sample sizes are often thinnest, can look convincingly calibrated while resting on too little data to support the claim at those extremes. The Brier score, a standard scoring rule for probabilistic forecasts, decomposes cleanly into a calibration term and a resolution term, and a vendor citing only an aggregate score without producing the underlying reliability curve is choosing not to show which of the two properties its model actually has. An institution evaluating any outcome-probability claim should ask for the reliability curve itself, with bin-level sample sizes, rather than accepting a single summary statistic as a substitute, since the summary alone cannot reveal where along the score range the underlying evidence actually thins out.
Working through a diligence process?
Institutional partners evaluating a position against this platform's outcome and duration models are welcome to reach out directly.
