The training population is the specific, defined set of resolved matters a model was fit against during its construction, and it is a distinct thing from the population a model is later scored against during out-of-sample validation, a distinction that matters because a model checked only against its own training population will appear calibrated by construction regardless of its genuine predictive value. A training population's size and representativeness relative to the population a model will actually be scoring in production determine how much confidence a calibration claim actually deserves: a model trained on a narrow, unrepresentative population can appear well calibrated within that population while performing poorly against the broader population it is deployed to score in practice. Disclosing a training population's size, date range, and the specific criteria used to define it is a basic transparency requirement for any calibration claim, since a claim that omits this information gives a counterparty no way to judge whether the underlying evidence actually supports the model's intended use, as distinct from simply supporting a narrower, less demanding use the model happened to be validated for and nothing more ambitious than that particular, narrowly scoped intended use case. A training population should also be re-examined whenever a model is retrained, since an expanded or contracted population changes what the model's calibration claim can honestly cover going forward.
Working through a diligence process?
Institutional partners evaluating a position against this platform's outcome and duration models are welcome to reach out directly.
