A genuine backtest requires three specific conditions, and the word is frequently used to describe evaluations that satisfy none of them. First, the model has to be frozen at a defined baseline, ideally a sealed one, before the scored population is evaluated. Second, the scored population has to be matters the model has not seen, drawn from a period after the freeze date, not from the training data the model was built on. Third, the evaluation's population, cutoff date, and success criteria should be preregistered, locked before any result is generated, so the test cannot be quietly adjusted after an unfavorable early result to find a framing that passes. A backtest that reruns a frozen model against historical data it was trained on is not a test of forecasting ability; it is a check that the model remembers its own training set, and it will look successful by construction regardless of the model's real predictive value. A genuine backtest is the only structure that reflects how a model will actually be used, forming an opinion about a matter whose outcome is not yet known, and any claim using the word 'backtested' should come with the freeze date, the scored population's date range, and confirmation that the scored population was excluded from training.
Working through a diligence process?
Institutional partners evaluating a position against this platform's outcome and duration models are welcome to reach out directly.
