Criterica Group — The institutional data science platform for regulated outcomes. A Splitifi company.
Data Provenance

Every model traces back to a real record.

How the corpus behind Criterica's intelligence is sourced, deduplicated, and labeled, and where the production model fleet draws its training population from.

1.1B+
Verified Unique Records
10,470
Production Predictive Models
16,302
Verified Judges
2.5M+
Cases Analyzed

The Corpus

Criterica's models are built from a single sourcing discipline: every record enters the corpus from a public court or regulatory source, not from a vendor summary or a third-party aggregation layer that has already interpreted the underlying filing. The corpus today holds over one billion verified unique court and regulatory records, spanning state and federal civil dockets, judicial decisions, and regulatory filings across every jurisdiction the platform covers. Scale alone is not the discipline. Every record in that corpus has been through the same identity-resolution and deduplication process before it is treated as a distinct, usable observation, which is what separates a verified count from a raw ingestion total.

Raw ingestion, before deduplication, runs far higher than the verified figure, because the same underlying case, ruling, or filing frequently appears in more than one source system, under more than one identifier, with formatting that differs enough to look like a separate record until it is resolved back to the single matter it actually represents. The verified count is what remains after that resolution work, and it is the number the platform stands behind for any institutional, diligence, or acquirer-facing use.

How Records Are Sourced

Sourcing starts at the primary record: a docket entry, a filed opinion, a regulatory disposition, retrieved from the court, agency, or regulatory system that generated it, or from a data provider whose own feed is itself traceable back to that primary system. A record's source is captured and retained alongside the record itself, not discarded once ingestion completes, because the ability to trace a fact back to where it came from is what makes the corpus usable for institutional diligence rather than only for internal modeling.

Coverage spans the litigation and regulatory activity that drives outcome, duration, and settlement questions across the verticals this platform serves: personal injury, mass tort, commercial, intellectual property, employment, insurance-adjacent claims activity, healthcare and lien matters, real estate and construction, government contracts, and regulatory and environmental proceedings. Each vertical draws on the overlapping court systems relevant to it rather than a bespoke, siloed source per vertical, which is part of why cross-vertical patterns, a judge who sits across case types, a venue whose behavior is consistent regardless of claim category, are visible in the corpus at all.

Deduplication and Identity Resolution

Deduplication resolves the same underlying matter, party, attorney, judge, or entity across the different identifiers, name formats, and filing conventions that different court and agency systems use for it. A single case can appear under a state docket number, a related federal removal number, and a secondary reporting service's own internal identifier, and identity resolution is the work of confirming those all refer to one matter before the corpus counts it once. The same resolution discipline applies to parties and counsel: an attorney or a repeat corporate litigant appearing under slightly different name formatting across filings is resolved to a single entity, which is what makes judge-level, venue-level, and party-level pattern analysis reliable rather than fragmented across artificial duplicates.

This work is continuous rather than a one-time pass. As new records are ingested, they are checked against the existing resolved corpus before being added as new observations, and the deduplication logic itself is periodically re-validated against the full corpus to catch cases where an earlier resolution pass missed a match or, less often, incorrectly merged two distinct matters that only appeared similar on the surface.

What Feeds the Model Fleet

Not every record in the corpus feeds a production model, and not every model draws on the same slice of the corpus. The platform runs 10,470 production predictive models, each trained on a data population, a target definition, and a filter selected for the specific outcome, duration, or settlement question it answers, and each registered against a signature, its target, filter, sample size, and validation performance, that prevents two models differing only in name from being counted as separate assets. Models sit alongside a larger set of governance artifacts and reference tables that support the fleet, verification logic, reference taxonomies, calibration utilities, which are disclosed separately from the predictive count rather than folded into it, because a governance artifact is not a model making a prediction and should never be counted as though it were.

The corpus also supports intelligence products that are not models in the predictive sense at all: judge-level profiles built from 16,302 verified judges' decisional histories, and case-level research drawing on 2.5 million or more cases analyzed across the platform's coverage. Both draw on the same underlying verified, deduplicated corpus described above rather than a separate, less-verified data source.

Labeling and Real-Data Discipline

Every production model in the fleet is trained on real data only. No production model is trained on synthetic, simulated, or artificially generated case records, a discipline the platform re-verified fleet-wide in a universal retrain, and one that applies without exception across every vertical and every jurisdiction the platform covers. This matters specifically because a model's outputs are only as trustworthy as the population it learned from, and a model trained in part on synthetic data cannot make a verifiable claim about what real courts, real judges, and real regulatory bodies actually do.

Labels themselves are handled with the same discipline the corpus demands elsewhere: an outcome label is only as good as the record it was extracted from, and where a label's basis is contested, ambiguous, or dependent on a data source with known limitations, the platform's practice is to withhold or qualify the claim rather than publish a precise-looking number resting on a shaky foundation. That discipline is what allows the platform to state its verified figures plainly rather than defensively.

Related
The Legal Asset Integrity Standard →The Regulated Outcome Lifecycle →Regulated Outcomes Markets →

Questions on Data Provenance

Diligence teams, capital partners, and institutional buyers evaluating the corpus behind Criterica's models.

I am a
Name
Organization
Email
Message