A raw record count is one of the easiest numbers in this business to inflate without anyone intending to, because the same underlying real-world event, a single docket entry, a single enforcement action, a single filed judgment, can enter a corpus more than once through more than one ingestion path, and reporting the raw total as though it represented the corpus's actual size overstates it, sometimes substantially, without any single step in the process having done anything an auditor would call dishonest.
The first mechanism is source overlap. Multiple harvesting processes pulling from different points in an ecosystem, a court's own electronic filing system, a commercial legal-data aggregator that repackages the same filings, a state-level public records portal that mirrors the same underlying court, will frequently capture the same underlying event more than once, represented with slightly different formatting, slightly different metadata, and no obvious shared key connecting the two representations back to the single real event they both describe.
The second mechanism is temporal re-harvesting. A periodic refresh process designed to catch new filings from a source will, unless carefully scoped, re-pull records that have not actually changed since the prior harvest, creating a second copy of the same event separated in the corpus by ingestion date rather than by source. This kind of duplication is easy to miss precisely because it does not look like the classic case of the same event appearing twice from two different systems; it looks like normal, expected growth in the corpus over time.
The third mechanism is the hardest to catch after the fact: datasets acquired, merged, or inherited from a prior initiative, brought into a corpus without independent verification of their own internal duplication, carry whatever duplication already existed inside them into the combined corpus, compounding any duplication the new harvesting process introduces on top. A merged dataset's history of how it was assembled is often undocumented by the time it reaches a new custodian, which means its duplication rate has to be independently investigated rather than assumed away.
Deduplication is not a data-hygiene nicety that can be deferred until a corpus is otherwise finished, because every claim built on top of the corpus inherits whatever duplication the corpus contains. A reported corpus size is the most visible casualty, but the quieter one is more consequential: a base rate computed from a sample that silently double-counts a share of its underlying events looks better-supported, carries a larger apparent sample size, than the true, deduplicated sample actually justifies, and a model or a claim built on that inflated sample size is overstating its own statistical confidence to everyone downstream, including to the institution building it.
Proper deduplication requires a defined notion of unique record identity that goes beyond an exact-match check on raw text, because the same underlying event frequently appears across sources with minor formatting differences, a docket number written with different punctuation, a party name abbreviated one way in one system and spelled out in another, that an exact-match process will treat as two different events. Deduplication done properly requires a verification pass built around this reality, run repeatedly as new data enters the corpus, not a one-time cleanup exercise performed once and assumed to hold indefinitely as the corpus keeps growing.
The right institutional posture, when a specific portion of a corpus has not yet been independently verified for duplication, is to hold any grand total that depends on that portion as provisional rather than citing a rounded, convenient figure and moving on. This is a more conservative posture than most organizations are comfortable adopting, because it means declining to state a headline number a marketing document would prefer to have, but it is the only posture consistent with actually knowing what the corpus contains rather than assuming it.
The discipline required here extends beyond a single cleanup effort into an ongoing practice, because a corpus that grows continuously through active harvesting accumulates new duplication risk with every additional source and every additional refresh cycle added over time. A deduplication process validated against the corpus as it existed a year ago says nothing certain about the corpus as it exists today if new sources or new harvesting paths have been added since, and an institution should treat deduplication as a standing operational process with its own monitoring, not as a milestone that was completed once and can be assumed to still hold.
This same discipline should extend to how an institution treats figures it has not yet been able to independently verify, particularly when those figures originate from data acquired through a merger, an inherited legacy system, or a source outside the institution's own harvesting pipeline. The correct posture is to flag such figures explicitly as unverified pending a defined verification process, rather than incorporating them into a reported total on the assumption that they are probably fine, because that assumption is precisely the gap through which the largest and least visible overstatements tend to enter a corpus's reported scale.
A useful institutional habit is publishing the deduplication and verification status alongside any headline corpus figure, stating explicitly which portions have been fully verified, which remain provisional, and what the known blockers to full verification are, so that anyone relying on the figure understands its current confidence level rather than assuming a single reported number carries uniform certainty across every source that contributed to it, a habit that costs an institution some polish today in exchange for a far stronger position the next time its figures face outside scrutiny.
A corpus's real value to an institution that depends on it for pricing risk is not its raw record count. It is its verified unique record count, the number that survives after every known source of duplication has been checked and removed, and an institution that reports the former as though it were the latter has overstated its own evidence base to itself well before that overstatement ever reaches a counterparty, which means the first person misled by an unverified corpus total is usually the institution that produced it.
