Research
Hidden Errors in Big Data: The Case of Property Records
Overview Research area: AI safety and ethics, with a focus on data quality, data provenance, and the use of brokered datasets in research and public-sector machine learning. The paper sits at the inte
- arXiv
- 2607.28827
- Published
- 2026-07-30
- Authors
- Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho
AI summary
Overview
Research area: AI safety and ethics, with a focus on data quality, data provenance, and the use of brokered datasets in research and public-sector machine learning. The paper sits at the intersection of responsible AI, applied econometrics, and public finance.
Technical level: Intermediate. The concepts are accessible to a general reader, but the analysis relies on statistical matching procedures, discrepancy measures, and standard property tax assessment metrics.
Scope: This paper audits two widely used commercial property data brokers (ATTOM and Cotality) against county administrative records for Cook County, Illinois, from 2018 to 2021, documenting sale price errors, coverage errors, shared errors across brokers, and the resulting bias in estimates of property tax regressivity.
What This Paper Is About
Governments, researchers, and AI developers increasingly rely on datasets purchased from commercial data brokers, but the decisions brokers make about sourcing, cleaning, and imputing data are usually hidden from end users. The authors take one high-stakes case—brokered U.S. property records, which feed studies of gentrification, inequality, and property taxation as well as machine learning property valuation models—and check whether the data match the underlying government records. Their goal is to quantify how accurate and representative these brokered datasets actually are, and whether the errors they find distort the economic statistics people build on top of them.
Key Contributions
-
A transaction-level audit of two major property data brokers. The authors match every arms-length, single-family home sale in Cook County from 2018 to 2021 across three sources—Cotality, ATTOM, and the county's own open data repository—and compare them record by record. The filtered Cook County data contain 152,060 transactions from ATTOM, 147,731 from Cotality, and 153,044 from ground truth.
-
A distinction between two error types, separately quantified for each broker. The paper separates imputation errors (the same transaction has a different recorded sale price in brokered versus ground-truth data, likely from brokers imputing prices from transfer tax payments when prices are not directly available) from coverage errors (transactions that exist in ground truth but are missing, duplicated, or misclassified in brokered data).
-
The first direct comparison of two brokered property databases to each other. Rather than treating one broker's data as a validation source for another's, the authors measure how much the brokers' errors overlap at the transaction level, finding substantial dependence that undermines cross-source replication as a robustness check.
-
A demonstration that these errors change a headline economic statistic. The authors compute three standard measures of property tax assessment regressivity—the price-related differential (PRD), the log coefficient, and the Suits Index—using brokered and ground-truth data, and decompose the gap into imputation versus coverage contributions.
Main Findings
-
Sale price errors affect 1–2% of matched sales. For 1–2% of matched sales in Cook County, broker-provided sale prices differ from ground-truth sale prices by more than 5%. Among ATTOM's 134,427 matched transactions, 1,976 had discrepancies above 5%; among Cotality's 129,606 matched transactions, 1,919 did. Conditional mean errors are large in dollar terms—for example, ATTOM's conditional mean error in 2018 was 260.80% (95% CI 137.63%–383.97%) or $398,934, and Cotality's conditional mean error in 2021 was 133.24% (95% CI 93.07%–173.41%) or $431,771.
-
Many price errors look like exact multiples, consistent with imputation error. The most common ratios of broker-reported to ground-truth sale prices are 2/3, 1/3, and 2. For ATTOM, the 2/3 ratio accounted for 998 discrepancies (50.51%) and 2.000 for 152 (7.69%); for Cotality, 2/3 accounted for 1,093 (56.96%) and 1/3 for 138 (7.19%). The authors confirmed in discussions with one broker that sale prices may be imputed using local transfer tax rates when not directly available, and Cook County has 75 distinct municipalities that independently set those rates—so using the wrong municipality's rate or confusing the buyer's share with the seller's share produces an exact multiple of the true price. There were also cases of dropped, transposed, or mistranscribed digits (for example, $7,000,000 instead of $700,000).
-
Coverage errors affect 12–15% of transactions. Ground-truth transactions fail to match to brokered data for several reasons. For ATTOM, 18,617 transactions (out of 153,044) were unmatched: 36.8% were missing records, 8.6% matched a duplicate sale, 49.8% involved a misreported broker filter (38.0% arms-length sale, 13.5% sale price above $10K, 13.5% single-family home, 2.6% not multiparcel), 4.0% failed fuzzy match criteria, and 0.9% were other. For Cotality, 23,438 transactions were unmatched: 56.0% missing, 8.9% duplicates, 32.0% misreported filters, 2.3% fuzzy match failures, and 0.7% other.
-
Assessed value errors are much less common than sale price errors. ATTOM had 371 matched transactions with assessed value discrepancies exceeding 5% of ground truth, and Cotality had 352. These do not follow a clear pattern such as substitution between pre- and post-appeal values; only 6 instances in each broker's data were attributable to rounding error.
-
The brokers make the same mistakes. Of 1,976 discrepant ATTOM transactions and 1,919 discrepant Cotality transactions, 1,512 overlap—roughly 3 out of every 4 discrepant transactions. Among those shared discrepancies, the broker-reported sale prices are identical in 99.8% of cases. On coverage, the brokers share 11,094 unmatched transactions, roughly 50–60% of each broker's total unmatched transactions, and the reasons for match failure largely agree: among shared unmatched transactions, if one is missing in Cotality's data there is a 97% chance (5,432/5,615) it is also missing in ATTOM's. A plausible mechanism is the FTC consent decree that required Cotality (then-CoreLogic) to share data with ATTOM (then-RealtyTrac), though the authors note they cannot rule out other explanations such as similar imputation and processing methods, or brokers selling data to one another.
-
County records are a valid benchmark. The authors manually checked disagreements against Cook County Clerk deed records for a random 2% sample. For discrepancies with Cotality, the deed price matched the county record 89.2% of the time (33 of 37 cases), yielding overall county accuracy of 99.8%; for ATTOM, the deed price matched 97.2% of the time (36 of 37), yielding 99.9% county accuracy. In a few cases brokers appear to correct genuine administrative mistakes or reflect prices before later county corrections.
-
Aggregate distributions hide the problem. Summary statistics for price, assessed value, and assessment ratios look broadly comparable across the three sources in Cook County. For example, mean sale price is $350,405 (ground truth), $356,661 (ATTOM), and $367,500 (Cotality), and mean assessed value is $30,059, $29,718, and $30,213 respectively. The visible differences are a slight leftward skew in Cotality's sale price distribution and a corresponding rightward skew in its assessment ratios. The discrepancies only become apparent at the transaction level.
-
Regressivity estimates shift depending on the data source, and the direction of bias is broker-specific. Regressivity estimates derived from broker-reported data differ significantly from ground-truth values for at least one broker in all four sample years. Cotality data tend to underestimate regressivity for the log coefficient and PRD measures, while ATTOM-derived estimates generally overstate regressivity across all three metrics. Differences narrow by the end of the sample period.
-
Coverage error usually dominates, imputation error is smaller but consistent. For ATTOM, coverage error dominates in 2018 and 2019, when imputation error rates were around 1%; in 2020 the share of transactions with imputation error rose to 3.7% and imputation error explained a larger portion of the gap; in 2021 the differences largely disappeared. For Cotality, coverage error predominates wherever differences are statistically significant. Imputation error introduces a consistent but modest attenuation bias, while coverage error varies in sign and magnitude across brokers and years.
-
Shared errors do not imply agreeing estimates. The substantial overlap in broker errors does not translate into agreement at the level of regressivity estimates—the remaining idiosyncratic coverage differences are large enough to produce substantial wedges between the two brokers' results.
-
Results generalize beyond Cook County. The authors report similar patterns using comparable data from New York City, New York, and Philadelphia, Pennsylvania, as robustness checks.
Methodology in Plain English
The authors take one county—Cook County, Illinois—where detailed property transaction records are published openly, and treat those county records as "ground truth" because they come directly from the administrative source and form the basis of what brokers resell. To support that assumption, they manually verify a random 2% sample of disagreements against the original deed records filed with the Cook County Clerk.
They then clean all three datasets the same way: keep only arms-length sales of single-family homes, drop sales under $10,000, drop multi-parcel sales and properties that sold more than once in a year (which may be duplicates or recording errors), and drop records missing address, sale date, or sale amount.
Next, they match records across datasets using property address, latitude and longitude, and sale date, with some tolerance for small textual differences. Any transaction where the broker and county sale prices differ by more than 5% of the county price counts as a discrepancy. They then look at the ratios between broker and county prices, which reveals the tell-tale pattern of exact multiples like 2/3 and 1/3. Anything the county has that a broker does not match becomes a coverage error, and the authors classify why each match failed using a concordance framework.
Finally, they compute three standard regressivity statistics used in the property tax literature—the price-related differential, the log coefficient, and the Suits Index—three ways: from the full ground-truth population, from ground truth with broker prices substituted in for matched transactions, and from the full broker population. The gap between the second and first estimates approximates the imputation error contribution; the gap between the third and second approximates the coverage error contribution.
Why This Matters
The paper shows that a dataset can look fine in aggregate while being wrong at the level of individual records, and that those individual errors are large enough to shift an economic statistic that shapes policy debates. Because brokers do not publish their data lineage, users have no practical way to detect these problems on their own—and because the errors are shared across brokers, the standard practice of checking one data source against another does not work here.
Real-world applications:
- Property tax assessment and regressivity research. Studies of whether assessments are regressive often use brokered data; the results show broker choice can move the answer, sometimes in opposite directions.
- Government use of commercial property data. ATTOM has contracted with the Department of Housing and Urban Development to provide foreclosure and sales data to inform agency policy, and both brokers supply data and software, including ML-based valuation models, that counties can use for assessments.
- Machine learning systems trained on or fed by brokered data. The paper frames property records as an example of a broader pattern in which data errors cascade through AI systems, making them harder to maintain and eroding institutional trust.
- Mortgage financing and private appraisals. These private-sector systems rely on the same brokers for property and transaction data.
Industry relevance: The paper speaks directly to the data brokerage industry, arguing that provenance and lineage disclosure is not a nicety but a precondition for users being able to assess data validity. It also notes that brokers frequently sell data to one another, creating propagation pathways for errors, and that the FTC's 2018 consent decree requiring Cotality (then-CoreLogic) to share data with ATTOM (then-RealtyTrac) provides one documented mechanism for error sharing. ATTOM claims coverage of more than 158 million parcels and Cotality advertises 99.9% market coverage—claims that, per this audit, coexist with 12–15% coverage error against a single county's records.
Future Directions
-
Expand the audit beyond three jurisdictions. The authors report that their findings generalize to two other large U.S. counties, but the main analysis rests on Cook County because most jurisdictions do not publish property data at scale. Extending the approach to more counties would clarify how much error rates vary with local reporting practices.
-
Trace the precise mechanisms of error propagation. The paper attributes shared errors most plausibly to the FTC-mandated data sharing agreement, but explicitly cannot rule out similar imputation and processing methods or inter-broker data sales. Identifying which mechanism dominates would help regulators target the right intervention.
-
Establish disclosure standards for data provenance and lineage. The authors call for open administrative data and transparency from brokers, but the paper does not specify what a disclosure standard should contain or how compliance would be verified.
-
Develop practical validation tools for end users. Since broker aggregation means users cannot currently validate claims of accuracy or comprehensiveness, there is an open question about what checks researchers and agencies can run—and what independent institutional infrastructure would be needed to support them.
Target Audience
This paper is most valuable to empirical researchers in economics, public finance, and urban policy who use brokered property data; to AI and machine learning practitioners concerned with dataset documentation and upstream data quality; to government agencies and county assessors that license broker data or valuation models; to data brokers and their regulators; and to anyone in the responsible AI community interested in how failures hidden in commercial data pipelines propagate into downstream systems and policy conclusions.
Authors’ abstract
Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between brokers: more often than not, brokers make identical reporting errors for the same transactions. Third, to illustrate the significance of these errors, we measure their impact on estimates of property tax regressivity, finding that they drive significant wedges between estimates depending on the data source. These findings generalize to two other large counties in the U.S., and highlight the crucial importance of open administrative data and transparency from brokers regarding data provenance and lineage.