Prism Digital Labs built a discovery product enabling users to find local programs matching their needs. Success hinged on one critical factor: data quality. The team sourced a vendor dataset containing 14,457 business records covering the entire country—an appealing shortcut that appeared comprehensive and cost-effective compared to manual API queries.
However, the dataset proved problematic. Despite structured fields (names, addresses, phone numbers, categories), contamination was rampant.
The Seduction of Volume
Large datasets create psychological appeal during development. Fourteen thousand records "feels" like progress. Yet upon examination, the broad "recreational" category swept in irrelevant entries: lakeside cabin rentals, roadside motels, RV parks, bed-and-breakfasts, psychic services, and adult venues.
The data isn't missing. It's contaminated—full of records that look right at a glance and are wrong on inspection.
Contamination proves more dangerous than absence because absence appears obvious while contamination hides.
Precision Over Recall
The team made a deliberate strategic choice: optimize for precision rather than recall. They preferred displaying a smaller, trustworthy catalog over a larger one containing errors. A single misplaced entry undermines confidence in the entire dataset.
Gartner research supported this approach, recommending organizations prioritize data quality in high-risk, customer-facing contexts where bad answers cause reputational damage.
Cleanup Methodology
The solution involved layered filtering:
- Stopped trusting category tags alone by requiring genuine youth/program signals within record names or descriptions rather than relying solely on categorical labels
- Added hard exclusions for business types wrongly absorbed by categories: cabins, cottages, chalets, motels, hotels, RV parks, vacation rentals, and adult venues
- Capped category sprawl by flagging records tagged with implausibly high category counts (legitimate programs typically occupy one or two categories)
- Caught suspicious combinations such as listings tagged simultaneously as youth camps and pest control services
- Applied AI strategically by running language models only on pre-filtered records to assign clean types and structured attributes, avoiding expensive processing of obviously contaminated data
Results
The vendor file of 14,457 records reduced to approximately 4,500 validated entries—a removal of nearly 70%. The discipline of exclusion transformed the dataset into a defensible product.
Business Implications
Data contamination affects numerous industries. Gartner estimates poor data quality costs organizations approximately $12.9 million annually, yet roughly 6 in 10 organizations don't measure that cost at all.
Critically, automation amplifies rather than fixes bad data. AI models cannot distinguish between legitimate programs and misclassified motels—they process whatever input receives, implementing "garbage in, garbage out" principles.
Key Takeaways
- Don't trust category labels — verify against actual content
- Decide false-positive tolerance upfront — bias toward precision for customer-facing or compliance work
- Filter using inexpensive rules before deploying expensive AI processing
- Document exclusions comprehensively to justify data-cleaning decisions
The fundamental principle: A clean, smaller dataset you trust beats a large one you have to second-guess.
— Amanda
Prism Digital Labs