Prism Digital Labs
  • Services
  • Experiments
  • Resources
  • About
Contact Us
Resources›Prompting Guides›We deleted nearly 70% of our data

We deleted nearly 70% of our data
on purpose

A clean, smaller dataset you trust beats a large one you have to second-guess. How we chose precision over volume — and why that matters for any AI project.

By Prism Digital Labs · June 30, 2026
Advanced SystemsAdvanced7 min read

Prism Digital Labs built a discovery product enabling users to find local programs matching their needs. Success hinged on one critical factor: data quality. The team sourced a vendor dataset containing 14,457 business records covering the entire country—an appealing shortcut that appeared comprehensive and cost-effective compared to manual API queries.

However, the dataset proved problematic. Despite structured fields (names, addresses, phone numbers, categories), contamination was rampant.

The problem

The Seduction of Volume

Large datasets create psychological appeal during development. Fourteen thousand records "feels" like progress. Yet upon examination, the broad "recreational" category swept in irrelevant entries: lakeside cabin rentals, roadside motels, RV parks, bed-and-breakfasts, psychic services, and adult venues.

The data isn't missing. It's contaminated—full of records that look right at a glance and are wrong on inspection.

Contamination proves more dangerous than absence because absence appears obvious while contamination hides.

The decision

Precision Over Recall

The team made a deliberate strategic choice: optimize for precision rather than recall. They preferred displaying a smaller, trustworthy catalog over a larger one containing errors. A single misplaced entry undermines confidence in the entire dataset.

Gartner research supported this approach, recommending organizations prioritize data quality in high-risk, customer-facing contexts where bad answers cause reputational damage.

The approach

Cleanup Methodology

The solution involved layered filtering:

  • Stopped trusting category tags alone by requiring genuine youth/program signals within record names or descriptions rather than relying solely on categorical labels
  • Added hard exclusions for business types wrongly absorbed by categories: cabins, cottages, chalets, motels, hotels, RV parks, vacation rentals, and adult venues
  • Capped category sprawl by flagging records tagged with implausibly high category counts (legitimate programs typically occupy one or two categories)
  • Caught suspicious combinations such as listings tagged simultaneously as youth camps and pest control services
  • Applied AI strategically by running language models only on pre-filtered records to assign clean types and structured attributes, avoiding expensive processing of obviously contaminated data
The result

Results

The vendor file of 14,457 records reduced to approximately 4,500 validated entries—a removal of nearly 70%. The discipline of exclusion transformed the dataset into a defensible product.

Why it matters

Business Implications

Data contamination affects numerous industries. Gartner estimates poor data quality costs organizations approximately $12.9 million annually, yet roughly 6 in 10 organizations don't measure that cost at all.

Critically, automation amplifies rather than fixes bad data. AI models cannot distinguish between legitimate programs and misclassified motels—they process whatever input receives, implementing "garbage in, garbage out" principles.

The takeaway

Key Takeaways

  • Don't trust category labels — verify against actual content
  • Decide false-positive tolerance upfront — bias toward precision for customer-facing or compliance work
  • Filter using inexpensive rules before deploying expensive AI processing
  • Document exclusions comprehensively to justify data-cleaning decisions

The fundamental principle: A clean, smaller dataset you trust beats a large one you have to second-guess.

— Amanda
Prism Digital Labs

✦

Prism Digital Labs · Original content developed through client training engagements. Last reviewed June 2026.

Share this article
In this section
Your AI tool isn't broken.
We deleted nearly 70% of our data
Your CRM already has the leads.
Teaching AI to greet you
Build, buy,
Put this into practice

The Prism Prompt Catalog has 25+ ready-to-use prompts built around these techniques.

Browse the catalog →

Sitting on data you're not sure you can trust?

We'll audit one dataset with you — spot the contamination, show you what to cut, and build the filtering logic to clean it. Free, 30 minutes, no pitch.

Book a time →
Prism Digital Labs

We turn AI hype into measurable business results — through strategy, custom builds, and tools that actually ship.

hello@prismdigitallabs.com
Services
  • AI Strategy
  • Custom Development
  • Book a Consultation
Company
  • About Us
  • Experiments
  • Contact
Resources
  • Prism Pulse
  • Prompt Catalog
Connect
Women-led
Fortune 500 background
© 2026 Prism Digital Labs. All rights reserved.
Back to top