Data enrichment sources and collection explained - DataImpulse

Data enrichment is the practice of taking a thin record, usually a domain or an email, and attaching everything else: company size, industry, location, technologies, funding, sometimes contacts. Looked at from the collection side rather than the sales side, it is a straightforward question of where each attribute comes from and how old it is.

This guide maps the four sources behind every vendor, explains why match rates mislead buyers, covers the decay curve nobody puts on a pricing page, and is specific about where proxies do and do not fit.


Key Facts

  • Every enrichment vendor draws from the same four wells: public web, official registries, contributed data, and partner exchanges. The difference is coverage and freshness, not magic.
  • Match rate is the most misleading number in the category. A 90 percent match on common fields and a 20 percent match on the field you actually need are quoted as the same headline.
  • Records decay at roughly a fifth to a third per year for people-level fields, because people change jobs. Company-level fields decay far more slowly.
  • Personal data is the boundary. Company records are commercial data; named contacts bring the GDPR and similar regimes into scope, whatever the vendor’s terms say.
  • Proxies belong to the public-web well only, and they change the exit address, not your lawful basis for holding a record.

Where does enriched data actually come from?

Four wells, and any vendor is a blend of them. We call it the 4-point enrichment model.

Source Typical fields Freshness
1. Public web Company description, location, technologies, hiring, news As fresh as the last crawl, which is the vendor’s real product
2. Official registries Legal entity, registration number, filings, ownership Authoritative but slow; updates follow filing cycles
3. Contributed data Contacts and firmographics submitted by users of a tool Fresh where the tool is popular, thin everywhere else
4. Partner exchanges Licensed sets bought or swapped between providers Inherited freshness, and the reason vendors share the same errors

For who sells these blends commercially, see our comparison of B2B data providers. That last row explains a common experience: three vendors returning the same wrong job title for the same person. They are not three independent observations, they are one observation resold three times. When you evaluate vendors, testing the same hundred records across all of them tells you more about overlap than any coverage chart.


Why is match rate the wrong headline number?

Because it averages fields of wildly different difficulty. Country and industry match easily and are nearly worthless; employee count, tech stack and direct contacts are hard and are the reason you are buying.

Ask for match rate per field on your own sample, not on the vendor’s. Two habits make that evaluation honest: send a list that reflects your real segment, including the awkward small companies rather than the Fortune 500 everyone covers, and measure accuracy separately from coverage by verifying a random subset by hand.

A vendor that fills 95 percent of rows with a plausible-looking employee count is not better than one that fills 60 percent and leaves the rest empty. It is worse, because you cannot tell which numbers are inferred.


How fast do enriched records decay?

Split the answer by level, because the two decay at completely different rates.

People-level fields go stale quickly: job changes alone invalidate a meaningful share of contact records every year, and titles drift even when people stay. Any outbound process built on a year-old contact list is mostly paying for bounces and annoyance.

Company-level fields are far more stable. Legal entity, headquarters and industry change rarely; technologies and hiring shift on a monthly-to-quarterly rhythm.

The design consequence is to enrich at the moment of use rather than maintaining a warehouse of everything. Enriching on demand costs more per record and far less in total, because you never pay to refresh the 90 percent of your database nobody touches.


Where do proxies fit, and where do they not?

They belong to the first well only: collecting public web evidence. They do not touch registries, contributed data or licensed sets, and they have nothing to do with whether you may hold a record.

Task Use Avoid when
Company-page and tech signals across many domains Rotating residential, one visit per domain The set is small, where a datacenter exit is cheaper
Region-dependent content, pricing or availability Residential pinned per country The field is region-independent
Official registries and open data portals Datacenter or no proxy Never over-engineer: these are meant to be read by machines
Anything behind a login Do not Always: that is the line where collection becomes a legal problem

DataImpulse residential is $1 per GB across 195 countries with country targeting included. We sell the exit layer, not a data product, and the honest framing is that a proxy makes public collection reproducible rather than permissible.


What are the privacy and quality limits?

Personal data. Named contacts, work emails and profile links identify individuals, which brings the GDPR, the UK regime, CCPA and others into scope regardless of what a vendor’s terms promise. Buying a list does not transfer a lawful basis to you, and the practical requirements, notice, purpose and retention, sit with the controller using the data.

Quality. Inference is unavoidable in this category, so what separates a usable dataset from a dangerous one is whether inferred fields are labelled. Keep the evidence and the detection date for every attribute, and never let an inferred value into an outbound template without a confidence threshold.

Terms. Honour robots.txt and site terms on the collection side, and prefer official sources where they exist. General information, not legal advice: see is web scraping legal.

Related: how tech stack detection works, 403 Forbidden when scraping.


Frequently Asked Questions

What is data enrichment?

Attaching additional attributes to a thin record, usually starting from a domain or an email: firmographics, location, technologies, hiring signals and sometimes contacts. The attributes come from public web collection, official registries, data contributed by users of a tool, and licensed partner sets.

Why do different enrichment vendors return the same wrong value?

Because they license from each other. Partner exchanges mean several vendors can be reselling one original observation, so three matching answers are not three independent confirmations. Testing the same sample across vendors reveals overlap faster than any coverage chart.

What match rate should I expect?

Ask per field on your own sample rather than accepting a headline number. Country and industry match easily; employee count, technologies and direct contacts are the hard fields and the reason you are paying. A vendor that fills every row with plausible values is hiding inference, not delivering coverage.

How quickly does enriched data go stale?

People-level fields decay fast because people change jobs; company-level fields are far more stable. The practical answer is to enrich at the moment of use rather than maintaining a fully refreshed warehouse, since most records in a database are never touched.

Do proxies make enrichment legal?

No. A proxy changes the exit address of a request and nothing about your lawful basis for collecting or holding a record. Personal data brings privacy law into scope regardless of how the data was obtained, and content behind a login is out of scope entirely.


Collect the public half reproducibly

Public-web enrichment is the well where the exit address matters. DataImpulse residential proxies give geo-accurate exits at $1 per GB across 195 countries, so region-dependent fields are collected the same way every run. Create an account and test one segment before scaling.

Related: technographic data · proxies for web scraping · is web scraping legal.

Last updated: September 17, 2026.


Share article: