Company data API provider types compared - DataImpulse

A company data API turns a domain, a name or a registration number into a structured record: legal entity, location, size, industry, sometimes funding and technology. It looks like a solved problem until two providers return different employee counts for the same business and neither is obviously wrong.

This guide covers the four provider types, what each is actually good for, why entity identity is harder than any individual field, and when collecting yourself beats buying.


Key Facts

  • The identifier is the product. Fields are easy to find; a stable company identifier that survives renames, mergers and locale differences is not.
  • Four provider types exist, and they answer different questions: registries, aggregators, web-derived providers and financial data vendors.
  • Registry data is authoritative and slow; web-derived data is fast and inferential. Most real pipelines need both and should label which is which.
  • Coverage claims are about domains, not companies. A provider with fifty million domains has far fewer real operating businesses behind them.
  • Rate limits decide architecture. Enriching on demand at the moment of use is usually cheaper and fresher than maintaining a mirrored database.

What kinds of company data provider exist?

Four, distinguished by where their data originates. We call it the 4-part provider model.

Type Origin Best for
1. Registry Official company registers and filings Legal entity, ownership, registration status, compliance checks
2. Aggregator Licensed combinations of many sources Breadth in one call, at the cost of inherited errors
3. Web-derived Crawled company sites, job boards, news Freshness, technology signals, hiring, descriptions
4. Financial Filings, credit files, market data Revenue, risk and creditworthiness, usually paid and jurisdiction-limited

A useful shortcut when evaluating: ask which type a specific field comes from. An employee count from a registry filing and an employee count inferred from a professional network are different measurements wearing the same label, and only one of them updates monthly.


Why is entity identity the hard part?

Because a company is not one thing. It is a legal entity, possibly several, a trading name, a set of domains, a group structure, and a presence in multiple jurisdictions, and these change independently.

Concretely: a business rebrands and keeps its old registration; a group operates twenty subsidiaries under one brand; a company’s legal name is written differently in two registries; a domain is sold to an unrelated business. Any of these breaks a naive join, and the failure is silent because the record still looks complete.

What to require from a provider: a stable internal identifier that persists across name changes, an explicit link to official registration numbers where they exist, and documented behaviour on mergers. Any provider whose primary key is effectively the domain will produce duplicates and orphans in your database within a year.


Should you use an API or collect the data?

Situation Use an API when Collect yourself when
Breadth across millions of companies Always: the maintenance is the cost, not the crawl Never realistic at this scale for one team
Legal entity and registration checks Always: registries are authoritative and often free Never scrape a registry that publishes an API
A narrow vertical with unusual fields The provider actually covers your niche Your fields are not in any product, which is common
Freshness on hiring, technology or pricing The provider states a refresh cadence you can verify You need it weekly, where crawling beats licensing
Region-specific content The provider collects from that region Pages differ by country and you can pin the exit

The realistic architecture for most teams is both: an API for identity and breadth, and your own collection for the few fields that decide your business and change faster than a licence refreshes. DataImpulse residential supports the second half at $1 per GB across 195 countries with country targeting included.


What are the limits?

Coverage is not what it sounds like. Counts are usually domains, and a large share of domains are parked, redirects or micro-sites rather than operating businesses. Ask for coverage within your segment, on your own sample.

Inferred fields must be labelled. Employee count, revenue band and industry are frequently modelled rather than filed. An unlabelled estimate in an outbound email is worse than a blank field, because it cannot be corrected by anyone reading it.

Personal data changes the rules. Company records are commercial data; the moment contacts are attached, privacy law applies and the lawful basis is yours, not the vendor’s.

Rate limits shape design. Most providers price per call and cap throughput, which makes on-demand enrichment at the moment of use both cheaper and fresher than mirroring the whole database. General information, not legal advice.

Related: data enrichment sources, technographic data, B2B data providers compared.


Frequently Asked Questions

What does a company data API return?

Typically legal entity details, location, industry classification, size estimates and a domain, with some providers adding funding, technology and contacts. Which fields are authoritative depends on the provider type: registry-sourced fields are filed, web-derived fields are inferred and fresher.

Why do providers disagree about employee count?

Because they measure different things. A registry filing reports a figure from an annual submission; a web-derived provider infers from professional network profiles or job posting volume. Both are labelled employee count and neither updates on the same schedule.

What should I check before choosing a provider?

Whether they issue a stable identifier that survives renames and mergers, whether they link to official registration numbers, how they handle merged entities, and their coverage measured on your own sample rather than on a global headline count.

Should I collect company data myself?

For breadth, no: maintaining coverage across millions of companies costs far more than a licence. For a narrow vertical, unusual fields, or anything needing weekly freshness, yes, because that is exactly what licensed datasets are worst at.

Do company data APIs include contacts?

Some do, and that is where the compliance picture changes. Company records are commercial data, while named contacts are personal data with privacy obligations that rest on you as the user, regardless of what the vendor’s terms assert.


The fields no licence refreshes fast enough

Hiring, technology and pricing move faster than licensed datasets refresh, and they render differently by country. DataImpulse residential proxies give country-pinned exits at $1 per GB across 195 countries. Create an account and collect the fields that decide your business.

Related: data enrichment · technographic data · B2B data providers.

Last updated: September 17, 2026.


Share article: