The CFOs Guide to Ethical Scraping

Just 5 years ago, web scraping was solely IT teams’ headache, that C-level executives had nothing to do with. Today, the situation is drastically different: web data has become a line item in the Profit & Loss. Banks and investment funds create strategies based on pricing signals, the labor market, and customer feedback, gathered from public sources. Pharmaceutical companies track competitors’ prices, patent registries, and tenders in real time. The consequence? CFO now sits at the table with legal counsel to make any decision regarding data gathering, because the price of a legal mistake is high and is paid not only with money, but with reputation as well. DataImpulse is an ethical proxy provider, and we know first-hand what risks shady scraping and tools can bear. 

This article is not about how to gather data, but about how to prevent fines, court proceedings, or reputational damages.

Key Facts:

  • Proxies are not a technical expense only – unethical proxies lead to gathering wrong data and making wrong decisions at the least. Or worse, they can lead to legal risks due to the use of illegal tools and reputational damages. Reputable proxies are an investment in your legal safety and image.  
  • Even if you scrape publicly available data, they still need to be collected within legal limitations.
  • Aggressive measures to get around detection systems can create an additional risk. You need to scrape ethically – sticking to laws like GDPR and platforms’ ToS, not just trying to prevent detection. That is the difference between scraping and ethical scraping.
  • Before deciding on a proxy provider, check IP origin, law compliance, and ISO certification. 

Why should you not see proxies as a technical expense only 

CFO sees a proxy bill as an operational cost, but they miss the unobvious expenses:

  • Costs on retries – Cheap proxies with a bad reputation get blocked in the first minutes. Teams do not get an error message – they get wrong or incomplete data that looks just fine. Then decisions are made based on that data. The real price of such a decision is not $500 for proxies, but hundreds and millions of dollars due to flawed market assessment and mispredicted launch. 
  • Legal expenses – hiQ Labs vs LinkedIn, Ryanair vs Aviation court cases are not hypothetical – they are real examples that show: even if you are after public data, it is still important how you get it. The infrastructure you use, the technical limitations you stick to, and transparency regarding the traffic source – everything may become the subject of a court proceeding.
    Regulated industries like Pharma face more pressure –
    GDPR, and other regulations and requirements regarding data. The financial industry has it even harder with rules regarding market abuse and insider information in case scraping accidentally touches non-public data. 
  • Reputational and contract expenses – Enterprise clients more and more often demand to know how you get data for analytics. Not-so-straightforward answer or absence of one may become a reason for contract termination for those who care about compliance.

Thinking of proxy infrastructure just in terms of IT-related expenses underestimates it as a tool for risk management. 

Why cheaper proxies are not the best solution 

The solution to buy the cheapest proxy pool looks wise only in the short run. In reality, such pools are often built using schemes that involve getting around receiving the direct consent of end-users. In other words, traffic is routed via gadgets of people that does not fully understand or even know at all that the third party is using their IP address. For compliance teams of companies operating in regulated industries, this is not just an ethical question – it is an issue of data supply chain, which is actually the same as a question regarding the origin of personal data.

If a regulating organ or auditor asks, “Where do the IP addresses a company gathers competitors’ prices with come from?”, an answer “We do not know” sounds worse than the absence of an answer altogether.

That is why the question of origin and authenticity of a proxy pool evolved from being a technical detail to the object of careful examination. 

What legal aspects does a CFO have to understand 

There are several things to keep in mind while building a data strategy – things that even a person without legal education can and should remember.

  • Data publicity is not equal to the absence of limitations 

Courts (for example, in the hiQ vs. LinkedIn case) acknowledged the right to scrape publicly available data without authorization. At the same time, though, platforms’ ToS and contracts create another layer of legal responsibility that does not depend on the technical “publicity” of data. 

  • Aggressive measures to get over protection systems increase legal risks

Getting over CAPTCHA, ignoring robot.txt files, and user-agent string faking – all are arguments that plaintiffs use to prove bad faith. This drastically distorts the image of a company that scraped data. 

  • GDPR and other laws are still to be obeyed

GDPR and other regulations are still active, even for public data. LinkedIn profiles, reviews with names, contacts – according to GDPR, those are personal data, regardless of the source. 

  • Additional regulations for the financial industry

MNPI (material non-public information), market manipulation rules, and MiFID II/SEC requirements regarding data quality for investment decisions meant that alternative data must have a documented, audited supply chain. 

  • Pharma faces pricing transparency and anti-competitive concerns 

Tracking competitors’ pricing is a legitimate practice. However, if the method of obtaining that data raises suspicions and looks untransparent, it may serve as evidence while investigating unfair competition. Actually, pharma deserves more attention, so there is a thorough article on why pharma and healthcare enterprises need secure IPs

The conclusion is simple: legal risk isn’t only in what data you obtain, but also in how you obtain that data and whether you can prove that your actions are legal. Ethically derived proxies for web scraping here are a must. 

What criteria should legal and compliance teams use to evaluate a proxy provider? 

When you decide whether to trust a particular vendor, it is not the price per GB that you should check first, but the following points: 

  • The transparency of IP origin. A vendor must confirm that residential IPs are obtained in an ethical way – after receiving direct consent from end users, so they opted in to allow the use of their IPs. No hidden SDKs in free apps must be involved. Here lies the difference between using a legal tool and becoming a potential accomplice in an intrusion case. 
  • Certifications and audits. ISO and log audits for a particular time period are what your legal team may need in 18 months, when a regulator will ask to show how exactly you obtained that data in March of the previous year.
  • Sticking to data protection laws. For financial and pharmaceutical companies operating in the EU, it is crucial that a provider sticks to GDPR and guarantees traffic routing and processing according to requirements. 

Why expensive proxies end up being cheaper 

At first glance, the comparison between a cheap proxy pool and an enterprise-level provider may not look advantageous for the latter. However, consider the price of developers’ time wasted on retries when cheap proxies get blocked, which is typical while working with low-quality providers. Add here the price of legal consultations for risk assessment and the price of insurance coverage. Calculate the price of delay in product launch. 

At the end of the day, enterprise-level solutions with clear IP origin and certifications almost always turn out to be budget-wise in the long run, than cheap providers – even if the initial price is higher. 

What should the CFO and Compliance check before deciding on a provider 

  1. Check the consent policy. Your compliance team needs to evaluate it before you start working with a vendor, not after. 
  2. Clarify whether a vendor provides log information for internal audits. 
  3. Calculate estimated costs, taking into consideration costs per retry and insurance coverage, not just traffic. 
  4. Include a proxy provider in the third-party risk vendors list, just like cloud hosting or a payment processor. 

To sum up 

Data is a competitive advantage impossible to ignore, but for CFOs and legal councils in regulated industries, the question is not about whether to use such data, but about how to prove a business obtained and used the data legally. When the data origin and the supporting tools origin are transparent and disclosed, when there are clear policies that guarantee sticking to data regulations, scraping turns from the source of risk into a manageable, predictable, and cheap process.

Frequently Asked Questions

Is web scraping legal?

Yes, web scraping is generally a legal practice. Yet, a lot depends on what niche you operate in, what tools you use, what data you scrape, and how you store it. To remain on the legal side, consider data regulations active in your location and platforms' ToS, use tools with transparent policies, and remember limitations specific to your niche.

If the data is publicly available, how can limitations apply?

Data may be public, but still be personal - for example, LinkedIn profiles, reviews with names, contacts - all of that may be classified as personal data, according to GDPR. Also, visibility does not remove the necessity to process, store, and share data according to regulations. That is why even for public data, limitations are still present, and caution is necessary.

Can a business be held responsible if a proxy provider’s IPs are not ethically derived or raise any other concerns?

There is a possibility of it. Especially if a company knew about the legal issues beforehand and knowingly used such proxies. That is why making sure IPs were derived with users' consent is a must. To avoid issues, choose legal providers like DataImpulse that do not resell proxies and only derive IPs with users' consent.

What is DataImpulse?

DataImpulse (dataimpulse.com) is an ethical proxy vendor that offers 90M+ residential, 16M+ mobile, and 20M+ datacenter IPs from 195 countries. The provider operates on a pay-per-GB pricing and sells non-expiring traffic, with residential proxies costing $1/GB. DataImpulse's proxies are GDPR-compliant, thus suitable for legal web scraping, ad verification, SERP monitoring, price tracking, and other use cases.

When is DataImpulse not an option?

DataImpulse can not help with accessing the government or bank sources. There is also no scraping API. Worth mentioning, the vendor does not provide addresses from Iran, Cuba, Belarus, the Russian Federation, North Korea, Syria, and temporarily occupied parts of Ukraine.

Does using a legal proxy provider eliminate all legal risks?

No. It significantly reduces exposure, but legal risks depend on a lot of factors, such as target platforms' rules, types of data collected, and the way you store and process them. A proxy vendor does not influence those factors; instead, it addresses the infrastructure layer of risk. In other words, a reputable proxy provider is necessary, but it is not everything.

Share article: