In this Article
A 403 Forbidden during web scraping means the server understood your request perfectly and refused to serve it. That distinction matters, because a 403 is a decision rather than a failure, and decisions have causes you can identify instead of retrying blindly.
403 forbidden web scraping problems are therefore a diagnosis exercise, not a retry exercise. This guide explains what actually produces a 403 for a crawler, gives a five-step ladder for finding which cause applies to you, lists the fixes that work at each step, and is direct about the cases where a 403 is the final word and the correct next move is an official API.
Key Facts
- 403 means refused, not broken. The server understood your request and declined to serve it, which makes it a decision you can diagnose rather than a bug you can retry away.
- Four different causes look identical. Authorisation, network reputation, client identity and request pacing all surface as the same status code, so the fix depends entirely on which one applies.
- Triage in five steps: classify the body, test authorisation, isolate the network, isolate the client, isolate the pace. Change one variable at a time.
- Our test, 2026-09-17: two of four vendor sites returned 403 to a plain datacenter request and 200 through a residential exit, while the other two returned 200 on both routes. The same status code, two different causes.
- Sometimes 403 is the answer. Behind a login, after repeated challenges, or on personal data, the correct response is an official API or a data licence rather than another retry strategy.
What does 403 Forbidden actually mean in scraping?
403 is the server saying the request was valid and it will not be fulfilled. Unlike a 404, the resource usually exists. Unlike a 429, no rate limit is being announced. Unlike a 401, no credentials are being requested. The server has simply decided that this client does not get this resource.
For a crawler, four very different situations collapse into that one status code.
| Cause | Typical fingerprint | What actually fixes it |
|---|---|---|
| Authorisation | Consistent 403 on a specific path, works while logged in | Permission or an API, not engineering |
| Network reputation | 403 from one address type, 200 from another, identical request | An exit address that matches the identity you claim |
| Client identity | 403 regardless of address, small body, no challenge script | TLS, header and runtime coherence |
| Pacing and pattern | Works at first, then 403 after a burst or a sweep | Lower concurrency, jitter, fewer requests |
Because the symptom is shared, guessing is expensive. The ladder below exists to turn the guess into a measurement.
How do you diagnose it with the 5-step triage ladder?
The 5-step 403 triage ladder changes one variable at a time and stops as soon as the cause is identified.
Step one, read the body. A few hundred bytes with a short message is a flat refusal. A full page with a script and a token is a challenge, which means the door is open but your client cannot answer. A 200 with an empty data container is a third case entirely, where you passed and the content simply never rendered.
Step two, test authorisation. Open the same URL in an ordinary browser without a session. If it refuses there too, the content is not public and no crawler configuration will change that.
Step three, isolate the network. Send the identical request from a different address type. If the result changes, the cause is reputation. If it does not, stop buying proxies and move on.
Step four, isolate the client. From the same address, repeat through a real browser engine. If the browser succeeds where the library failed, the cause is TLS, header or runtime identity.
Step five, isolate the pace. Drop to a single worker with jitter and a small page count. If that succeeds, the cause was volume, and rotating addresses at the original rate will only spread the problem.
Two rules make the ladder work: change one variable per test, and record the status code and body size for every run. Teams that change five things at once learn nothing from the result.
What did our own test show about the same status code?
On 17 September 2026 we sent one GET request per cell, identical browser-like headers, no JavaScript engine, comparing a datacenter address with a DataImpulse residential exit in the United States. The targets were four public vendor marketing pages.
| Target | Datacenter route | Residential route | Cause indicated |
|---|---|---|---|
| datadome.co | 403, 771 bytes | 200, 446 KB | Network reputation |
| humansecurity.com | 403, 5.7 KB | 200, 1.99 MB | Network reputation |
| kasada.io | 200, 47 KB | 200, 47 KB | Not gated at this layer |
| imperva.com | 200, 249 KB | 200, 249 KB | Not gated at this layer |
Two 403s with identical status codes, one cause identified in a single test, and two sites where the same change would have been wasted money. Twenty minutes later the residential route to humansecurity.com returned 403 as well, which is the other lesson: a route that works is a state, not a property.
Caveats: one request per cell, marketing pages rather than protected endpoints, no JavaScript engine, and results move over time.
Which fixes work at each step?
Match the fix to the cause you identified. Applying all of them at once is how teams end up maintaining a fragile stack that nobody understands.
Authorisation. Ask for access, use the documented API, or accept that the data is not available to you. There is no configuration fix for a permission decision.
Network reputation. Use an exit type that matches your claim: residential for consumer-facing pages, datacenter for APIs and open data, mobile for app-first targets. Bind one session to one exit rather than rotating mid-session. Our guide to proxy rotation best practices covers the session model, and static versus rotating proxies covers the choice.
Client identity. Make every layer agree: TLS stack, header set and order, HTTP version, cookie and redirect handling, and a real runtime where the site expects one. Wiring is in our Playwright proxy tutorial and Puppeteer guide.
Pacing. Low single-digit concurrency per host, jitter between requests, exponential backoff on the first rejection, a circuit breaker that halts the job, and caching so that repeat reads never leave your machine.
If the 403 came from a named protection, the vendor-specific guides go deeper: DataDome, PerimeterX and HUMAN, Kasada, and Imperva and Incapsula.
Which exit type should you use once you know the cause?
Only after the ladder points at network reputation does the exit type become the decision. Use this table then, and not before.
| Situation | Use | Avoid when |
|---|---|---|
| Consumer-facing public pages, moderate volume | Rotating residential, one session per logical task | The workflow depends on a token issued earlier, which rotation invalidates |
| Multi-step flows, search state, authorised logins | Sticky residential for the length of the flow | Single-request reads, where sticky costs more and buys nothing |
| APIs, documentation, open data portals | Datacenter | The host treats hosting ranges as automation, which our test showed is common on consumer sites |
| App-first and carrier-sensitive endpoints | Mobile | Budget is tight and the target does not distinguish carrier traffic |
| 403 persists identically across every exit type | Stop and fix the client | Always: more addresses cannot fix a client-identity problem |
DataImpulse residential is $1 per GB across 195 countries with country targeting included, HTTP, HTTPS and SOCKS5, rotating or sticky sessions, and state, city, ZIP or ASN targeting on the paid tier. We do not sell static ISP addresses and we are not a managed unblocking API. The wider comparison is in best proxies for web scraping.
What mistakes make a 403 worse?
Four habits turn a recoverable situation into a burned target.
- Retrying immediately and often. A refusal answered with fifty retries per minute converts a soft classification into a durable one, and it affects every other user of your address pool.
- Rotating addresses instead of diagnosing. If the cause was client identity, rotation multiplies the evidence against you across a whole pool rather than fixing anything.
- Copying a header dump once and freezing it. Browser header sets change with every release, so last year’s perfect capture becomes this year’s anomaly.
- Ignoring robots.txt and the terms. Beyond the ethics, crawling paths a site has asked you to avoid is the fastest way to earn the reputation that causes the next 403.
The general pattern behind all four is optimising for getting through today rather than for a pipeline that still works next quarter.
When is a 403 the final answer?
Sometimes the correct engineering decision is to stop. Treat a 403 as final when the content sits behind a login or paywall, when a documented API covers the same fields, when challenges keep appearing rather than blocks, when the data includes personal information you have no lawful basis to process, or when the only remaining ideas target the protection itself rather than your own client.
Terms of service are a contract, unauthorised access to systems is a criminal matter in many jurisdictions including under the Computer Fraud and Abuse Act in the United States, and personal data brings the GDPR and similar regimes into scope regardless of how public the page looked. This is general information rather than legal advice; our overview of whether web scraping is legal has the longer version.
We do not cover CAPTCHA solving, challenge reverse engineering or token replay, and we would not recommend building a pipeline that depends on them: they carry the legal risk and they break first when a vendor updates.
Frequently Asked Questions
What does 403 Forbidden mean when scraping?
It means the server understood your request and refused to serve it. The resource usually exists, no credentials are being requested, and no rate limit is being announced. For a crawler the cause is one of four things: authorisation, network reputation, client identity, or request pacing.
How do I fix a 403 when scraping?
Diagnose before you change anything. Read the response body to tell a flat refusal from a challenge, check whether the page is public at all in a normal browser, then test the same request from a different address type, then from a real browser engine, then at one worker with jitter. Each test isolates one cause, and the fix follows from which test changed the outcome.
Why do I get 403 with a browser user agent?
Because the user agent is a single string among many signals. Your TLS handshake, header order, HTTP version, cookie handling and pacing are read together, and a browser user agent on a Python TLS stack is a contradiction. Consistency across layers is what matters, not any one value.
Does changing IP fix a 403?
Only when the cause is network reputation. In our test on 17 September 2026 a different exit flipped 403 to 200 on two of four sites and changed nothing on the other two. Test the address as one variable before you build a rotation strategy around it.
What is the difference between 403 and 429?
A 429 announces that you exceeded a rate limit and usually tells you when to retry, so the fix is pacing. A 403 makes no such promise: it is a refusal that can come from permissions, reputation, client identity or pattern. Treating a 403 as a rate limit and simply waiting is why many crawlers stall for hours without learning anything.
Should I use a headless browser to avoid 403s?
Only if the site requires script execution or checks runtime evidence. A browser engine is heavier and slower per page, so use it where it is needed and keep an HTTP client for endpoints that do not need it. If a library gets a 403 and a browser does not, the cause was client identity.
Diagnose first, then choose the exit
If the triage points at network reputation, DataImpulse residential proxies give geo-accurate exits at $1 per GB across 195 countries, rotating or sticky, with HTTP, HTTPS and SOCKS5 and country targeting included. Create an account and test a single target before scaling, or see the setup patterns in our web scraping use case.
Last updated: September 17, 2026.

State/City/Zip/ASN Targeting 



