Scrapy Proxy Setup: Authentication, Rotation and Error Handling
Scrapy proxy setup starts with the downloader, where each request receives its proxy configuration before being sent. A useful crawler also needs bounded retries, per-domain concurrency, and diagnostics that distinguish proxy failures from destination responses. This guide uses the built-in HTTP/1.1 download handler and current Scrapy 2.19 settings. The asynchronous start() method shown here was introduced in 2.13; older projects may still use start_requests().
For a lighter script, compare the Python proxies overview and Requests proxy tutorial. Scrapy is useful when scheduling, extraction, and crawl statistics belong in one framework.
Set one authenticated proxy on a request
Install Scrapy in your project’s environment and check its version with scrapy version. Supply BP_PROXY_HOST, BP_PROXY_PORT, BP_PROXY_USER, BP_PROXY_PASSWORD, and BP_TARGET_URL through the environment. The target should initially be an endpoint you control.
import os
from urllib.parse import quote
import scrapy
class ProxyCheckSpider(scrapy.Spider):
name = "proxy_check"
custom_settings = {
"DOWNLOAD_TIMEOUT": 20,
"DOWNLOAD_VERIFY_CERTIFICATES": True,
"ROBOTSTXT_OBEY": True,
"RETRY_TIMES": 1,
}
async def start(self):
user = quote(os.environ["BP_PROXY_USER"], safe="")
password = quote(os.environ["BP_PROXY_PASSWORD"], safe="")
host = os.environ["BP_PROXY_HOST"]
port = int(os.environ["BP_PROXY_PORT"])
proxy = f"http://{user}:{password}@{host}:{port}"
yield scrapy.Request(os.environ["BP_TARGET_URL"],
meta={"proxy": proxy})
def parse(self, response):
yield {"status": response.status}
The enabled HttpProxyMiddleware interprets the proxy metadata and handles credentials in the URL, as its official implementation shows. Encode the two credential components separately, once. Do not put a proxy password in a destination Authorization header. The authentication guide explains endpoint authorization, and the proxy formatter helps check URL structure.
Scrapy can also read http_proxy, https_proxy, and no_proxy. An explicit meta["proxy"] takes precedence and ignores no_proxy for that request. Keep one deliberate configuration path while diagnosing routing. The middleware reference documents that precedence.
Match the proxy scheme to the handler
An http:// proxy URL can carry requests to an https:// destination. The downloader asks the proxy to establish a CONNECT tunnel, then negotiates destination TLS through it. An https:// proxy URL instead requests TLS on the connection to the proxy itself. Choose the scheme supplied for the endpoint rather than copying the destination’s scheme.
Support depends on the download handler. Scrapy’s HTTP/1.1 handler supports HTTP proxies for HTTPS destinations, while its HTTPS-proxy support is more limited. The built-in HTTP/2 handler does not support the proxy metadata. SOCKS support exists with the optional HTTPX handler in current documentation; it is not a feature of the HTTP/1.1 example here. Consult the download handler reference before swapping handlers.
Set DOWNLOAD_VERIFY_CERTIFICATES=True on current Scrapy: its documented default is false. Older 2.13 deployments instead need a verifying TLS context factory, such as scrapy.core.downloader.contextfactory.BrowserLikeContextFactory via DOWNLOADER_CLIENTCONTEXTFACTORY. Check your installed version’s settings before copying TLS configuration.
Assign a proxy before HttpProxyMiddleware
For a simple pool, select an endpoint in downloader middleware with order 740, before the built-in proxy middleware at 750. Supply a nonempty list of encoded authenticated proxy URLs through your project’s BP_PROXY_URLS setting, populated from your secret configuration. The component builder calls from_crawler() to construct this middleware.
from itertools import cycle
class ProxyPoolMiddleware:
@classmethod
def from_crawler(cls, crawler):
proxies = crawler.settings.getlist("BP_PROXY_URLS")
if not proxies:
raise ValueError("BP_PROXY_URLS must contain at least one URL")
return cls(proxies)
def __init__(self, proxies):
self.pool = cycle(enumerate(proxies))
def process_request(self, request):
index, proxy = next(self.pool)
request.meta["proxy"] = proxy
request.meta["bp_proxy_id"] = f"pool-{index}"
Register the class’s actual module path in DOWNLOADER_MIDDLEWARES, for example {"myproject.middlewares.ProxyPoolMiddleware": 740}. The downloadable script below builds its pool from separate endpoint and credential environment variables instead.
This round-robin policy selects again when a retry enters the downloader. It is a starting pattern, not a health-aware allocator: it does not quarantine unreachable endpoints or enforce a per-proxy connection budget. Track failures against bp_proxy_id if you extend it. Avoid selecting a proxy inside parse(), because requests for robots files and other downloader work also need routing.
For authenticated browsing or a workflow that binds cookies to an egress address, pin a proxy to that session instead. Rotating an address while reusing session state can break the workflow. Decide which unit stays stable before implementing the pool.
Bound retries and destination load
DOWNLOAD_TIMEOUT = 20
CONCURRENT_REQUESTS = 4
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 1
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_TARGET_CONCURRENCY = 1
RETRY_TIMES = 1
RETRY_HTTP_CODES = [408, 500, 502, 503, 504]
One retry means at most two download attempts for a request, excluding other separately scheduled work. Scrapy’s retry implementation handles configured transient responses and supported download exceptions; it does not establish that another proxy will solve the failure. The deliberate status list above excludes 407 and 429. Authentication needs correction; rate limiting needs a separate delay or rescheduling policy that considers Retry-After.
AutoThrottle adjusts delay using observed latency and respects configured limits. It does not make a proxy pool a reason to increase pressure on a destination. Keep robots handling enabled and choose crawl permissions and request rates appropriate to the target. See AutoThrottle and the settings reference for their scope.
Inspect failures without exposing credentials
| Signal | Useful next action |
|---|---|
| HTTP 407 or failed CONNECT | Check proxy authentication, endpoint details, and allowed destination port. |
| Download timeout or connection refusal | Test endpoint reachability separately, then compare per-proxy failures. |
| HTTP 403 | Identify whether the response came from the destination or an intermediary. |
| HTTP 429 | Reduce destination load and apply an explicit rate-limit policy. |
| Response absent from parse() | Check HTTP status filtering, robots decisions, offsite rules, and the errback. |
Use an errback for exhausted download failures. Record the exception type, attempt count, and pool label; avoid printing the complete failure object, authenticated proxy URL, headers, or target query string. Scrapy’s normal logging may include request details, so review logging before processing sensitive URLs. The downloadable script disables framework logging and prints a small diagnostic summary.
To inspect a non-success HTTP response, allow its status explicitly using handle_httpstatus_list, or use handle_httpstatus_all for a diagnostic request. Those options affect delivery to callbacks; they do not make an unsuccessful status a success. The request/response reference explains metadata and errbacks. The proxy error guide covers the wider troubleshooting flow.
Run the downloadable spider
Download the Scrapy proxy spider. It reads BP_PROXY_ENDPOINTS_JSON as a JSON array of credential-free URLs such as ["http://proxy.example:8080"], plus BP_PROXY_USER, BP_PROXY_PASSWORD, and BP_TARGET_URL. It validates endpoint format with python -B scrapy-proxy-spider.py --check. Remove --check to run the crawler in an environment with Scrapy 2.19 installed.
Compare the egress reported by your controlled target and review failures with one endpoint before enlarging the pool. The proxy tester provides an additional endpoint check. A completed crawl process can contain failed requests, so inspect its summaries rather than treating process completion as proof of successful downloads.
Frequently asked questions
Does setting HTTPPROXY_ENABLED rotate addresses?
No. It enables the built-in proxy middleware. Your application or another middleware must choose each endpoint.
Should I copy all request metadata to follow-up requests?
Avoid copying it indiscriminately. Internal metadata such as retry counts can affect the new request’s behavior. Transfer only the state that belongs to that follow-up.
Will a proxy bypass robots handling?
No. Proxy routing and robots decisions are separate. Keep robots handling configured for the crawl.
Can I reuse this configuration in HTTPX or aiohttp?
The concepts carry over, but the APIs differ. Use the HTTPX guide, aiohttp guide, or developer examples hub for their client settings.
For a pool with endpoint details you can assign explicitly, review BuyProxies dedicated proxy plans and size your crawl around the target’s request limits.
Official sources
- Scrapy downloader middleware
- Scrapy settings
- Scrapy download handlers
- Scrapy requests and responses
- Scrapy spider start method
- Scrapy AutoThrottle
- Scrapy 2.13 certificate verification settings
- Scrapy proxy middleware implementation
- Scrapy retry middleware implementation
- Scrapy component construction
Technical references reviewed October 2, 2026. See our editorial policy and testing methodology.
