Skip to main content

Scrapy Proxy Headers

PyPI version Documentation

Send custom headers to proxies and receive proxy response headers in Scrapy.

The Problem

When making HTTPS requests through a proxy, Scrapy cannot send custom headers to the proxy itself. This is because HTTPS requests create an encrypted tunnel (via HTTP CONNECT) - any headers you add to request.headers are encrypted and only visible to the destination server, not the proxy.

┌──────────┐     CONNECT      ┌───────┐     Encrypted     ┌────────────┐
│  Scrapy  │ ───────────────► │ Proxy │ ════════════════► │ Target URL │
└──────────┘  (unencrypted)   └───────┘    (tunnel)       └────────────┘
                  │                              │
           Proxy headers             request.headers
           go HERE                   go here (encrypted)

This extension solves the problem by:

  1. Sending custom headers to the proxy during the CONNECT handshake
  2. Capturing response headers from the proxy's CONNECT response
  3. Making those headers available in your spider

Installation

pip install scrapy-proxy-headers

Quick Start

1. Configure the Download Handler

In your Scrapy settings.py:

DOWNLOAD_HANDLERS = {
    "https": "scrapy_proxy_headers.HTTP11ProxyDownloadHandler"
}

Or in your spider's custom_settings:

class MySpider(scrapy.Spider):
    custom_settings = {
        "DOWNLOAD_HANDLERS": {
            "https": "scrapy_proxy_headers.HTTP11ProxyDownloadHandler"
        }
    }

2. Send Proxy Headers

Use request.meta["proxy_headers"] to send headers to the proxy:

import scrapy

class MySpider(scrapy.Spider):
    name = "example"
    
    def start_requests(self):
        yield scrapy.Request(
            url="https://api.ipify.org?format=json",
            meta={
                "proxy": "http://your-proxy:port",
                "proxy_headers": {"X-ProxyMesh-Country": "US"}
            }
        )
    
    def parse(self, response):
        # Proxy response headers are available in response.headers
        proxy_ip = response.headers.get("X-ProxyMesh-IP")
        self.logger.info(f"Proxy IP: {proxy_ip}")

3. Receive Proxy Response Headers

Headers from the proxy's CONNECT response are automatically merged into response.headers:

def parse(self, response):
    # Access headers sent by the proxy
    proxy_ip = response.headers.get(b"X-ProxyMesh-IP")
    if proxy_ip:
        print(f"Request made through IP: {proxy_ip.decode()}")

Complete Example

import scrapy

class ProxyHeadersSpider(scrapy.Spider):
    name = "proxy_headers_demo"
    
    custom_settings = {
        "DOWNLOAD_HANDLERS": {
            "https": "scrapy_proxy_headers.HTTP11ProxyDownloadHandler"
        }
    }
    
    def start_requests(self):
        yield scrapy.Request(
            url="https://api.ipify.org?format=json",
            meta={
                "proxy": "http://us.proxymesh.com:31280",
                "proxy_headers": {"X-ProxyMesh-Country": "US"}
            },
            callback=self.parse_ip
        )
    
    def parse_ip(self, response):
        data = response.json()
        proxy_ip = response.headers.get(b"X-ProxyMesh-IP")
        
        self.logger.info(f"Public IP: {data['ip']}")
        if proxy_ip:
            self.logger.info(f"Proxy IP: {proxy_ip.decode()}")
        
        yield {
            "public_ip": data["ip"],
            "proxy_ip": proxy_ip.decode() if proxy_ip else None
        }

How It Works

  1. HTTP11ProxyDownloadHandler - Custom download handler that manages proxy header caching
  2. ScrapyProxyHeadersAgent - Agent that reads proxy_headers from request meta
  3. TunnelingHeadersAgent - Sends custom headers in the CONNECT request
  4. TunnelingHeadersTCP4ClientEndpoint - Captures proxy response headers from CONNECT response

The handler also caches proxy response headers by proxy URL. This ensures headers remain available even when Scrapy reuses existing tunnel connections for subsequent requests.

Test Harness

A test harness is included to verify proxy header functionality:

# Basic test
PROXY_URL=http://your-proxy:port TEST_URL=https://api.ipify.org python test_proxy_headers.py

# With custom proxy header
PROXY_URL=http://your-proxy:port \
PROXY_HEADER=X-ProxyMesh-IP \
SEND_PROXY_HEADER=X-ProxyMesh-Country \
SEND_PROXY_VALUE=US \
python test_proxy_headers.py

# Verbose output
python test_proxy_headers.py -v

Environment Variables

Variable Description Default
PROXY_URL Proxy URL (also checks HTTPS_PROXY) Required
TEST_URL URL to request https://api.ipify.org?format=json
PROXY_HEADER Response header to check for X-ProxyMesh-IP
SEND_PROXY_HEADER Header name to send to proxy Optional
SEND_PROXY_VALUE Value for the send header Optional

Documentation

Full documentation is available at scrapy-proxy-headers.readthedocs.io.

Use Cases

  • Geographic targeting: Send X-ProxyMesh-Country to route through specific countries
  • Session consistency: Request the same IP across multiple requests
  • Debugging: Capture proxy response headers to see which IP was assigned
  • Load balancing: Use proxy headers to control request distribution

Requirements

  • Python 3.8+
  • Scrapy 2.14.2+

License

BSD License - see LICENSE for details.

Links

Metadata

Release files for scrapy-proxy-headers 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scrapy-proxy-headers 0.2.1
File Size Uploaded
scrapy_proxy_headers-0.2.1.tar.gz 7.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scrapy-proxy-headers 0.2.1
File Interpreter ABI Platform
scrapy_proxy_headers-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 15.4 kB

Release files / scrapy_proxy_headers-0.2.1.tar.gz

Download URL scrapy_proxy_headers-0.2.1.tar.gz
Size 7.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f1a4f36d4cc3386da9eed10861aff5c9a8059a3e6d6a7e60ea6f2c77bd78ac25
BLAKE2b-256 checksum
How to use checksums
503c4f102e722999fdde3f7bb865812ad61778ac7a76e48278ad9ab7c1882ae8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.

Transparency log

Release files / scrapy_proxy_headers-0.2.1-py3-none-any.whl

Download URL scrapy_proxy_headers-0.2.1-py3-none-any.whl
Size 8.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
227766da401a612eaeaa756e008925056fb35ad81348ace25cd51625c0afa7da
BLAKE2b-256 checksum
How to use checksums
ece7aff3583b855b1acb521bebbf8cec42487daff022c29aca6fa306fca66cf3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page