Proxy choice for crawling comes down to how much scrutiny the targets apply. Paying for residential traffic against a site that never checks is waste; using datacentre addresses against one that does is a failed job.
Planning A Crawl
- Sort targets by how strictly they filter, then assign address types.
- Estimate in bytes including sub-resources, not in page counts.
- Track complete records parsed, not HTTP 200s.
Matching Address Type To Target
Datacentre space is cheap, fast and entirely adequate for public endpoints, documentation, and sites with no interest in filtering. Residential traffic costs more per gigabyte and is what survives against retail, travel and marketplace sites that grade their visitors. Most real pipelines want both, sorted by target, rather than one type applied uniformly.
Building The Volume Estimate
Pages per day multiplied by realistic page weight multiplied by running days. Realistic means including images, scripts and fonts unless they are explicitly blocked, which typically puts actual transfer several times above the HTML size. An estimate built from HTML alone will be short by a wide margin and the shortfall always appears mid-project.
Measuring Success Properly
Success rate by HTTP status overstates how well a crawl is going, because degraded responses return 200. Measure the share of fetches that parse into complete records, and track bytes per complete record alongside it. Those two together show both cost efficiency and data quality, and they move before a hard block appears.
Where Budget Actually Goes
In most crawls the majority of transferred bytes are not the data being collected. Images, fonts, analytics and third-party scripts routinely account for the bulk of it, which means the largest available saving is a resource-blocking rule rather than a cheaper plan. It is worth measuring that split once before comparing providers, because it commonly changes the required volume by a factor large enough to move which tier you should be buying.
Per Gigabyte
Volume pricing suits breadth-first collection.
Large Exit Pool
Spread load widely so per-address limits are not the ceiling.
Geo Selection
Country and city targeting for localised result sets.
⬇ Download
Get Started →