head to head · open source
scrapy vs proxy_pool
scrapy has 64,388 GitHub stars, 11,962 forks, 394 open issues and last shipped yesterday. proxy_pool has 23,707 stars, 5,418 forks, 298 open issues and last shipped 3 months ago. scrapy leads on adoption by 172% (64,388 vs 23,707 stars). scrapy is written in Python under BSD-3-Clause; proxy_pool is written in Python under MIT. scrapy has attracted 19% as many forks as stars, proxy_pool 23%. scrapy was the more recently maintained of the two, and both are self-hostable with no licence fee. The two share 1 topic tag (crawler), so they are genuine substitutes rather than adjacent tools.
Two open source projects, one decision. Both are free and self-hostable — the differences are community size, license terms, language stack and release pace.
← all 8884 open source comparisons
Side by side
| scrapy | proxy_pool | |
|---|---|---|
| GitHub stars | ★ 64K | ★ 24K |
| License | BSD-3-Clause | MIT |
| Written in | Python | Python |
| Last push | 2026-09-17 | 2026-06-15 |
| Forks | ⑂ 12K | ⑂ 5.4K |
| Self-hosting | Yes | Yes |
| Data ownership | Your server | Your server |
pick scrapy if
- You weight community size — 64K stars and counting
- You want the BSD-3-Clause license terms
- Your stack matches Python
- You value the larger contributor base for long-term maintenance
pick proxy_pool if
- You want the proxy_pool feature set and don't need the biggest community
- You prefer the MIT license terms
- Your stack matches Python
- You evaluated both and proxy_pool fits your workflow better
About scrapy
Scrapy is a web scraping framework written in Python, distributed under the BSD 3 Clause license, and described in its own README as a tool to extract structured data from websites. It is cross platform and requires Python 3.10 or newer. The project is maintained by Zyte, formerly known as Scrapinghub, together with a broader group of contributors, and it lives in the Python ecosystem as an installable library rather than a hosted service. Its repository is roughly seventeen years old, which places it among the longer running projects in the web scraping space.
read the full scrapy overview →
About proxy_pool
proxy pool is a self hosted Python proxy IP pool for web spiders that continuously collects, checks, and serves rotating HTTP proxies from public free proxy sources through a Redis backed pool.
read the full proxy_pool overview →
More in Data & Analytics
Related comparisons
More Data Extraction & Web Scraping projects
Compare either of these against the rest of the Data Extraction & Web Scraping field.
Frequently asked questions
Is scrapy or proxy_pool more popular?
scrapy has 64,388 GitHub stars and proxy_pool has 23,707. scrapy has the larger community by that measure.
Are scrapy and proxy_pool free?
Both are open source. scrapy is licensed under BSD-3-Clause and proxy_pool under MIT. Neither carries a licence fee.
What is the difference between scrapy and proxy_pool?
scrapy is written in Python and proxy_pool in Python. The practical differences are community size, licence terms, language stack and release cadence — all compared in the table above.
Which should I choose, scrapy or proxy_pool?
Choose scrapy if you want the larger community (64,388 stars) or its BSD-3-Clause licence terms. Choose proxy_pool if its feature set, stack or MIT licence fits better. Both are self-hostable.