Web scraping and crawling

13 repos

Libraries and frameworks for automated web data extraction, including both general-purpose scrapers and specialized tools that combine crawling with AI-powered content processing. The cluster spans multiple languages (Python, Go, Ruby) and includes established frameworks like Scrapy alongside newer projects integrating LLMs for intelligent data extraction, making it useful for anyone building web automation, data collection, or content parsing systems.

Python · 6
Go · 3
JavaScript · 1
Jupyter Notebook · 1
Ruby · 1
TypeScript · 1
web-scraping ·369,255
crawler ·363,378
scraping ·360,936
web-scraper ·296,541
ai-scraping ·289,815
webscraping ·289,815
data-extraction ·289,815
ai ·259,206
llm ·216,608
web-crawler ·210,942