Crawl Open Source Projects

Browse 198 Crawl open source projects, ranked by GitHub stars. Find the most popular Crawl tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 198 projects
8,240 stars

kangvcar/InfoSpider

INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。

Metrics details
Stars8,240
5,042 stars

lc/gau

Fetch known URLs from AlienVault's Open Threat Exchange, the Wayback Machine, and Common Crawl.

Metrics details
Stars5,042
4,670 stars

201206030/novel-plus

novel-plus 是一个多端(PC、WAP)阅读 、功能完善的小说 CMS 系统。包括小说推荐、小说检索、小说排行、小说阅读、小说书架、小说评论、小说爬虫、会员中心、作家专区、充值订阅、新闻发布等功能。

Metrics details
Stars4,670
4,620 stars

yasserg/crawler4j

Open Source Web Crawler for Java

Metrics details
Stars4,620
4,402 stars

hanc00l/wooyun_public

This repo is archived. Thanks for wooyun! 乌云公开漏洞、知识库爬虫和搜索 crawl and search for wooyun.org public bug(vulnerability) and drops

Metrics details
Stars4,402
3,371 stars

kajweb/dict

英语字典 英语词库 字典词库 四级单词 六级单词 考研单词 雅思 托福 SAT GMAT TOEFL GRE

Metrics details
Stars3,371
3,371 stars

wkunzhi/Python3-Spider

Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️

Metrics details
Stars3,371
3,280 stars

internetarchive/heritrix3

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.

Metrics details
Stars3,280
2,985 stars

jaeles-project/gospider

Gospider - Fast web spider written in Go

Metrics details
Stars2,985
2,919 stars

crawl/crawl

Dungeon Crawl: Stone Soup official repository

Metrics details
Stars2,919
2,828 stars

spatie/crawler

https://spatie.be/docs/crawler

Metrics details
Stars2,828
2,618 stars

spatie/laravel-sitemap

Create and generate sitemaps with ease

Metrics details
Stars2,618
2,470 stars

fhamborg/news-please

news-please - an integrated web crawler and information extractor for news that just works

Metrics details
Stars2,470
2,309 stars

sjdirect/abot

Cross Platform C# web crawler framework built for speed and flexibility. Please star this project! +1.

Metrics details
Stars2,309
2,182 stars

ReaJason/xhs

基于小红书 Web 端进行的请求封装。https://reajason.github.io/xhs/

Metrics details
Stars2,182
2,161 stars

munificent/hauberk

A web-based roguelike written in Dart.

Metrics details
Stars2,161
2,054 stars

PuerkitoBio/gocrawl

Polite, slim and concurrent web crawler.

Metrics details
Stars2,054
2,052 stars

yahoo/gryffin

Gryffin is a large scale web security scanning platform.

Metrics details
Stars2,052
1,873 stars

coder-hxl/x-crawl

Flexible Node.js AI-assisted crawler library

Metrics details
Stars1,873
1,868 stars

eldraco/domain_analyzer

Analyze the security of any domain by finding all the information possible. Made in python.

Metrics details
Stars1,868
1,821 stars

hu17889/go_spider

[爬虫框架 (golang)] An awesome Go concurrent Crawler(spider) framework. The crawler is flexible and modular. It can be expanded to an Individualized crawler easily or you can use the default crawl components only.

Metrics details
Stars1,821
1,743 stars

DanMcInerney/xsscrapy

XSS spider - 66/66 wavsep XSS detected

Metrics details
Stars1,743
1,725 stars

thecodrr/fdir

⚡ The fastest directory crawler & globbing library for NodeJS. Crawls 1m files in < 1s

Metrics details
Stars1,725
1,656 stars

geelen/react-snapshot

A zero-configuration static pre-renderer for React apps

Metrics details
Stars1,656
1,602 stars

chriskite/anemone

Anemone web-spider framework

Metrics details
Stars1,602
1,602 stars

markdalgleish/static-site-generator-webpack-plugin

Minimal, unopinionated static site generator powered by webpack

Metrics details
Stars1,602
1,600 stars

ArchiveTeam/grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

Metrics details
Stars1,600
1,462 stars

sethblack/python-seo-analyzer

An SEO tool that analyzes the structure of a site, crawls the site, count words in the body of the site and warns of any technical SEO issues.

Metrics details
Stars1,462
1,405 stars

darbra/sperm

浏览过的精彩逆向文章汇总,值得一看

Metrics details
Stars1,405
1,399 stars

zhuweiyou/weixin-game-helper

微信小游戏辅助合集(加减大师、包你懂我、大家来找茬腾讯版、头脑王者、好友画我、悦动音符、我最在行、星途WeGoing、猜画小歌、知乎答题王、腾讯中国象棋、跳一跳、题多多黄金版)

Metrics details
Stars1,399
1,345 stars

mvdbos/php-spider

A configurable and extensible PHP web spider

Metrics details
Stars1,345
1,332 stars

scrapinghub/frontera

A scalable frontier for web crawlers

Metrics details
Stars1,332
1,225 stars

istresearch/scrapy-cluster

This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.

Metrics details
Stars1,225
1,185 stars

timwhitez/crawlergo_x_XRAY

360/0Kee-Team/crawlergo动态爬虫结合长亭XRAY扫描器的被动扫描功能

Metrics details
Stars1,185
1,156 stars

DoTheEvo/ANGRYsearch

Linux file search, instant results as you type

Metrics details
Stars1,156
1,087 stars

webrecorder/browsertrix-crawler

Run a high-fidelity browser-based web archiving crawler in a single Docker container

Metrics details
Stars1,087
1,028 stars

LoseNine/Crack-JS-Spider

JS破解逆向,破解JS反爬虫加密参数,已破解极验滑块w(2022.2.19),QQ音乐sign(2022.2.13),拼多多anti_content,boss直聘zp_token,知乎x-zse-96,酷狗kg_mid/dfid,唯品会mars_cid,中国裁判文书网(2020-06-30更新),淘宝密码,天安保险登录,b站登录,房天下登录,WPS登录,微博登录,有道翻译,网易登录,微信公众号登录,空中网登录,今目标登录,学生信息管理系统登录,共赢金融登录,重庆科技资源共享平台登录,网易云音乐下载,一键解析视频链接,财联社登录。

Metrics details
Stars1,028
958 stars

fredwu/crawler

A high performance web crawler / scraper in Elixir.

Metrics details
Stars958
932 stars

liuhuanyong/PersonRelationKnowledgeGraph

ChinesePersonRelationGraph, person relationship extraction based on nlp methods.中文人物关系知识图谱项目,内容包括中文人物关系图谱构建,基于知识库的数据回标,基于远程监督与bootstrapping方法的人物关系抽取,基于知识图谱的知识问答等应用。

Metrics details
Stars932
882 stars

WayneDW/Sentiment-Analysis-in-Event-Driven-Stock-Price-Movement-Prediction

Use NLP to predict stock price movement associated with news

Metrics details
Stars882
863 stars

soskek/bookcorpus

Crawl BookCorpus

Metrics details
Stars863
813 stars

openzim/zimit

Make a ZIM file from any Web site and surf offline!

Metrics details
Stars813
767 stars

mattsse/voyager

crawl and scrape web pages in rust

Metrics details
Stars767
720 stars

arkadiyt/bounty-targets

This project crawls bug bounty platform scopes (like Hackerone/Bugcrowd/Intigriti/etc) hourly and dumps them into the bounty-targets-data repo

Metrics details
Stars720
717 stars

tj/staticgen

Static website generator that lets you use HTTP servers and frameworks you already know

Metrics details
Stars717
712 stars

anvaka/word2vec-graph

Exploring word2vec embeddings as a graph of nearest neighbors

Metrics details
Stars712
691 stars

rugantio/fbcrawl

A Facebook crawler

Metrics details
Stars691
674 stars

bit4woo/domain_hunter

A Burp Suite Extension that try to find all sub-domain, similar-domain and related-domain of an organization automatically! 基于流量自动收集整个企业或组织的子域名、相似域名、相关域名的burp插件

Metrics details
Stars674
673 stars

oooldtoy/SSTAP_ip_crawl_tool

一个自动获取游戏远程ip,并自动写成SSTAP/NETCH规则文件的脚本

Metrics details
Stars673
673 stars

hartleybrody/public-amazon-crawler

Metrics details
Stars673
639 stars

liip/TheA11yMachine

The A11y Machine is an automated accessibility testing tool which crawls and tests pages of any web application to produce detailed reports.

Metrics details
Stars639
633 stars

yacy/yacy_grid_crawler

Crawler Microservice for the YaCy Grid

Metrics details
Stars633
625 stars

philschmid/clipper.js

HTML to Markdown converter and crawler.

Metrics details
Stars625
623 stars

markowanga/stweet

Advanced python library to scrap Twitter (tweets, users) from unofficial API

Metrics details
Stars623
615 stars

s0md3v/Orbit

Blockchain Transactions Investigation Tool

Metrics details
Stars615
600 stars

spatie/http-status-check

CLI tool to crawl a website and check HTTP status codes

Metrics details
Stars600
531 stars

dataapiman/data-api

(更新)数据接口,淘宝(带精确预售量、精确月销量),拼多多,小红书,微信公众号,大众点评,快手,京东,饿了么,B站,知乎,微博,Bigo,TEMU,得物、贝壳,shopee,百度指数,等数据接口;大模型训练预料

Metrics details
Stars531
510 stars

scalingexcellence/scrapybook

Scrapy Book Code

Metrics details
Stars510
508 stars

commoncrawl/commoncrawl

Common Crawl support library to access 2008-2012 crawl archives (ARC files)

Metrics details
Stars508
495 stars

blackfireio/player

Blackfire Player is a powerful Web Crawling, Web Testing, and Web Scraper application. It provides a nice DSL to crawl HTTP services, assert responses, and extract data from HTML/XML/JSON responses.

Metrics details
Stars495
1-60 of 198 projects
Get A Weekly Email With Trending Crawl Projects
Stay updated on Crawl plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.