Crawl Open Source Projects
Browse 198 Crawl open source projects, ranked by GitHub stars. Find the most popular Crawl tools and libraries.
kangvcar/InfoSpider
INFO-SPIDER 是一个集众多数据源于一身的爬虫工具箱🧰,旨在安全快捷的帮助用户拿回自己的数据,工具代码开源,流程透明。支持数据源包括GitHub、QQ邮箱、网易邮箱、阿里邮箱、新浪邮箱、Hotmail邮箱、Outlook邮箱、京东、淘宝、支付宝、中国移动、中国联通、中国电信、知乎、哔哩哔哩、网易云音乐、QQ好友、QQ群、生成朋友圈相册、浏览器浏览历史、12306、博客园、CSDN博客、开源中国博客、简书。
Metrics details
| Stars | 8,240 |
lc/gau
Fetch known URLs from AlienVault's Open Threat Exchange, the Wayback Machine, and Common Crawl.
Metrics details
| Stars | 5,042 |
201206030/novel-plus
novel-plus 是一个多端(PC、WAP)阅读 、功能完善的小说 CMS 系统。包括小说推荐、小说检索、小说排行、小说阅读、小说书架、小说评论、小说爬虫、会员中心、作家专区、充值订阅、新闻发布等功能。
Metrics details
| Stars | 4,670 |
yasserg/crawler4j
Open Source Web Crawler for Java
Metrics details
| Stars | 4,620 |
hanc00l/wooyun_public
This repo is archived. Thanks for wooyun! 乌云公开漏洞、知识库爬虫和搜索 crawl and search for wooyun.org public bug(vulnerability) and drops
Metrics details
| Stars | 4,402 |
kajweb/dict
英语字典 英语词库 字典词库 四级单词 六级单词 考研单词 雅思 托福 SAT GMAT TOEFL GRE
Metrics details
| Stars | 3,371 |
wkunzhi/Python3-Spider
Python爬虫实战 - 模拟登陆各大网站 包含但不限于:滑块验证、拼多多、美团、百度、bilibili、大众点评、淘宝,如果喜欢请start ❤️
Metrics details
| Stars | 3,371 |
internetarchive/heritrix3
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.
Metrics details
| Stars | 3,280 |
jaeles-project/gospider
Gospider - Fast web spider written in Go
Metrics details
| Stars | 2,985 |
crawl/crawl
Dungeon Crawl: Stone Soup official repository
Metrics details
| Stars | 2,919 |
spatie/crawler
https://spatie.be/docs/crawler
Metrics details
| Stars | 2,828 |
spatie/laravel-sitemap
Create and generate sitemaps with ease
Metrics details
| Stars | 2,618 |
fhamborg/news-please
news-please - an integrated web crawler and information extractor for news that just works
Metrics details
| Stars | 2,470 |
sjdirect/abot
Cross Platform C# web crawler framework built for speed and flexibility. Please star this project! +1.
Metrics details
| Stars | 2,309 |
ReaJason/xhs
基于小红书 Web 端进行的请求封装。https://reajason.github.io/xhs/
Metrics details
| Stars | 2,182 |
munificent/hauberk
A web-based roguelike written in Dart.
Metrics details
| Stars | 2,161 |
PuerkitoBio/gocrawl
Polite, slim and concurrent web crawler.
Metrics details
| Stars | 2,054 |
yahoo/gryffin
Gryffin is a large scale web security scanning platform.
Metrics details
| Stars | 2,052 |
coder-hxl/x-crawl
Flexible Node.js AI-assisted crawler library
Metrics details
| Stars | 1,873 |
eldraco/domain_analyzer
Analyze the security of any domain by finding all the information possible. Made in python.
Metrics details
| Stars | 1,868 |
hu17889/go_spider
[爬虫框架 (golang)] An awesome Go concurrent Crawler(spider) framework. The crawler is flexible and modular. It can be expanded to an Individualized crawler easily or you can use the default crawl components only.
Metrics details
| Stars | 1,821 |
DanMcInerney/xsscrapy
XSS spider - 66/66 wavsep XSS detected
Metrics details
| Stars | 1,743 |
thecodrr/fdir
⚡ The fastest directory crawler & globbing library for NodeJS. Crawls 1m files in < 1s
Metrics details
| Stars | 1,725 |
geelen/react-snapshot
A zero-configuration static pre-renderer for React apps
Metrics details
| Stars | 1,656 |
chriskite/anemone
Anemone web-spider framework
Metrics details
| Stars | 1,602 |
markdalgleish/static-site-generator-webpack-plugin
Minimal, unopinionated static site generator powered by webpack
Metrics details
| Stars | 1,602 |
ArchiveTeam/grab-site
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
Metrics details
| Stars | 1,600 |
sethblack/python-seo-analyzer
An SEO tool that analyzes the structure of a site, crawls the site, count words in the body of the site and warns of any technical SEO issues.
Metrics details
| Stars | 1,462 |
darbra/sperm
浏览过的精彩逆向文章汇总,值得一看
Metrics details
| Stars | 1,405 |
zhuweiyou/weixin-game-helper
微信小游戏辅助合集(加减大师、包你懂我、大家来找茬腾讯版、头脑王者、好友画我、悦动音符、我最在行、星途WeGoing、猜画小歌、知乎答题王、腾讯中国象棋、跳一跳、题多多黄金版)
Metrics details
| Stars | 1,399 |
mvdbos/php-spider
A configurable and extensible PHP web spider
Metrics details
| Stars | 1,345 |
scrapinghub/frontera
A scalable frontier for web crawlers
Metrics details
| Stars | 1,332 |
istresearch/scrapy-cluster
This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.
Metrics details
| Stars | 1,225 |
timwhitez/crawlergo_x_XRAY
360/0Kee-Team/crawlergo动态爬虫结合长亭XRAY扫描器的被动扫描功能
Metrics details
| Stars | 1,185 |
DoTheEvo/ANGRYsearch
Linux file search, instant results as you type
Metrics details
| Stars | 1,156 |
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
Metrics details
| Stars | 1,087 |
LoseNine/Crack-JS-Spider
JS破解逆向,破解JS反爬虫加密参数,已破解极验滑块w(2022.2.19),QQ音乐sign(2022.2.13),拼多多anti_content,boss直聘zp_token,知乎x-zse-96,酷狗kg_mid/dfid,唯品会mars_cid,中国裁判文书网(2020-06-30更新),淘宝密码,天安保险登录,b站登录,房天下登录,WPS登录,微博登录,有道翻译,网易登录,微信公众号登录,空中网登录,今目标登录,学生信息管理系统登录,共赢金融登录,重庆科技资源共享平台登录,网易云音乐下载,一键解析视频链接,财联社登录。
Metrics details
| Stars | 1,028 |
fredwu/crawler
A high performance web crawler / scraper in Elixir.
Metrics details
| Stars | 958 |
liuhuanyong/PersonRelationKnowledgeGraph
ChinesePersonRelationGraph, person relationship extraction based on nlp methods.中文人物关系知识图谱项目,内容包括中文人物关系图谱构建,基于知识库的数据回标,基于远程监督与bootstrapping方法的人物关系抽取,基于知识图谱的知识问答等应用。
Metrics details
| Stars | 932 |
WayneDW/Sentiment-Analysis-in-Event-Driven-Stock-Price-Movement-Prediction
Use NLP to predict stock price movement associated with news
Metrics details
| Stars | 882 |
soskek/bookcorpus
Crawl BookCorpus
Metrics details
| Stars | 863 |
openzim/zimit
Make a ZIM file from any Web site and surf offline!
Metrics details
| Stars | 813 |
mattsse/voyager
crawl and scrape web pages in rust
Metrics details
| Stars | 767 |
arkadiyt/bounty-targets
This project crawls bug bounty platform scopes (like Hackerone/Bugcrowd/Intigriti/etc) hourly and dumps them into the bounty-targets-data repo
Metrics details
| Stars | 720 |
tj/staticgen
Static website generator that lets you use HTTP servers and frameworks you already know
Metrics details
| Stars | 717 |
anvaka/word2vec-graph
Exploring word2vec embeddings as a graph of nearest neighbors
Metrics details
| Stars | 712 |
rugantio/fbcrawl
A Facebook crawler
Metrics details
| Stars | 691 |
bit4woo/domain_hunter
A Burp Suite Extension that try to find all sub-domain, similar-domain and related-domain of an organization automatically! 基于流量自动收集整个企业或组织的子域名、相似域名、相关域名的burp插件
Metrics details
| Stars | 674 |
oooldtoy/SSTAP_ip_crawl_tool
一个自动获取游戏远程ip,并自动写成SSTAP/NETCH规则文件的脚本
Metrics details
| Stars | 673 |
hartleybrody/public-amazon-crawler
Metrics details
| Stars | 673 |
liip/TheA11yMachine
The A11y Machine is an automated accessibility testing tool which crawls and tests pages of any web application to produce detailed reports.
Metrics details
| Stars | 639 |
yacy/yacy_grid_crawler
Crawler Microservice for the YaCy Grid
Metrics details
| Stars | 633 |
philschmid/clipper.js
HTML to Markdown converter and crawler.
Metrics details
| Stars | 625 |
markowanga/stweet
Advanced python library to scrap Twitter (tweets, users) from unofficial API
Metrics details
| Stars | 623 |
s0md3v/Orbit
Blockchain Transactions Investigation Tool
Metrics details
| Stars | 615 |
spatie/http-status-check
CLI tool to crawl a website and check HTTP status codes
Metrics details
| Stars | 600 |
dataapiman/data-api
(更新)数据接口,淘宝(带精确预售量、精确月销量),拼多多,小红书,微信公众号,大众点评,快手,京东,饿了么,B站,知乎,微博,Bigo,TEMU,得物、贝壳,shopee,百度指数,等数据接口;大模型训练预料
Metrics details
| Stars | 531 |
scalingexcellence/scrapybook
Scrapy Book Code
Metrics details
| Stars | 510 |
commoncrawl/commoncrawl
Common Crawl support library to access 2008-2012 crawl archives (ARC files)
Metrics details
| Stars | 508 |
blackfireio/player
Blackfire Player is a powerful Web Crawling, Web Testing, and Web Scraper application. It provides a nice DSL to crawl HTTP services, assert responses, and extract data from HTML/XML/JSON responses.
Metrics details
| Stars | 495 |
