Big Data Open Source Projects
Browse 319 Big Data open source projects, ranked by GitHub stars. Find the most popular Big Data tools and libraries.
binhnguyennus/awesome-scalability
The Patterns of Scalable, Reliable, and Performant Large-Scale Systems
Metrics details
| Stars | 72,527 |
ClickHouse/ClickHouse
ClickHouse® is a real-time analytics database management system
Metrics details
| Stars | 48,745 |
apache/spark
Apache Spark - A unified analytics engine for large-scale data processing
Metrics details
| Stars | 43,662 |
donnemartin/data-science-ipython-notebooks
Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.
Metrics details
| Stars | 29,248 |
apache/flink
Apache Flink
Metrics details
| Stars | 26,197 |
n0shake/Public-APIs
📚 A public list of APIs from round the web.
Metrics details
| Stars | 23,656 |
amark/gun
An open source cybersecurity protocol for syncing decentralized graph data.
Metrics details
| Stars | 19,076 |
questdb/questdb
QuestDB is a high performance, open-source, time-series database
Metrics details
| Stars | 17,185 |
heibaiying/BigData-Notes
大数据入门指南 :star:
Metrics details
| Stars | 16,922 |
prestodb/presto
The official home of the Presto distributed SQL query engine for big data
Metrics details
| Stars | 16,719 |
andkret/Cookbook
The Data Engineering Cookbook
Metrics details
| Stars | 15,180 |
trinodb/trino
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Metrics details
| Stars | 13,048 |
apache/predictionio
PredictionIO, a machine learning server for developers and ML engineers.
Metrics details
| Stars | 12,521 |
vesoft-inc/nebula
A distributed, fast open-source graph database featuring horizontal scalability and high availability
Metrics details
| Stars | 12,299 |
provectus/kafka-ui
Open-Source Web UI for Apache Kafka Management
Metrics details
| Stars | 12,231 |
yahoo/CMAK
CMAK is a tool for managing Apache Kafka clusters
Metrics details
| Stars | 11,926 |
StarRocks/starrocks
The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.
Metrics details
| Stars | 11,912 |
quickwit-oss/quickwit
Cloud-native OSS search engine for observability
Metrics details
| Stars | 11,420 |
cython/cython
The most widely used Python to C compiler
Metrics details
| Stars | 10,801 |
risingwavelabs/risingwave
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
Metrics details
| Stars | 9,176 |
catboost/catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
Metrics details
| Stars | 9,030 |
apache/arrow-datafusion
Apache Arrow DataFusion SQL Query Engine
Metrics details
| Stars | 8,975 |
delta-io/delta
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
Metrics details
| Stars | 8,917 |
apache/beam
Apache Beam is a unified programming model for Batch and Streaming data processing.
Metrics details
| Stars | 8,636 |
h2oai/h2o-3
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.
Metrics details
| Stars | 7,499 |
arkime/arkime
Arkime is an open source, large scale, full packet capturing, indexing, and database system.
Metrics details
| Stars | 7,415 |
feast-dev/feast
The Open Source Feature Store for AI/ML
Metrics details
| Stars | 7,142 |
vespa-engine/vespa
The AI search platform
Metrics details
| Stars | 7,021 |
apache/couchdb
Seamless multi-primary syncing database with an intuitive HTTP/JSON API, designed for reliability
Metrics details
| Stars | 6,929 |
apache/zeppelin
Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.
Metrics details
| Stars | 6,644 |
hazelcast/hazelcast
Hazelcast is a unified real-time data platform combining stream processing with a fast data store, allowing customers to act instantly on data-in-motion for real-time insights.
Metrics details
| Stars | 6,591 |
apache/iotdb
Apache IoTDB
Metrics details
| Stars | 6,368 |
pachyderm/pachyderm
Data-Centric Pipelines and Data Versioning
Metrics details
| Stars | 6,297 |
apache/hive
Apache Hive
Metrics details
| Stars | 5,994 |
Eventual-Inc/Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Metrics details
| Stars | 5,641 |
microsoft/SynapseML
Simple and Distributed Machine Learning
Metrics details
| Stars | 5,231 |
apache/calcite
Apache Calcite
Metrics details
| Stars | 5,158 |
apache/ignite
Apache Ignite
Metrics details
| Stars | 5,073 |
tschellenbach/Stream-Framework
Stream Framework is a Python library, which allows you to build news feed, activity streams and notification systems using Cassandra and/or Redis. The authors of Stream-Framework also provide a cloud service for feed technology:
Metrics details
| Stars | 4,744 |
tangbc/vue-virtual-scroll-list
⚡️A vue component support big amount data list with high render performance and efficient.
Metrics details
| Stars | 4,512 |
rom1504/img2dataset
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
Metrics details
| Stars | 4,435 |
crate/crate
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
Metrics details
| Stars | 4,416 |
alibaba/fastjson2
🚄 FASTJSON2 is a Java JSON library with excellent performance.
Metrics details
| Stars | 4,320 |
Moataz-Elmesmary/Data-Science-Roadmap
Data Science Roadmap from A to Z
Metrics details
| Stars | 4,311 |
alibaba/GraphScope
🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计算系统
Metrics details
| Stars | 3,556 |
databricks/koalas
Koalas: pandas API on Apache Spark
Metrics details
| Stars | 3,371 |
apache/incubator-paimon
Apache Paimon(incubating) is a streaming data lake platform that supports high-speed data ingestion, change data tracking and efficient real-time analytics.
Metrics details
| Stars | 3,334 |
root-project/root
The official repository for ROOT: analyzing, storing and visualizing big data, scientifically
Metrics details
| Stars | 3,258 |
lakesoul-io/LakeSoul
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.
Metrics details
| Stars | 3,243 |
apache/incubator-hugegraph
A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)
Metrics details
| Stars | 3,127 |
TuiQiao/CBoard
An easy to use, self-service open BI reporting and BI dashboard platform.
Metrics details
| Stars | 3,095 |
apache/parquet-mr
Apache Parquet
Metrics details
| Stars | 3,069 |
alldatacenter/alldata
🔥🔥 AllData可定义数据中台,以数据平台为底座,以数据中台为桥梁,以机器学习平台为工厂,以大模型应用为上游产品,提供全链路数字化解决方案。产品正式演示体验、社群咨询、商务采购:https://docs.qq.com/doc/DVHlkSEtvVXVCdEFo
Metrics details
| Stars | 3,053 |
apache/flume
Mirror of Apache Flume
Metrics details
| Stars | 2,565 |
FeatureBaseDB/featurebase
A crazy fast analytical database, built on bitmaps. Perfect for ML applications. Learn more at: http://docs.featurebase.com/. Start a Docker instance: https://hub.docker.com/r/featurebasedb/featurebase
Metrics details
| Stars | 2,525 |
apache/parquet-format
Apache Parquet Format
Metrics details
| Stars | 2,498 |
man-group/ArcticDB
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
Metrics details
| Stars | 2,429 |
jostmey/NakedTensor
Bare bone examples of machine learning in TensorFlow
Metrics details
| Stars | 2,403 |
apache/ambari
Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.
Metrics details
| Stars | 2,306 |
ytsaurus/ytsaurus
YTsaurus is a scalable and fault-tolerant open-source big data platform.
Metrics details
| Stars | 2,198 |
