Big Data Open Source Projects

Browse 319 Big Data open source projects, ranked by GitHub stars. Find the most popular Big Data tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 319 projects
72,527 stars

binhnguyennus/awesome-scalability

The Patterns of Scalable, Reliable, and Performant Large-Scale Systems

Metrics details
Stars72,527
48,745 stars

ClickHouse/ClickHouse

ClickHouse® is a real-time analytics database management system

Metrics details
Stars48,745
43,662 stars

apache/spark

Apache Spark - A unified analytics engine for large-scale data processing

Metrics details
Stars43,662
29,248 stars

donnemartin/data-science-ipython-notebooks

Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

Metrics details
Stars29,248
26,197 stars

apache/flink

Apache Flink

Metrics details
Stars26,197
23,656 stars

n0shake/Public-APIs

📚 A public list of APIs from round the web.

Metrics details
Stars23,656
19,076 stars

amark/gun

An open source cybersecurity protocol for syncing decentralized graph data.

Metrics details
Stars19,076
17,185 stars

questdb/questdb

QuestDB is a high performance, open-source, time-series database

Metrics details
Stars17,185
16,922 stars

heibaiying/BigData-Notes

大数据入门指南 :star:

Metrics details
Stars16,922
16,719 stars

prestodb/presto

The official home of the Presto distributed SQL query engine for big data

Metrics details
Stars16,719
15,180 stars

andkret/Cookbook

The Data Engineering Cookbook

Metrics details
Stars15,180
13,048 stars

trinodb/trino

Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)

Metrics details
Stars13,048
12,521 stars

apache/predictionio

PredictionIO, a machine learning server for developers and ML engineers.

Metrics details
Stars12,521
12,299 stars

vesoft-inc/nebula

A distributed, fast open-source graph database featuring horizontal scalability and high availability

Metrics details
Stars12,299
12,231 stars

provectus/kafka-ui

Open-Source Web UI for Apache Kafka Management

Metrics details
Stars12,231
11,926 stars

yahoo/CMAK

CMAK is a tool for managing Apache Kafka clusters

Metrics details
Stars11,926
11,912 stars

StarRocks/starrocks

The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.

Metrics details
Stars11,912
11,420 stars

quickwit-oss/quickwit

Cloud-native OSS search engine for observability

Metrics details
Stars11,420
10,801 stars

cython/cython

The most widely used Python to C compiler

Metrics details
Stars10,801
9,176 stars

risingwavelabs/risingwave

Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.

Metrics details
Stars9,176
9,030 stars

catboost/catboost

A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.

Metrics details
Stars9,030
8,975 stars

apache/arrow-datafusion

Apache Arrow DataFusion SQL Query Engine

Metrics details
Stars8,975
8,917 stars

delta-io/delta

An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs

Metrics details
Stars8,917
8,636 stars

apache/beam

Apache Beam is a unified programming model for Batch and Streaming data processing.

Metrics details
Stars8,636
7,499 stars

h2oai/h2o-3

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

Metrics details
Stars7,499
7,415 stars

arkime/arkime

Arkime is an open source, large scale, full packet capturing, indexing, and database system.

Metrics details
Stars7,415
7,142 stars

feast-dev/feast

The Open Source Feature Store for AI/ML

Metrics details
Stars7,142
7,021 stars

vespa-engine/vespa

The AI search platform

Metrics details
Stars7,021
6,929 stars

apache/couchdb

Seamless multi-primary syncing database with an intuitive HTTP/JSON API, designed for reliability

Metrics details
Stars6,929
6,644 stars

apache/zeppelin

Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.

Metrics details
Stars6,644
6,591 stars

hazelcast/hazelcast

Hazelcast is a unified real-time data platform combining stream processing with a fast data store, allowing customers to act instantly on data-in-motion for real-time insights.

Metrics details
Stars6,591
6,368 stars

apache/iotdb

Apache IoTDB

Metrics details
Stars6,368
6,297 stars

pachyderm/pachyderm

Data-Centric Pipelines and Data Versioning

Metrics details
Stars6,297
5,994 stars

apache/hive

Apache Hive

Metrics details
Stars5,994
5,641 stars

Eventual-Inc/Daft

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

Metrics details
Stars5,641
5,231 stars

microsoft/SynapseML

Simple and Distributed Machine Learning

Metrics details
Stars5,231
5,158 stars

apache/calcite

Apache Calcite

Metrics details
Stars5,158
5,073 stars

apache/ignite

Apache Ignite

Metrics details
Stars5,073
4,744 stars

tschellenbach/Stream-Framework

Stream Framework is a Python library, which allows you to build news feed, activity streams and notification systems using Cassandra and/or Redis. The authors of Stream-Framework also provide a cloud service for feed technology:

Metrics details
Stars4,744
4,512 stars

tangbc/vue-virtual-scroll-list

⚡️A vue component support big amount data list with high render performance and efficient.

Metrics details
Stars4,512
4,435 stars

rom1504/img2dataset

Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.

Metrics details
Stars4,435
4,416 stars

crate/crate

CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.

Metrics details
Stars4,416
4,320 stars

alibaba/fastjson2

🚄 FASTJSON2 is a Java JSON library with excellent performance.

Metrics details
Stars4,320
4,311 stars

Moataz-Elmesmary/Data-Science-Roadmap

Data Science Roadmap from A to Z

Metrics details
Stars4,311
3,556 stars

alibaba/GraphScope

🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计算系统

Metrics details
Stars3,556
3,371 stars

databricks/koalas

Koalas: pandas API on Apache Spark

Metrics details
Stars3,371
3,334 stars

apache/incubator-paimon

Apache Paimon(incubating) is a streaming data lake platform that supports high-speed data ingestion, change data tracking and efficient real-time analytics.

Metrics details
Stars3,334
3,258 stars

root-project/root

The official repository for ROOT: analyzing, storing and visualizing big data, scientifically

Metrics details
Stars3,258
3,243 stars

lakesoul-io/LakeSoul

LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.

Metrics details
Stars3,243
3,127 stars

apache/incubator-hugegraph

A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)

Metrics details
Stars3,127
3,095 stars

TuiQiao/CBoard

An easy to use, self-service open BI reporting and BI dashboard platform.

Metrics details
Stars3,095
3,069 stars

apache/parquet-mr

Apache Parquet

Metrics details
Stars3,069
3,053 stars

alldatacenter/alldata

🔥🔥 AllData可定义数据中台,以数据平台为底座,以数据中台为桥梁,以机器学习平台为工厂,以大模型应用为上游产品,提供全链路数字化解决方案。产品正式演示体验、社群咨询、商务采购:https://docs.qq.com/doc/DVHlkSEtvVXVCdEFo

Metrics details
Stars3,053
2,565 stars

apache/flume

Mirror of Apache Flume

Metrics details
Stars2,565
2,525 stars

FeatureBaseDB/featurebase

A crazy fast analytical database, built on bitmaps. Perfect for ML applications. Learn more at: http://docs.featurebase.com/. Start a Docker instance: https://hub.docker.com/r/featurebasedb/featurebase

Metrics details
Stars2,525
2,498 stars

apache/parquet-format

Apache Parquet Format

Metrics details
Stars2,498
2,429 stars

man-group/ArcticDB

ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.

Metrics details
Stars2,429
2,403 stars

jostmey/NakedTensor

Bare bone examples of machine learning in TensorFlow

Metrics details
Stars2,403
2,306 stars

apache/ambari

Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.

Metrics details
Stars2,306
2,198 stars

ytsaurus/ytsaurus

YTsaurus is a scalable and fault-tolerant open-source big data platform.

Metrics details
Stars2,198
1-60 of 319 projects
Get A Weekly Email With Trending Big Data Projects
Stay updated on Big Data plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.