Dataframe Open Source Projects
Browse 217 Dataframe open source projects, ranked by GitHub stars. Find the most popular Dataframe tools and libraries.
pola-rs/polars
Extremely fast Query Engine for DataFrames, written in Rust
Metrics details
| Stars | 39,053 |
Kanaries/pygwalker
PyGWalker: Turn your dataframe into an interactive UI for visual analysis
Metrics details
| Stars | 15,912 |
modin-project/modin
Modin: Scale your Pandas workflows by changing a single line of code
Metrics details
| Stars | 10,394 |
rapidsai/cudf
cuDF - GPU DataFrame Library
Metrics details
| Stars | 9,704 |
apache/arrow-datafusion
Apache Arrow DataFusion SQL Query Engine
Metrics details
| Stars | 8,975 |
vaexio/vaex
Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀
Metrics details
| Stars | 8,509 |
haifengl/smile
Statistical Machine Intelligence & Learning Engine
Metrics details
| Stars | 6,403 |
ujjwalkarn/DataSciencePython
common data analysis and machine learning tasks using python
Metrics details
| Stars | 5,799 |
Eventual-Inc/Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Metrics details
| Stars | 5,641 |
javascriptdata/danfojs
Danfo.js is an open source, JavaScript library providing high performance, intuitive, and easy to use data structures for manipulating and processing structured data.
Metrics details
| Stars | 5,051 |
sngyai/Sequoia
A股自动选股程序,实现了海龟交易法则、缠中说禅牛市买点,以及其他若干种技术形态
Metrics details
| Stars | 5,021 |
lk-geimfari/mimesis
Mimesis is a Python library for generating fake but realistic data in multiple languages and locales.
Metrics details
| Stars | 4,831 |
twopirllc/pandas-ta
Technical Analysis Indicators - Pandas TA is an easy to use Python 3 Pandas Extension with 130+ Indicators
Metrics details
| Stars | 4,337 |
jtablesaw/tablesaw
Java dataframe and visualization library
Metrics details
| Stars | 3,758 |
databricks/koalas
Koalas: pandas API on Apache Spark
Metrics details
| Stars | 3,371 |
adamerose/PandasGUI
A GUI for Pandas DataFrames
Metrics details
| Stars | 3,259 |
pydata/pandas-datareader
Extract data from a wide range of Internet sources into a pandas DataFrame.
Metrics details
| Stars | 3,229 |
fbdesignpro/sweetviz
Visualize and compare datasets, target values and associations, with one line of code.
Metrics details
| Stars | 3,115 |
ajcr/100-pandas-puzzles
100 data puzzles for pandas, ranging from short and simple to super tricky (60% complete)
Metrics details
| Stars | 2,982 |
hosseinmoein/DataFrame
C++ DataFrame for statistical, financial, and ML analysis in modern C++
Metrics details
| Stars | 2,967 |
scikit-learn-contrib/sklearn-pandas
Pandas integration with sklearn
Metrics details
| Stars | 2,849 |
mars-project/mars
Mars is a tensor-based unified framework for large-scale data computation which scales numpy, pandas, scikit-learn and Python functions.
Metrics details
| Stars | 2,741 |
jmcarpenter2/swifter
A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner
Metrics details
| Stars | 2,640 |
sfu-db/connector-x
Fastest library to load data from DB to DataFrames in Rust and Python
Metrics details
| Stars | 2,636 |
DAGWorks-Inc/hamilton
Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage and metadata. Runs and scales everywhere python does.
Metrics details
| Stars | 2,545 |
chezou/tabula-py
Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame
Metrics details
| Stars | 2,315 |
approximatelabs/sketch
AI code-writing assistant that understands data content
Metrics details
| Stars | 2,283 |
peerchemist/finta
Common financial technical indicators implemented in Pandas.
Metrics details
| Stars | 2,262 |
justmarkham/pandas-videos
Jupyter notebook and datasets from the pandas video series
Metrics details
| Stars | 2,250 |
ballista-compute/ballista
Distributed compute platform implemented in Rust, and powered by Apache Arrow.
Metrics details
| Stars | 2,244 |
alexhallam/tv
📺(tv) Tidy Viewer is a cross-platform CLI csv pretty printer that uses column styling to maximize viewer enjoyment.
Metrics details
| Stars | 2,162 |
apache/arrow-ballista
Apache Arrow Ballista Distributed Query Engine
Metrics details
| Stars | 2,081 |
shramos/Awesome-Cybersecurity-Datasets
A curated list of amazingly awesome Cybersecurity datasets
Metrics details
| Stars | 2,044 |
wrobstory/vincent
A Python to Vega translator
Metrics details
| Stars | 2,023 |
AutoViML/AutoViz
Automatically Visualize any dataset, any size with a single line of code. Created by Ram Seshadri. Collaborators Welcome. Permission Granted upon Request.
Metrics details
| Stars | 1,905 |
dask/dask-tutorial
Dask tutorial
Metrics details
| Stars | 1,856 |
Lumiwealth/lumibot
Backtestable AI trading agents and Python algorithmic trading strategies for stocks, options, crypto, futures, forex, SEC filings, FRED macro data, and real brokers.
Metrics details
| Stars | 1,827 |
fonnesbeck/statistical-analysis-python-tutorial
Statistical Data Analysis in Python
Metrics details
| Stars | 1,736 |
alpacahq/marketstore
DataFrame Server for Financial Timeseries Data
Metrics details
| Stars | 1,687 |
uwdata/arquero
Query processing and transformation of array-backed data tables.
Metrics details
| Stars | 1,530 |
pyjanitor-devs/pyjanitor
Clean APIs for data cleaning. Python implementation of R package Janitor
Metrics details
| Stars | 1,500 |
michaelchu/optopsy
A nimble options backtesting library for Python
Metrics details
| Stars | 1,420 |
spark-examples/pyspark-examples
Pyspark RDD, DataFrame and Dataset Examples in Python language
Metrics details
| Stars | 1,363 |
holoviz/hvplot
A high-level plotting API for pandas, dask, xarray, and networkx built on HoloViews
Metrics details
| Stars | 1,355 |
JoinQuant/jqdatasdk
简单易用的量化金融数据包(easy utility for getting financial market data of China)
Metrics details
| Stars | 1,349 |
rocketlaunchr/dataframe-go
DataFrames for Go: For statistics, machine-learning, and data manipulation/exploration
Metrics details
| Stars | 1,288 |
palantir/pyspark-style-guide
This is a guide to PySpark code style presenting common situations and the associated best practices based on the most frequent recurring topics across the PySpark repos we've encountered.
Metrics details
| Stars | 1,254 |
Kotlin/kotlin-jupyter
Kotlin kernel for Jupyter/IPython
Metrics details
| Stars | 1,225 |
graphframes/graphframes
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
Metrics details
| Stars | 1,193 |
machow/siuba
Python library for using dplyr like syntax with pandas and SQL
Metrics details
| Stars | 1,184 |
8080labs/ppscore
Predictive Power Score (PPS) in Python
Metrics details
| Stars | 1,170 |
comet-ml/kangas
🦘 Explore multimedia datasets at scale
Metrics details
| Stars | 1,075 |
sharebook-kr/pykrx
KRX 주식 정보 스크래핑
Metrics details
| Stars | 1,062 |
Kotlin/dataframe
Kotlin DataFrame: typesafe in-memory structured data processing for JVM
Metrics details
| Stars | 1,052 |
freqtrade/technical
Various indicators developed or collected for the Freqtrade
Metrics details
| Stars | 1,020 |
mm-mansour/Fast-Pandas
Benchmark for different operations in pandas against various dataframe sizes.
Metrics details
| Stars | 959 |
RedisLabs/spark-redis
A connector for Spark that allows reading and writing to/from Redis cluster
Metrics details
| Stars | 947 |
microsoft/Mobius
C# and F# language binding and extensions to Apache Spark
Metrics details
| Stars | 947 |
kieferk/dfply
dplyr-style piping operations for pandas dataframes
Metrics details
| Stars | 895 |
stitchfix/hamilton
A scalable general purpose micro-framework for defining dataflows. THIS REPOSITORY HAS BEEN MOVED TO www.github.com/dagworks-inc/hamilton
Metrics details
| Stars | 860 |
