Cuda Open Source Projects
Browse 1,016 Cuda open source projects, ranked by GitHub stars. Find the most popular Cuda tools and libraries.
nagadomi/waifu2x
Image Super-Resolution for Anime-Style Art
Metrics details
| Stars | 28,206 |
hashcat/hashcat
World's fastest and most advanced password recovery utility
Metrics details
| Stars | 26,356 |
NVIDIA/nvidia-docker
Build and run Docker containers leveraging NVIDIA GPUs
Metrics details
| Stars | 17,581 |
NVlabs/instant-ngp
Instant neural graphics primitives: lightning fast NeRF and more
Metrics details
| Stars | 17,494 |
kaldi-asr/kaldi
kaldi-asr/kaldi is the official location of the Kaldi project.
Metrics details
| Stars | 15,431 |
vosen/ZLUDA
CUDA on non-NVIDIA GPUs
Metrics details
| Stars | 14,621 |
cyrildiagne/ar-cutpaste
Cut and paste your surroundings using AR
Metrics details
| Stars | 14,578 |
isl-org/Open3D
Open3D: A Modern Library for 3D Data Processing
Metrics details
| Stars | 13,808 |
NVIDIA/TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
Metrics details
| Stars | 13,167 |
srush/GPU-Puzzles
Solve puzzles. Learn CUDA.
Metrics details
| Stars | 12,331 |
cupy/cupy
NumPy & SciPy for GPU
Metrics details
| Stars | 12,170 |
DefTruth/cuda-learn-note
🎉CUDA 笔记 / 高频面试题汇总 / C++笔记,个人笔记,更新随缘: sgemm、sgemv、warp reduce、block reduce、dot product、elementwise、softmax、layernorm、rmsnorm、hist etc.
Metrics details
| Stars | 11,533 |
jobbole/awesome-cpp-cn
C++ 资源大全中文版,标准库、Web应用框架、人工智能、数据库、图片处理、机器学习、日志、代码分析等。由「开源前哨」和「CPP开发者」微信公号团队维护更新。
Metrics details
| Stars | 11,150 |
numba/numba
NumPy aware dynamic Python compiler using LLVM
Metrics details
| Stars | 11,083 |
NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
Metrics details
| Stars | 10,106 |
xmrig/xmrig
RandomX, KawPow, CryptoNight and GhostRider unified CPU/GPU miner and RandomX benchmark
Metrics details
| Stars | 10,046 |
rapidsai/cudf
cuDF - GPU DataFrame Library
Metrics details
| Stars | 9,704 |
replicate/cog
Containers for machine learning
Metrics details
| Stars | 9,445 |
Oneflow-Inc/oneflow
OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.
Metrics details
| Stars | 9,411 |
NVIDIA/cuda-samples
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
Metrics details
| Stars | 9,404 |
catboost/catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
Metrics details
| Stars | 9,030 |
NVIDIA/apex
A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch
Metrics details
| Stars | 8,986 |
kroma-network/tachyon
Modular ZK(Zero Knowledge) backend accelerated by GPU
Metrics details
| Stars | 7,660 |
hybridgroup/gocv
Go package for computer vision using OpenCV 4 and beyond. Includes support for DNN, CUDA, OpenCV Contrib, and OpenVINO.
Metrics details
| Stars | 7,473 |
XuehaiPan/nvitop
An interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management.
Metrics details
| Stars | 7,058 |
open-mmlab/mmcv
OpenMMLab Computer Vision Foundation
Metrics details
| Stars | 6,452 |
luanfujun/deep-painterly-harmonization
Code and data for paper "Deep Painterly Harmonization": https://arxiv.org/abs/1804.03189
Metrics details
| Stars | 6,044 |
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
Metrics details
| Stars | 5,983 |
ethereum-mining/ethminer
Ethereum miner with OpenCL, CUDA and stratum support
Metrics details
| Stars | 5,930 |
gorgonia/gorgonia
Gorgonia is a library that helps facilitate machine learning in Go.
Metrics details
| Stars | 5,922 |
chainer/chainer
A flexible framework of neural networks for deep learning
Metrics details
| Stars | 5,919 |
autumnai/leaf
Open Machine Intelligence Framework for Hackers. (GPU/CPU)
Metrics details
| Stars | 5,543 |
chrxh/alien
ALIEN is a CUDA-powered artificial life simulation program.
Metrics details
| Stars | 5,454 |
shader-slang/slang
Making it easier to work with shaders
Metrics details
| Stars | 5,444 |
nerfstudio-project/gsplat
CUDA accelerated rasterization of gaussian splatting
Metrics details
| Stars | 5,410 |
rapidsai/cuml
cuML - RAPIDS Machine Learning Library
Metrics details
| Stars | 5,230 |
NVIDIAGameWorks/kaolin
A PyTorch Library for Accelerating 3D Deep Learning Research
Metrics details
| Stars | 5,141 |
NVIDIA/thrust
[ARCHIVED] The C++ parallel algorithms library. See https://github.com/NVIDIA/cccl
Metrics details
| Stars | 5,004 |
arrayfire/arrayfire
ArrayFire: a general purpose GPU library.
Metrics details
| Stars | 4,896 |
NVIDIA/nccl
Optimized primitives for collective multi-GPU communication
Metrics details
| Stars | 4,893 |
MegEngine/MegEngine
MegEngine 是一个快速、可拓展、易于使用且支持自动求导的深度学习框架
Metrics details
| Stars | 4,807 |
senguptaumd/Background-Matting
Background Matting: The World is Your Green Screen
Metrics details
| Stars | 4,767 |
OpenNMT/CTranslate2
Fast inference engine for Transformer models
Metrics details
| Stars | 4,577 |
OAID/Tengine
Tengine is a lite, high performance, modular inference engine for embedded device
Metrics details
| Stars | 4,526 |
NVlabs/tiny-cuda-nn
Lightning fast C++/CUDA neural network framework
Metrics details
| Stars | 4,513 |
leoxiaobin/deep-high-resolution-net.pytorch
The project is an official implementation of our CVPR2019 paper "Deep High-Resolution Representation Learning for Human Pose Estimation"
Metrics details
| Stars | 4,479 |
wookayin/gpustat
📊 A simple command-line utility for querying and monitoring GPU status
Metrics details
| Stars | 4,387 |
ROCm/HIP
HIP: C++ Heterogeneous-Compute Interface for Portability
Metrics details
| Stars | 4,375 |
pytorch/serve
Serve, optimize and scale PyTorch models in production
Metrics details
| Stars | 4,350 |
baidu-research/warp-ctc
Fast parallel CTC.
Metrics details
| Stars | 4,067 |
princeton-vl/RAFT
Metrics details
| Stars | 4,065 |
openxla/iree
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Metrics details
| Stars | 3,840 |
hujie-frank/SENet
Squeeze-and-Excitation Networks
Metrics details
| Stars | 3,641 |
NVIDIA/TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Metrics details
| Stars | 3,435 |
MrNeRF/gaussian-splatting-cuda
3D Gaussian Splatting, reimagined: Unleashing unmatched speed with C++ and CUDA from the ground up!
Metrics details
| Stars | 3,389 |
nihui/opencv-mobile
The minimal opencv for Android, iOS, ARM Linux, Windows, Linux, MacOS, HarmonyOS, WebAssembly, watchOS, tvOS, visionOS
Metrics details
| Stars | 3,324 |
Celtoys/Remotery
Single C file, Realtime CPU/GPU Profiler with Remote Web Viewer
Metrics details
| Stars | 3,307 |
bytedance/lightseq
LightSeq: A High Performance Library for Sequence Processing and Generation
Metrics details
| Stars | 3,296 |
roflcoopter/viseron
Self-hosted, local only NVR and AI Computer Vision software. With features such as object detection, motion detection, face recognition and more, it gives you the power to keep an eye on your home, office or any other place you want to monitor.
Metrics details
| Stars | 3,288 |
Jittor/jittor
Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
Metrics details
| Stars | 3,227 |
