Quantization Open Source Projects

Browse 187 Quantization open source projects, ranked by GitHub stars. Find the most popular Quantization tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 187 projects
73,376 stars

hiyouga/LLaMA-Factory

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Metrics details
Stars73,376
24,368 stars

SYSTRAN/faster-whisper

Faster Whisper transcription with CTranslate2

Metrics details
Stars24,368
18,942 stars

ymcui/Chinese-LLaMA-Alpaca

中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)

Metrics details
Stars18,942
18,106 stars

UFund-Me/Qbot

[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs: https://ufund-me.github.io/Qbot ✨ :news: qbot-mini: https://github.com/Charmve/iQuant

Metrics details
Stars18,106
5,723 stars

kornelski/pngquant

Lossy PNG compressor — pngquant command based on libimagequant library

Metrics details
Stars5,723
5,703 stars

mozilla/mozjpeg

Improved JPEG encoder.

Metrics details
Stars5,703
5,072 stars

AutoGPTQ/AutoGPTQ

An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.

Metrics details
Stars5,072
4,577 stars

OpenNMT/CTranslate2

Fast inference engine for Transformer models

Metrics details
Stars4,577
4,506 stars

PINTO0309/PINTO_model_zoo

A repository for storing models that have been inter-converted between various frameworks. Supported frameworks are TensorFlow, PyTorch, ONNX, OpenVINO, TFJS, TFTRT, TensorFlowLite (Float32/16/INT8), EdgeTPU, CoreML.

Metrics details
Stars4,506
4,252 stars

IntelLabs/distiller

Neural Network Distiller by Intel AI Lab: a Python package for neural network compression research. https://intellabs.github.io/distiller

Metrics details
Stars4,252
3,982 stars

lucidrains/vector-quantize-pytorch

Vector (and Scalar) Quantization, in Pytorch

Metrics details
Stars3,982
3,479 stars

DingXiaoH/RepVGG

RepVGG: Making VGG-style ConvNets Great Again

Metrics details
Stars3,479
3,447 stars

huggingface/optimum

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

Metrics details
Stars3,447
3,162 stars

huawei-noah/Pretrained-Language-Model

Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.

Metrics details
Stars3,162
3,159 stars

neuralmagic/deepsparse

Sparsity-aware deep learning inference runtime for CPUs

Metrics details
Stars3,159
2,933 stars

IntelLabs/nlp-architect

A model library for exploring state-of-the-art deep learning topologies and techniques for optimizing Natural Language Processing neural networks

Metrics details
Stars2,933
2,908 stars

Tencent/PocketFlow

An Automatic Model Compression (AutoMC) framework for developing smaller and faster AI applications.

Metrics details
Stars2,908
2,714 stars

aaron-xichen/pytorch-playground

Base pretrained models and datasets in pytorch (MNIST, SVHN, CIFAR10, CIFAR100, STL10, AlexNet, VGG16, VGG19, ResNet, Inception, SqueezeNet)

Metrics details
Stars2,714
2,683 stars

intel/neural-compressor

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

Metrics details
Stars2,683
2,617 stars

quic/aimet

AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.

Metrics details
Stars2,617
2,406 stars

htqin/awesome-model-quantization

A list of papers, docs, codes about model quantization. This repo is aimed to provide the info for model quantization research, we are continuously improving the project. Welcome to PR the works (papers, repositories) that are missed by the repo.

Metrics details
Stars2,406
2,329 stars

dvmazur/mixtral-offloading

Run Mixtral-8x7B models in Colab or consumer desktops

Metrics details
Stars2,329
2,266 stars

666DZY666/micronet

micronet, a model compression and deploy lib. compression: 1、quantization: quantization-aware-training(QAT), High-Bit(>2b)(DoReFa/Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference)、Low-Bit(≤2b)/Ternary and Binary(TWN/BNN/XNOR-Net); post-training-quantization(PTQ), 8-bit(tensorrt); 2、 pruning: normal、regular and group convolutional channel pruning; 3、 group convolution structure; 4、batch-normalization fuse for quantization. deploy: tensorrt, fp32/fp16/int8(ptq-calibration)、op-adapt(upsample)、dynamic_shape

Metrics details
Stars2,266
2,121 stars

CesiumGS/gltf-pipeline

Content pipeline tools for optimizing glTF assets. :globe_with_meridians:

Metrics details
Stars2,121
2,061 stars

DanBloomberg/leptonica

Leptonica is an open source library containing software that is broadly useful for image processing and image analysis applications. The official github repository for Leptonica is: danbloomberg/leptonica. See leptonica.org for more documentation.

Metrics details
Stars2,061
2,014 stars

intel/intel-extension-for-pytorch

A Python package for extending the official PyTorch that can easily obtain performance on Intel platform

Metrics details
Stars2,014
1,806 stars

openppl-public/ppq

PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.

Metrics details
Stars1,806
1,675 stars

open-mmlab/mmrazor

OpenMMLab Model Compression Toolbox and Benchmark.

Metrics details
Stars1,675
1,611 stars

PaddlePaddle/PaddleSlim

PaddleSlim is an open-source library for deep model compression and architecture search.

Metrics details
Stars1,611
1,575 stars

tensorflow/model-optimization

A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.

Metrics details
Stars1,575
1,575 stars

saharNooby/rwkv.cpp

INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model

Metrics details
Stars1,575
1,553 stars

Xilinx/brevitas

Brevitas: neural network quantization in PyTorch

Metrics details
Stars1,553
1,404 stars

RahulSChand/gpu_poor

Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization

Metrics details
Stars1,404
1,304 stars

huawei-noah/Efficient-Computing

Efficient computing methods developed by Huawei Noah's Ark Lab

Metrics details
Stars1,304
1,274 stars

openvinotoolkit/training_extensions

Train, Evaluate, Optimize, Deploy Computer Vision Models via OpenVINO™

Metrics details
Stars1,274
1,182 stars

openvinotoolkit/nncf

Neural Network Compression Framework for enhanced OpenVINO™ inference

Metrics details
Stars1,182
1,109 stars

IST-DASLab/marlin

FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

Metrics details
Stars1,109
1,045 stars

huggingface/quanto

A pytorch Quantization Toolkit

Metrics details
Stars1,045
1,027 stars

Xilinx/finn

Dataflow compiler for QNN inference on FPGAs

Metrics details
Stars1,027
1,026 stars

gigwegbe/tinyml-papers-and-projects

This is a list of interesting papers and projects about TinyML.

Metrics details
Stars1,026
981 stars

PINTO0309/onnx2tf

A tool for converting ONNX files to LiteRT/TFLite/TensorFlow, PyTorch native code (nn.Module), TorchScript (.pt), state_dict (.pt), Exported Program (.pt2), and Dynamo ONNX. It also supports direct conversion from LiteRT to PyTorch.

Metrics details
Stars981
958 stars

mit-han-lab/TinyChatEngine

TinyChatEngine: On-Device LLM Inference Library

Metrics details
Stars958
951 stars

mit-han-lab/tinyengine

[NeurIPS 2020] MCUNet: Tiny Deep Learning on IoT Devices; [NeurIPS 2021] MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning; [NeurIPS 2022] MCUNetV3: On-Device Training Under 256KB Memory

Metrics details
Stars951
948 stars

mobiusml/hqq

Official implementation of Half-Quadratic Quantization (HQQ)

Metrics details
Stars948
921 stars

ImageOptim/libimagequant

Palette quantization library that powers pngquant and other PNG optimizers

Metrics details
Stars921
901 stars

OpenGVLab/OmniQuant

[ICLR2024 spotlight] OmniQuant is a simple and powerful quantization technique for LLMs.

Metrics details
Stars901
856 stars

guan-yuan/awesome-AutoML-and-Lightweight-Models

A list of high-quality (newest) AutoML works and lightweight models including 1.) Neural Architecture Search, 2.) Lightweight Structures, 3.) Model Compression, Quantization and Acceleration, 4.) Hyperparameter Optimization, 5.) Automated Feature Engineering.

Metrics details
Stars856
834 stars

Zhen-Dong/Awesome-Quantization-Papers

List of papers related to neural network quantization in recent AI conferences and journals.

Metrics details
Stars834
769 stars

csarron/awesome-emdl

Embedded and mobile deep learning research resources

Metrics details
Stars769
722 stars

SqueezeAILab/SqueezeLLM

[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization

Metrics details
Stars722
681 stars

DeepVAC/deepvac

PyTorch Project Specification.

Metrics details
Stars681
679 stars

hailo-ai/hailo_model_zoo

The Hailo Model Zoo includes pre-trained models and a full building and evaluation environment

Metrics details
Stars679
675 stars

square/gifencoder

A pure Java library implementing the GIF89a specification. Suitable for use on Android.

Metrics details
Stars675
671 stars

songhan/Deep-Compression-AlexNet

Deep Compression on AlexNet

Metrics details
Stars671
650 stars

SforAiDl/KD_Lib

A Pytorch Knowledge Distillation library for benchmarking and extending works in the domains of Knowledge Distillation, Pruning, and Quantization.

Metrics details
Stars650
647 stars

achuthasubhash/Complete-Life-Cycle-of-a-Data-Science-Project

Complete-Life-Cycle-of-a-Data-Science-Project

Metrics details
Stars647
630 stars

facebookresearch/kill-the-bits

Code for: "And the bit goes down: Revisiting the quantization of neural networks"

Metrics details
Stars630
606 stars

huggingface/optimum-intel

🤗 Optimum Intel: Accelerate inference with Intel optimization tools

Metrics details
Stars606
588 stars

Ki6an/fastT5

⚡ boost inference speed of T5 models by 5x & reduce the model size by 3x.

Metrics details
Stars588
585 stars

slavabarkov/tidy

Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art vision-language pretrained CLIP model and ONNX Runtime inference engine

Metrics details
Stars585
1-60 of 187 projects
Get A Weekly Email With Trending Quantization Projects
Stay updated on Quantization plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.