Clip Open Source Projects
Browse 100 Clip open source projects, ranked by GitHub stars. Find the most popular Clip tools and libraries.
OFA-Sys/Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
Metrics details
| Stars | 5,977 |
marqo-ai/marqo
Ecommerce Search and Discovery - marqo.ai
Metrics details
| Stars | 5,017 |
easychen/pushdeer
开放源码的无App推送服务,iOS14+扫码即用。亦支持快应用/iOS和Mac客户端、Android客户端、自制设备
Metrics details
| Stars | 5,014 |
open-compass/VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Metrics details
| Stars | 4,291 |
open-mmlab/mmpretrain
OpenMMLab Pre-training Toolbox and Benchmark
Metrics details
| Stars | 3,842 |
yuanzhoulvpi2017/zero_nlp
中文nlp解决方案(大模型、数据、模型、训练、推理)
Metrics details
| Stars | 3,830 |
jingyi0000/VLM_survey
Collection of AWESOME vision-language models for vision tasks
Metrics details
| Stars | 3,129 |
pharmapsychotic/clip-interrogator
Image to prompt with BLIP and CLIP
Metrics details
| Stars | 2,978 |
rom1504/clip-retrieval
Easily compute clip embeddings and build a clip retrieval system with them
Metrics details
| Stars | 2,786 |
QIN2DIM/hcaptcha-challenger
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
Metrics details
| Stars | 2,380 |
RuffianZhong/RWidgetHelper
Android UI 快速开发,专治原生控件各种不服
Metrics details
| Stars | 1,963 |
HFrost0/bilix
⚡️Lightning-fast async download tool for bilibili and more
Metrics details
| Stars | 1,781 |
roboflow/awesome-openai-vision-api-experiments
Must-have resource for anyone who wants to experiment with and build on the OpenAI vision API 🔥
Metrics details
| Stars | 1,689 |
mbzuai-oryx/Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
Metrics details
| Stars | 1,503 |
unum-cloud/uform
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
Metrics details
| Stars | 1,242 |
yzhuoning/Awesome-CLIP
Awesome list for research on CLIP (Contrastive Language-Image Pre-Training).
Metrics details
| Stars | 1,229 |
EdVince/Stable-Diffusion-NCNN
Stable Diffusion in NCNN with c++, supported txt2img and img2img
Metrics details
| Stars | 1,066 |
haltakov/natural-language-image-search
Search photos on Unsplash using natural language
Metrics details
| Stars | 1,041 |
ArrowLuo/CLIP4Clip
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
Metrics details
| Stars | 1,028 |
haltakov/natural-language-youtube-search
Search inside YouTube videos using natural language
Metrics details
| Stars | 935 |
hila-chefer/Transformer-MM-Explainability
[ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers, a novel method to visualize any Transformer-based network. Including examples for DETR, VQA.
Metrics details
| Stars | 911 |
omerbt/Text2LIVE
Official Pytorch Implementation for "Text2LIVE: Text-Driven Layered Image and Video Editing" (ECCV 2022 Oral)
Metrics details
| Stars | 888 |
pengsongyou/openscene
[CVPR'23] OpenScene: 3D Scene Understanding with Open Vocabularies
Metrics details
| Stars | 838 |
eps696/aphantasia
CLIP + FFT/DWT/RGB = text to image/video
Metrics details
| Stars | 790 |
PaddlePaddle/PaddleMIX
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
Metrics details
| Stars | 724 |
v-iashin/video_features
Extract video features from raw videos using multiple GPUs. We support RAFT flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, and TIMM models.
Metrics details
| Stars | 653 |
SkyWorkAIGC/SkyPaint-AI-Diffusion
基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model can generate high-quality images in several modern art styles.
Metrics details
| Stars | 647 |
SkalskiP/awesome-foundation-and-multimodal-models
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
Metrics details
| Stars | 637 |
leondgarse/keras_cv_attention_models
Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fasternet,fastervit,fastvit,flexivit,gcvit,ghostnet,gpvit,hornet,hiera,iformer,inceptionnext,lcnet,levit,maxvit,mobilevit,moganet,nat,nfnets,pvt,swin,tinynet,tinyvit,uniformer,volo,vanillanet,yolor,yolov7,yolov8,yolox,gpt2,llama2, alias kecam
Metrics details
| Stars | 627 |
pablosichert/react-truncate
React component for truncating multi-line spans and adding an ellipsis.
Metrics details
| Stars | 592 |
jina-ai/now
🧞 No-code tool for creating a neural search solution in minutes
Metrics details
| Stars | 588 |
slavabarkov/tidy
Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art vision-language pretrained CLIP model and ONNX Runtime inference engine
Metrics details
| Stars | 585 |
monatis/clip.cpp
CLIP inference in plain C/C++ with no extra dependencies
Metrics details
| Stars | 564 |
cliport/cliport
CLIPort: What and Where Pathways for Robotic Manipulation
Metrics details
| Stars | 546 |
patrickjohncyh/fashion-clip
FashionCLIP is a CLIP-like model fine-tuned for the fashion domain.
Metrics details
| Stars | 528 |
greyovo/PicQuery
🔍 Search local images with natural language on Android, powered by OpenAI's CLIP model. / 在 Android 上用自然语言搜索本地图片 (基于 OpenAI 的 CLIP 模型)
Metrics details
| Stars | 506 |
keshiim/ZMJImageEditor
ZMJImageEditor is a picture editing component like WeChat. It is powerful and easy to integrate, supporting rendering, text, rotation, tailoring, mapping and other functions. (ZMJImageEditor 是一个和微信一样图片编辑的组件,功能强大,极易集成,支持绘制、文字、旋转、剪裁、贴图等功能)
Metrics details
| Stars | 504 |
IceClear/CLIP-IQA
[AAAI 2023] Exploring CLIP for Assessing the Look and Feel of Images
Metrics details
| Stars | 490 |
poloclub/diffusion-explainer
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
Metrics details
| Stars | 482 |
vkgo/OCRAutoScore
OCR自动化阅卷项目
Metrics details
| Stars | 481 |
xmed-lab/CLIP_Surgery
[Pattern Recognition 25] CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
Metrics details
| Stars | 479 |
PathologyFoundation/plip
Pathology Language and Image Pre-Training (PLIP) is the first vision and language foundation model for Pathology AI (Nature Medicine). PLIP is a large-scale pre-trained model that can be used to extract visual and language features from pathology images and text description. The model is a fine-tuned version of the original CLIP model.
Metrics details
| Stars | 381 |
OpenGVLab/Instruct2Act
Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
Metrics details
| Stars | 374 |
Chrisvin/EasyReveal
Android Easy Reveal Library
Metrics details
| Stars | 358 |
liruiw/GenSim
Generating Robotic Simulation Tasks via Large Language Models
Metrics details
| Stars | 350 |
zcf0508/autocut-client
AutoCut Client
Metrics details
| Stars | 341 |
mu-cai/ViP-LLaVA
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
Metrics details
| Stars | 338 |
cyclomon/CLIPstyler
Official Pytorch implementation of "CLIPstyler:Image Style Transfer with a Single Text Condition" (CVPR 2022)
Metrics details
| Stars | 327 |
Taited/clip-score
Quick scripts to calculate CLIP text-image similarity
Metrics details
| Stars | 319 |
MohamadZeina/Disco_Diffusion_Local
Getting the latest versions of Disco Diffusion to work locally, instead of colab. Including how I run this on Windows, despite some Linux only dependencies ;)
Metrics details
| Stars | 315 |
mertyg/vision-language-models-are-bows
Experiments and data for the paper "When and why vision-language models behave like bags-of-words, and what to do about it?" Oral @ ICLR 2023
Metrics details
| Stars | 294 |
PaddlePaddle/PASSL
PASSL包含 SimCLR,MoCo v1/v2,BYOL,CLIP,PixPro,simsiam, SwAV, BEiT,MAE 等图像自监督算法以及 Vision Transformer,DEiT,Swin Transformer,CvT,T2T-ViT,MLP-Mixer,XCiT,ConvNeXt,PVTv2 等基础视觉算法
Metrics details
| Stars | 290 |
EdVince/CLIP-ImageSearch-NCNN
CLIP⚡NCNN⚡基于自然语言的图片搜索(Image Search)⚡以字搜图⚡x86⚡Android
Metrics details
| Stars | 281 |
ByChelsea/VAND-APRIL-GAN
[CVPR 2023 Workshop] VAND Challenge: 1st Place on Zero-shot AD and 4th Place on Few-shot AD
Metrics details
| Stars | 269 |
Imageomics/bioclip
This is the repository for the BioCLIP model and the TreeOfLife-10M dataset [CVPR'24 Oral, Best Student Paper].
Metrics details
| Stars | 268 |
yxuansu/MAGIC
Language Models Can See: Plugging Visual Controls in Text Generation
Metrics details
| Stars | 261 |
chao1224/MoleculeSTM
Multi-modal Molecule Structure-text Model for Text-based Editing and Retrieval, Nat Mach Intell 2023 (https://www.nature.com/articles/s42256-023-00759-6)
Metrics details
| Stars | 259 |
j-min/CLIP-Caption-Reward
PyTorch code for "Fine-grained Image Captioning with CLIP Reward" (Findings of NAACL 2022)
Metrics details
| Stars | 246 |
zwx8981/LIQE
[CVPR2023] Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning Perspective
Metrics details
| Stars | 241 |
Lednik7/CLIP-ONNX
It is a simple library to speed up CLIP inference up to 3x (K80 GPU)
Metrics details
| Stars | 234 |
