Gensim Open Source Projects

Browse 70 Gensim open source projects, ranked by GitHub stars. Find the most popular Gensim tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 70 projects
16,464 stars

piskvorky/gensim

Topic Modelling for Humans

Metrics details
Stars16,464
2,756 stars

phanein/deepwalk

DeepWalk - Deep Learning for Graphs

Metrics details
Stars2,756
1,730 stars

adashofdata/nlp-in-python-tutorial

comparing stand up comedians using natural language processing

Metrics details
Stars1,730
1,696 stars

dipanjanS/text-analytics-with-python

Learn how to process, classify, cluster, summarize, understand syntax, semantics and sentiment of text data with the power of Python! This repository contains code and datasets used in my book, "Text Analytics with Python" published by Apress/Springer.

Metrics details
Stars1,696
1,678 stars

explosion/sense2vec

🦆 Contextually-keyed word vectors

Metrics details
Stars1,678
1,665 stars

plasticityai/magnitude

A fast, efficient universal vector embedding utility package.

Metrics details
Stars1,665
1,629 stars

msgi/nlp-journey

Documents, papers and codes related to Natural Language Processing, including Topic Model, Word Embedding, Named Entity Recognition, Text Classificatin, Text Generation, Text Similarity, Machine Translation),etc.

Metrics details
Stars1,629
1,183 stars

kavgan/nlp-in-practice

Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.

Metrics details
Stars1,183
1,083 stars

guoday/Tencent2020_Rank1st

The code for 2020 Tencent College Algorithm Contest, and the online result ranks 1st.

Metrics details
Stars1,083
1,077 stars

hundredblocks/concrete_NLP_tutorial

An NLP workshop about concrete solutions to real problems

Metrics details
Stars1,077
1,056 stars

RaRe-Technologies/gensim-data

Data repository for pretrained NLP models and NLP corpora.

Metrics details
Stars1,056
682 stars

linanqiu/word2vec-sentiments

Tutorial for Sentiment Analysis using Doc2Vec in gensim (or "getting 87% accuracy in sentiment analysis in under 100 lines of code")

Metrics details
Stars682
653 stars

jhlau/doc2vec

Python scripts for training/testing paragraph vectors

Metrics details
Stars653
625 stars

oborchers/Fast_Sentence_Embeddings

Compute Sentence Embeddings Fast!

Metrics details
Stars625
602 stars

idio/wiki2vec

Generating Vectors for DBpedia Entities via Word2Vec and Wikipedia Dumps. Questions? https://gitter.im/idio-opensource/Lobby

Metrics details
Stars602
523 stars

AimeeLee77/wiki_zh_word2vec

利用Python构建Wiki中文语料词向量模型试验

Metrics details
Stars523
522 stars

zake7749/word2vec-tutorial

中文詞向量訓練教學

Metrics details
Stars522
460 stars

clayandgithub/zh_cnn_text_classify

基于CNN的中文文本分类算法(可应用于垃圾邮件过滤、情感分析等场景)

Metrics details
Stars460
448 stars

cjymz886/text-cnn

嵌入Word2vec词向量的CNN中文文本分类

Metrics details
Stars448
442 stars

dsfsi/textaugment

TextAugment: Text Augmentation Library

Metrics details
Stars442
423 stars

bakrianoo/aravec

AraVec is a pre-trained distributed word representation (word embedding) open source project which aims to provide the Arabic NLP research community with free to use and powerful word embedding models.

Metrics details
Stars423
416 stars

ThoughtRiver/lmdb-embeddings

Fast word vectors with little memory usage in Python

Metrics details
Stars416
389 stars

FanhuaandLuomu/BiLstm_CNN_CRF_CWS

BiLstm+CNN+CRF 法律文档(合同类案件)领域分词(100篇标注样本)

Metrics details
Stars389
377 stars

sdadas/polish-nlp-resources

Pre-trained models and language resources for Natural Language Processing in Polish

Metrics details
Stars377
368 stars

pskun/finance_news_analysis

金融新闻数据挖掘分析

Metrics details
Stars368
356 stars

5hirish/adam_qas

ADAM - A Question Answering System. Inspired from IBM Watson

Metrics details
Stars356
347 stars

AICoE/log-anomaly-detector

Log Anomaly Detection - Machine learning to detect abnormal events logs

Metrics details
Stars347
302 stars

textpipe/textpipe

Textpipe: clean and extract metadata from text

Metrics details
Stars302
295 stars

30lm32/ml-projects

ML based projects such as Spam Classification, Time Series Analysis, Text Classification using Random Forest, Deep Learning, Bayesian, Xgboost in Python

Metrics details
Stars295
292 stars

cjymz886/sentence-similarity

对四种句子/文本相似度计算方法进行实验与比较

Metrics details
Stars292
282 stars

hecongqing/2018-daguan-competition

2018年"达观杯"文本智能处理挑战赛-长文本分类-rank4

Metrics details
Stars282
258 stars

benedekrozemberczki/GEMSEC

The TensorFlow reference implementation of 'GEMSEC: Graph Embedding with Self Clustering' (ASONAM 2019).

Metrics details
Stars258
257 stars

DevinZ1993/Chinese-Poetry-Generation

An ML-based Chinese Poem Generator

Metrics details
Stars257
252 stars

yourh/AttentionXML

Implementation for "AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text Classification"

Metrics details
Stars252
250 stars

oxford-cs-deepnlp-2017/practical-1

Oxford Deep NLP 2017 course - Practical 1: word2vec

Metrics details
Stars250
244 stars

davidberenstein1957/concise-concepts

This repository contains an easy and intuitive approach to few-shot NER using most similar expansion over spaCy embeddings. Now with entity scoring.

Metrics details
Stars244
243 stars

devmount/GermanWordEmbeddings

Toolkit to obtain and preprocess German text corpora, train models and evaluate them with generated testsets. Built with Gensim and Tensorflow.

Metrics details
Stars243
224 stars

alisonmitchell/Stock-Prediction

Technical and sentiment analysis to predict the stock market with machine learning models based on historical time series data and news article sentiment collected using APIs and web scraping.

Metrics details
Stars224
222 stars

akoksal/Turkish-Word2Vec

Pre-trained Word2Vec Model for Turkish

Metrics details
Stars222
214 stars

bainingchao/PyDataPreprocessing

《Python数据预处理技术与实践》源码下载

Metrics details
Stars214
214 stars

benedekrozemberczki/Splitter

A Pytorch implementation of "Splitter: Learning Node Representations that Capture Multiple Social Contexts" (WWW 2019).

Metrics details
Stars214
208 stars

platisd/duplicate-code-detection-tool

A simple Python3 tool to detect similarities between files within a repository

Metrics details
Stars208
207 stars

columbia-applied-data-science/rosetta

Tools, wrappers, etc... for data science with a concentration on text processing

Metrics details
Stars207
204 stars

akutuzov/webvectors

Web-ify your word2vec: framework to serve distributional semantic models online

Metrics details
Stars204
204 stars

niitsuma/word2vec-keras-in-gensim

word2vec uisng keras inside gensim

Metrics details
Stars204
198 stars

giacbrd/ShallowLearn

An experiment about re-implementing supervised learning models based on shallow neural network approaches (e.g. fastText) with some additional exclusive features and nice API. Written in Python and fully compatible with Scikit-learn.

Metrics details
Stars198
187 stars

avidale/compress-fasttext

Tools for shrinking fastText models (in gensim format)

Metrics details
Stars187
186 stars

benedekrozemberczki/MUSAE

The reference implementation of "Multi-scale Attributed Node Embedding". (Journal of Complex Networks 2021)

Metrics details
Stars186
186 stars

jsksxs360/Word2Vec

对 ansj 编写的 Word2VEC_java 的进一步包装,同时实现了常用的词语相似度和句子相似度计算。

Metrics details
Stars186
180 stars

Disiok/poetry-seq2seq

Chinese Poetry Generation

Metrics details
Stars180
177 stars

WorksApplications/chiVe

Japanese word embedding with Sudachi and NWJC 🌿

Metrics details
Stars177
169 stars

benedekrozemberczki/role2vec

A scalable Gensim implementation of "Learning Role-based Graph Embeddings" (IJCAI 2018).

Metrics details
Stars169
161 stars

taozhijiang/chinese_nlp

Chinese Natural Language Processing tools and examples

Metrics details
Stars161
157 stars

RaRe-Technologies/w2v_server_googlenews

Code for the word2vec HTTP server running at https://rare-technologies.com/word2vec-tutorial/#bonus_app

Metrics details
Stars157
157 stars

PrashantRanjan09/WordEmbeddings-Elmo-Fasttext-Word2Vec

Using pre trained word embeddings (Fasttext, Word2Vec)

Metrics details
Stars157
153 stars

nlpjoe/daguan-classify-2018

2018达观杯长文本分类智能处理挑战赛 18解决方案

Metrics details
Stars153
140 stars

jingcheng-du/Gene2vec

Gene2Vec: Distributed Representation of Genes Based on Co-Expression

Metrics details
Stars140
136 stars

jhlau/topically-driven-language-model

Tensorflow code to train TDLM

Metrics details
Stars136
135 stars

dipanjanS/nlp_workshop_odsc_europe20

Extensive tutorials for the Advanced NLP Workshop in Open Data Science Conference Europe 2020. We will leverage machine learning, deep learning and deep transfer learning to learn and solve popular tasks using NLP including NER, Classification, Recommendation \ Information Retrieval, Summarization, Classification, Language Translation, Q&A and Topic Models.

Metrics details
Stars135
129 stars

ATEC2018/deep-siamese-text-similarity

基于siamese-lstm的中文句子相似度计算

Metrics details
Stars129
1-60 of 70 projects
Get A Weekly Email With Trending Gensim Projects
Stay updated on Gensim plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.