Gensim Open Source Projects
Browse 70 Gensim open source projects, ranked by GitHub stars. Find the most popular Gensim tools and libraries.
piskvorky/gensim
Topic Modelling for Humans
Metrics details
| Stars | 16,464 |
phanein/deepwalk
DeepWalk - Deep Learning for Graphs
Metrics details
| Stars | 2,756 |
adashofdata/nlp-in-python-tutorial
comparing stand up comedians using natural language processing
Metrics details
| Stars | 1,730 |
dipanjanS/text-analytics-with-python
Learn how to process, classify, cluster, summarize, understand syntax, semantics and sentiment of text data with the power of Python! This repository contains code and datasets used in my book, "Text Analytics with Python" published by Apress/Springer.
Metrics details
| Stars | 1,696 |
explosion/sense2vec
🦆 Contextually-keyed word vectors
Metrics details
| Stars | 1,678 |
plasticityai/magnitude
A fast, efficient universal vector embedding utility package.
Metrics details
| Stars | 1,665 |
msgi/nlp-journey
Documents, papers and codes related to Natural Language Processing, including Topic Model, Word Embedding, Named Entity Recognition, Text Classificatin, Text Generation, Text Similarity, Machine Translation),etc.
Metrics details
| Stars | 1,629 |
kavgan/nlp-in-practice
Starter code to solve real world text data problems. Includes: Gensim Word2Vec, phrase embeddings, Text Classification with Logistic Regression, word count with pyspark, simple text preprocessing, pre-trained embeddings and more.
Metrics details
| Stars | 1,183 |
guoday/Tencent2020_Rank1st
The code for 2020 Tencent College Algorithm Contest, and the online result ranks 1st.
Metrics details
| Stars | 1,083 |
hundredblocks/concrete_NLP_tutorial
An NLP workshop about concrete solutions to real problems
Metrics details
| Stars | 1,077 |
RaRe-Technologies/gensim-data
Data repository for pretrained NLP models and NLP corpora.
Metrics details
| Stars | 1,056 |
linanqiu/word2vec-sentiments
Tutorial for Sentiment Analysis using Doc2Vec in gensim (or "getting 87% accuracy in sentiment analysis in under 100 lines of code")
Metrics details
| Stars | 682 |
jhlau/doc2vec
Python scripts for training/testing paragraph vectors
Metrics details
| Stars | 653 |
oborchers/Fast_Sentence_Embeddings
Compute Sentence Embeddings Fast!
Metrics details
| Stars | 625 |
idio/wiki2vec
Generating Vectors for DBpedia Entities via Word2Vec and Wikipedia Dumps. Questions? https://gitter.im/idio-opensource/Lobby
Metrics details
| Stars | 602 |
AimeeLee77/wiki_zh_word2vec
利用Python构建Wiki中文语料词向量模型试验
Metrics details
| Stars | 523 |
zake7749/word2vec-tutorial
中文詞向量訓練教學
Metrics details
| Stars | 522 |
clayandgithub/zh_cnn_text_classify
基于CNN的中文文本分类算法(可应用于垃圾邮件过滤、情感分析等场景)
Metrics details
| Stars | 460 |
cjymz886/text-cnn
嵌入Word2vec词向量的CNN中文文本分类
Metrics details
| Stars | 448 |
dsfsi/textaugment
TextAugment: Text Augmentation Library
Metrics details
| Stars | 442 |
bakrianoo/aravec
AraVec is a pre-trained distributed word representation (word embedding) open source project which aims to provide the Arabic NLP research community with free to use and powerful word embedding models.
Metrics details
| Stars | 423 |
ThoughtRiver/lmdb-embeddings
Fast word vectors with little memory usage in Python
Metrics details
| Stars | 416 |
FanhuaandLuomu/BiLstm_CNN_CRF_CWS
BiLstm+CNN+CRF 法律文档(合同类案件)领域分词(100篇标注样本)
Metrics details
| Stars | 389 |
sdadas/polish-nlp-resources
Pre-trained models and language resources for Natural Language Processing in Polish
Metrics details
| Stars | 377 |
pskun/finance_news_analysis
金融新闻数据挖掘分析
Metrics details
| Stars | 368 |
5hirish/adam_qas
ADAM - A Question Answering System. Inspired from IBM Watson
Metrics details
| Stars | 356 |
AICoE/log-anomaly-detector
Log Anomaly Detection - Machine learning to detect abnormal events logs
Metrics details
| Stars | 347 |
textpipe/textpipe
Textpipe: clean and extract metadata from text
Metrics details
| Stars | 302 |
30lm32/ml-projects
ML based projects such as Spam Classification, Time Series Analysis, Text Classification using Random Forest, Deep Learning, Bayesian, Xgboost in Python
Metrics details
| Stars | 295 |
cjymz886/sentence-similarity
对四种句子/文本相似度计算方法进行实验与比较
Metrics details
| Stars | 292 |
hecongqing/2018-daguan-competition
2018年"达观杯"文本智能处理挑战赛-长文本分类-rank4
Metrics details
| Stars | 282 |
benedekrozemberczki/GEMSEC
The TensorFlow reference implementation of 'GEMSEC: Graph Embedding with Self Clustering' (ASONAM 2019).
Metrics details
| Stars | 258 |
DevinZ1993/Chinese-Poetry-Generation
An ML-based Chinese Poem Generator
Metrics details
| Stars | 257 |
yourh/AttentionXML
Implementation for "AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text Classification"
Metrics details
| Stars | 252 |
oxford-cs-deepnlp-2017/practical-1
Oxford Deep NLP 2017 course - Practical 1: word2vec
Metrics details
| Stars | 250 |
davidberenstein1957/concise-concepts
This repository contains an easy and intuitive approach to few-shot NER using most similar expansion over spaCy embeddings. Now with entity scoring.
Metrics details
| Stars | 244 |
devmount/GermanWordEmbeddings
Toolkit to obtain and preprocess German text corpora, train models and evaluate them with generated testsets. Built with Gensim and Tensorflow.
Metrics details
| Stars | 243 |
alisonmitchell/Stock-Prediction
Technical and sentiment analysis to predict the stock market with machine learning models based on historical time series data and news article sentiment collected using APIs and web scraping.
Metrics details
| Stars | 224 |
akoksal/Turkish-Word2Vec
Pre-trained Word2Vec Model for Turkish
Metrics details
| Stars | 222 |
bainingchao/PyDataPreprocessing
《Python数据预处理技术与实践》源码下载
Metrics details
| Stars | 214 |
benedekrozemberczki/Splitter
A Pytorch implementation of "Splitter: Learning Node Representations that Capture Multiple Social Contexts" (WWW 2019).
Metrics details
| Stars | 214 |
platisd/duplicate-code-detection-tool
A simple Python3 tool to detect similarities between files within a repository
Metrics details
| Stars | 208 |
columbia-applied-data-science/rosetta
Tools, wrappers, etc... for data science with a concentration on text processing
Metrics details
| Stars | 207 |
akutuzov/webvectors
Web-ify your word2vec: framework to serve distributional semantic models online
Metrics details
| Stars | 204 |
niitsuma/word2vec-keras-in-gensim
word2vec uisng keras inside gensim
Metrics details
| Stars | 204 |
giacbrd/ShallowLearn
An experiment about re-implementing supervised learning models based on shallow neural network approaches (e.g. fastText) with some additional exclusive features and nice API. Written in Python and fully compatible with Scikit-learn.
Metrics details
| Stars | 198 |
avidale/compress-fasttext
Tools for shrinking fastText models (in gensim format)
Metrics details
| Stars | 187 |
benedekrozemberczki/MUSAE
The reference implementation of "Multi-scale Attributed Node Embedding". (Journal of Complex Networks 2021)
Metrics details
| Stars | 186 |
jsksxs360/Word2Vec
对 ansj 编写的 Word2VEC_java 的进一步包装,同时实现了常用的词语相似度和句子相似度计算。
Metrics details
| Stars | 186 |
Disiok/poetry-seq2seq
Chinese Poetry Generation
Metrics details
| Stars | 180 |
WorksApplications/chiVe
Japanese word embedding with Sudachi and NWJC 🌿
Metrics details
| Stars | 177 |
benedekrozemberczki/role2vec
A scalable Gensim implementation of "Learning Role-based Graph Embeddings" (IJCAI 2018).
Metrics details
| Stars | 169 |
taozhijiang/chinese_nlp
Chinese Natural Language Processing tools and examples
Metrics details
| Stars | 161 |
RaRe-Technologies/w2v_server_googlenews
Code for the word2vec HTTP server running at https://rare-technologies.com/word2vec-tutorial/#bonus_app
Metrics details
| Stars | 157 |
PrashantRanjan09/WordEmbeddings-Elmo-Fasttext-Word2Vec
Using pre trained word embeddings (Fasttext, Word2Vec)
Metrics details
| Stars | 157 |
nlpjoe/daguan-classify-2018
2018达观杯长文本分类智能处理挑战赛 18解决方案
Metrics details
| Stars | 153 |
jingcheng-du/Gene2vec
Gene2Vec: Distributed Representation of Genes Based on Co-Expression
Metrics details
| Stars | 140 |
jhlau/topically-driven-language-model
Tensorflow code to train TDLM
Metrics details
| Stars | 136 |
dipanjanS/nlp_workshop_odsc_europe20
Extensive tutorials for the Advanced NLP Workshop in Open Data Science Conference Europe 2020. We will leverage machine learning, deep learning and deep transfer learning to learn and solve popular tasks using NLP including NER, Classification, Recommendation \ Information Retrieval, Summarization, Classification, Language Translation, Q&A and Topic Models.
Metrics details
| Stars | 135 |
ATEC2018/deep-siamese-text-similarity
基于siamese-lstm的中文句子相似度计算
Metrics details
| Stars | 129 |
