Asr Open Source Projects
Browse 178 Asr open source projects, ranked by GitHub stars. Find the most popular Asr tools and libraries.
m-bain/whisperX
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Metrics details
| Stars | 23,157 |
NVIDIA/NeMo
NeMo: a toolkit for conversational AI
Metrics details
| Stars | 17,773 |
kaldi-asr/kaldi
kaldi-asr/kaldi is the official location of the Kaldi project.
Metrics details
| Stars | 15,431 |
alphacep/vosk-api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Metrics details
| Stars | 14,964 |
k2-fsa/sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
Metrics details
| Stars | 13,651 |
PaddlePaddle/PaddleSpeech
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
Metrics details
| Stars | 12,650 |
speechbrain/speechbrain
A PyTorch-based Speech Toolkit
Metrics details
| Stars | 11,695 |
espnet/espnet
End-to-End Speech Processing Toolkit
Metrics details
| Stars | 9,896 |
jdepoix/youtube-transcript-api
This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!
Metrics details
| Stars | 7,935 |
wzpan/wukong-robot
🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。
Metrics details
| Stars | 7,121 |
snakers4/silero-models
Silero Models: pre-trained text-to-speech models made embarrassingly simple
Metrics details
| Stars | 6,017 |
xiangyuecn/Recorder
html5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文字 H5版语音通话聊天示例 DTMF编码解码
Metrics details
| Stars | 5,622 |
MahmoudAshraf97/whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Metrics details
| Stars | 5,600 |
wenet-e2e/wenet
Production First and Production Ready End-to-End Speech Recognition Toolkit
Metrics details
| Stars | 5,170 |
ahmetoner/whisper-asr-webservice
OpenAI Whisper ASR Webservice API
Metrics details
| Stars | 3,304 |
Purfview/whisper-standalone-win
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
Metrics details
| Stars | 3,117 |
CheshireCC/faster-whisper-GUI
faster_whisper GUI with PySide6
Metrics details
| Stars | 2,980 |
tensorflow/lingvo
Lingvo
Metrics details
| Stars | 2,860 |
linto-ai/whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
Metrics details
| Stars | 2,824 |
rhasspy/rhasspy
Offline private voice assistant for many human languages
Metrics details
| Stars | 2,750 |
coqui-ai/STT
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
Metrics details
| Stars | 2,591 |
mravanelli/pytorch-kaldi
pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.
Metrics details
| Stars | 2,398 |
syhw/wer_are_we
Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.
Metrics details
| Stars | 1,864 |
jovotech/jovo-framework
🔈 The React for Voice and Chat: Build Apps for Alexa, Messenger, Instagram, the Web, and more
Metrics details
| Stars | 1,670 |
Delta-ML/delta
DELTA is a deep learning based natural language and speech processing platform. LF AI & DATA Projects: https://lfaidata.foundation/projects/delta/
Metrics details
| Stars | 1,606 |
mkiol/dsnote
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
Metrics details
| Stars | 1,541 |
google/live-transcribe-speech-engine
Live Transcribe is an Android application that provides real-time captioning for people who are deaf or hard of hearing. This repository contains the Android client libraries for communicating with Google's Cloud Speech API that are used in Live Transcribe.
Metrics details
| Stars | 1,497 |
alphacep/vosk-server
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
Metrics details
| Stars | 1,259 |
mravanelli/SincNet
SincNet is a neural architecture for efficiently processing raw audio samples.
Metrics details
| Stars | 1,241 |
yeyupiaoling/Whisper-Finetune
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
Metrics details
| Stars | 1,218 |
sooftware/conformer
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
Metrics details
| Stars | 1,130 |
alphacep/vosk-android-demo
Offline speech recognition for Android with Vosk library.
Metrics details
| Stars | 1,053 |
pykaldi/pykaldi
A Python wrapper for Kaldi
Metrics details
| Stars | 1,038 |
athena-team/athena
an open-source implementation of sequence-to-sequence based speech processing engine
Metrics details
| Stars | 969 |
k2-fsa/sherpa
Speech-to-text server framework with next-gen Kaldi
Metrics details
| Stars | 960 |
freewym/espresso
Espresso: A Fast End-to-End Neural Speech Recognition Toolkit
Metrics details
| Stars | 939 |
songys/AwesomeKorean_Data
한국어 데이터 세트 링크
Metrics details
| Stars | 920 |
innovatorved/whisper.api
This project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model.
Metrics details
| Stars | 916 |
yeyupiaoling/PPASR
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
Metrics details
| Stars | 872 |
srvk/eesen
The official repository of the Eesen project
Metrics details
| Stars | 834 |
snakers4/open_stt
Open STT
Metrics details
| Stars | 826 |
kaituoxu/Speech-Transformer
A PyTorch implementation of Speech Transformer, an End-to-End ASR with Transformer network on Mandarin Chinese.
Metrics details
| Stars | 810 |
Ailln/cn2an
📦 快速转化「中文数字」和「阿拉伯数字」~ (最新特性:分数,日期、温度等转化)
Metrics details
| Stars | 764 |
yeyupiaoling/PaddlePaddle-DeepSpeech
基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。
Metrics details
| Stars | 761 |
Macoron/whisper.unity
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
Metrics details
| Stars | 749 |
speechio/chinese_text_normalization
Chinese text normalization for speech processing
Metrics details
| Stars | 733 |
yeyupiaoling/MASR
Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。
Metrics details
| Stars | 726 |
openspeech-team/openspeech
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
Metrics details
| Stars | 716 |
DmitryRyumin/INTERSPEECH-2023-Papers
INTERSPEECH 2023 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023 conference. Explore the latest advances in speech and language processing. Code included. Star the repository to support the advancement of speech technology!
Metrics details
| Stars | 685 |
iceychris/LibreASR
:speech_balloon: An On-Premises, Streaming Speech Recognition System
Metrics details
| Stars | 679 |
vilassn/whisper_android
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
Metrics details
| Stars | 677 |
yinruiqing/pyannote-whisper
Metrics details
| Stars | 675 |
Picovoice/cheetah
On-device streaming speech-to-text engine powered by deep learning
Metrics details
| Stars | 669 |
abhirooptalasila/AutoSub
A CLI script to generate subtitle files (SRT/VTT/TXT) for any video using either DeepSpeech or Coqui
Metrics details
| Stars | 650 |
sooftware/kospeech
Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.
Metrics details
| Stars | 637 |
zw76859420/ASR_Theory
语音识别理论、论文和PPT
Metrics details
| Stars | 618 |
RapidAI/RapidASR
📣 商用级开源语音自动识别程序库,开箱即用,全平台支持,中英文混合识别。A Cross-platform implementation of ASR inference. It's based on ONNXRuntime and FunASR. We provide a set of easier APIs to call ASR models.
Metrics details
| Stars | 608 |
hirofumi0810/neural_sp
End-to-end ASR/LM implementation with PyTorch
Metrics details
| Stars | 594 |
shashikg/WhisperS2T
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
Metrics details
| Stars | 577 |
SpeechColab/Leaderboard
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
Metrics details
| Stars | 547 |
