Asr Open Source Projects

Browse 178 Asr open source projects, ranked by GitHub stars. Find the most popular Asr tools and libraries.

Share your experience:✍️ Write a Post❓ Ask a Question
1-60 of 178 projects
23,157 stars

m-bain/whisperX

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Metrics details
Stars23,157
17,773 stars

NVIDIA/NeMo

NeMo: a toolkit for conversational AI

Metrics details
Stars17,773
15,431 stars

kaldi-asr/kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

Metrics details
Stars15,431
14,964 stars

alphacep/vosk-api

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

Metrics details
Stars14,964
13,651 stars

k2-fsa/sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

Metrics details
Stars13,651
12,650 stars

PaddlePaddle/PaddleSpeech

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.

Metrics details
Stars12,650
11,695 stars

speechbrain/speechbrain

A PyTorch-based Speech Toolkit

Metrics details
Stars11,695
9,896 stars

espnet/espnet

End-to-End Speech Processing Toolkit

Metrics details
Stars9,896
7,935 stars

jdepoix/youtube-transcript-api

This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!

Metrics details
Stars7,935
7,121 stars

wzpan/wukong-robot

🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。

Metrics details
Stars7,121
6,017 stars

snakers4/silero-models

Silero Models: pre-trained text-to-speech models made embarrassingly simple

Metrics details
Stars6,017
5,622 stars

xiangyuecn/Recorder

html5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文字 H5版语音通话聊天示例 DTMF编码解码

Metrics details
Stars5,622
5,600 stars

MahmoudAshraf97/whisper-diarization

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

Metrics details
Stars5,600
5,170 stars

wenet-e2e/wenet

Production First and Production Ready End-to-End Speech Recognition Toolkit

Metrics details
Stars5,170
3,304 stars

ahmetoner/whisper-asr-webservice

OpenAI Whisper ASR Webservice API

Metrics details
Stars3,304
3,117 stars

Purfview/whisper-standalone-win

Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.

Metrics details
Stars3,117
2,980 stars

CheshireCC/faster-whisper-GUI

faster_whisper GUI with PySide6

Metrics details
Stars2,980
2,860 stars

tensorflow/lingvo

Lingvo

Metrics details
Stars2,860
2,824 stars

linto-ai/whisper-timestamped

Multilingual Automatic Speech Recognition with word-level timestamps and confidence

Metrics details
Stars2,824
2,750 stars

rhasspy/rhasspy

Offline private voice assistant for many human languages

Metrics details
Stars2,750
2,591 stars

coqui-ai/STT

🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.

Metrics details
Stars2,591
2,398 stars

mravanelli/pytorch-kaldi

pytorch-kaldi is a project for developing state-of-the-art DNN/RNN hybrid speech recognition systems. The DNN part is managed by pytorch, while feature extraction, label computation, and decoding are performed with the kaldi toolkit.

Metrics details
Stars2,398
1,864 stars

syhw/wer_are_we

Attempt at tracking states of the arts and recent results (bibliography) on speech recognition.

Metrics details
Stars1,864
1,670 stars

jovotech/jovo-framework

🔈 The React for Voice and Chat: Build Apps for Alexa, Messenger, Instagram, the Web, and more

Metrics details
Stars1,670
1,606 stars

Delta-ML/delta

DELTA is a deep learning based natural language and speech processing platform. LF AI & DATA Projects: https://lfaidata.foundation/projects/delta/

Metrics details
Stars1,606
1,541 stars

mkiol/dsnote

Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.

Metrics details
Stars1,541
1,497 stars

google/live-transcribe-speech-engine

Live Transcribe is an Android application that provides real-time captioning for people who are deaf or hard of hearing. This repository contains the Android client libraries for communicating with Google's Cloud Speech API that are used in Live Transcribe.

Metrics details
Stars1,497
1,259 stars

alphacep/vosk-server

WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries

Metrics details
Stars1,259
1,241 stars

mravanelli/SincNet

SincNet is a neural architecture for efficiently processing raw audio samples.

Metrics details
Stars1,241
1,218 stars

yeyupiaoling/Whisper-Finetune

Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment

Metrics details
Stars1,218
1,130 stars

sooftware/conformer

[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)

Metrics details
Stars1,130
1,053 stars

alphacep/vosk-android-demo

Offline speech recognition for Android with Vosk library.

Metrics details
Stars1,053
1,038 stars

pykaldi/pykaldi

A Python wrapper for Kaldi

Metrics details
Stars1,038
969 stars

athena-team/athena

an open-source implementation of sequence-to-sequence based speech processing engine

Metrics details
Stars969
960 stars

k2-fsa/sherpa

Speech-to-text server framework with next-gen Kaldi

Metrics details
Stars960
939 stars

freewym/espresso

Espresso: A Fast End-to-End Neural Speech Recognition Toolkit

Metrics details
Stars939
920 stars

songys/AwesomeKorean_Data

한국어 데이터 세트 링크

Metrics details
Stars920
916 stars

innovatorved/whisper.api

This project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model.

Metrics details
Stars916
872 stars

yeyupiaoling/PPASR

基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型

Metrics details
Stars872
834 stars

srvk/eesen

The official repository of the Eesen project

Metrics details
Stars834
826 stars

snakers4/open_stt

Open STT

Metrics details
Stars826
810 stars

kaituoxu/Speech-Transformer

A PyTorch implementation of Speech Transformer, an End-to-End ASR with Transformer network on Mandarin Chinese.

Metrics details
Stars810
764 stars

Ailln/cn2an

📦 快速转化「中文数字」和「阿拉伯数字」~ (最新特性:分数,日期、温度等转化)

Metrics details
Stars764
761 stars

yeyupiaoling/PaddlePaddle-DeepSpeech

基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。

Metrics details
Stars761
749 stars

Macoron/whisper.unity

Running speech to text model (whisper.cpp) in Unity3d on your local machine.

Metrics details
Stars749
733 stars

speechio/chinese_text_normalization

Chinese text normalization for speech processing

Metrics details
Stars733
726 stars

yeyupiaoling/MASR

Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。

Metrics details
Stars726
716 stars

openspeech-team/openspeech

Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.

Metrics details
Stars716
685 stars

DmitryRyumin/INTERSPEECH-2023-Papers

INTERSPEECH 2023 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023 conference. Explore the latest advances in speech and language processing. Code included. Star the repository to support the advancement of speech technology!

Metrics details
Stars685
679 stars

iceychris/LibreASR

:speech_balloon: An On-Premises, Streaming Speech Recognition System

Metrics details
Stars679
677 stars

vilassn/whisper_android

Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android

Metrics details
Stars677
675 stars

yinruiqing/pyannote-whisper

Metrics details
Stars675
669 stars

Picovoice/cheetah

On-device streaming speech-to-text engine powered by deep learning

Metrics details
Stars669
650 stars

abhirooptalasila/AutoSub

A CLI script to generate subtitle files (SRT/VTT/TXT) for any video using either DeepSpeech or Coqui

Metrics details
Stars650
637 stars

sooftware/kospeech

Open-Source Toolkit for End-to-End Korean Automatic Speech Recognition leveraging PyTorch and Hydra.

Metrics details
Stars637
618 stars

zw76859420/ASR_Theory

语音识别理论、论文和PPT

Metrics details
Stars618
608 stars

RapidAI/RapidASR

📣 商用级开源语音自动识别程序库,开箱即用,全平台支持,中英文混合识别。A Cross-platform implementation of ASR inference. It's based on ONNXRuntime and FunASR. We provide a set of easier APIs to call ASR models.

Metrics details
Stars608
594 stars

hirofumi0810/neural_sp

End-to-end ASR/LM implementation with PyTorch

Metrics details
Stars594
577 stars

shashikg/WhisperS2T

An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

Metrics details
Stars577
547 stars

SpeechColab/Leaderboard

SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.

Metrics details
Stars547
1-60 of 178 projects
Get A Weekly Email With Trending Asr Projects
Stay updated on Asr plus related topics you pick below.

Copyright 2018-2026 Awesome Open Source.  All rights reserved.