Automatic Speech Recognition
ASR models for multilingual transcription and streaming speech recognition, with notes on use cases and licenses.
Automatic Speech Recognition • 2B • Updated • 1.76M • • 1.13kNote Multilingual ASR: 30 languages and 22 Chinese dialects. Offline and streaming inference; streaming uses the vLLM backend. Apache-2.0.
nvidia/parakeet-tdt-0.6b-v3
Automatic Speech Recognition • 0.6B • Updated • 562k • • 1.16kNote Compact 0.6B multilingual ASR covering 25 European languages. Useful transcription baseline in the NeMo/Transformers ecosystem. CC-BY-4.0.
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.4M • • 3.41kNote Widely supported multilingual transcription baseline. Faster variant of Whisper large-v3. MIT.
kyutai/stt-1b-en_fr
Automatic Speech Recognition • 1.0B • Updated • 137Note Native streaming ASR for English and French. Produces punctuated transcripts while audio arrives. CC-BY-4.0.
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 1.06M • • 1.15kNote Native cache-aware streaming ASR with configurable chunk sizes and punctuation. Multilingual coverage is tiered; some locales require fine-tuning. OpenMDW-1.1.
mistralai/Voxtral-Mini-4B-Realtime-2602
Automatic Speech Recognition • 4B • Updated • 1.82M • 995Note Native streaming multilingual ASR with a causal audio encoder and configurable transcription delay. Useful for latency/accuracy comparisons. Apache-2.0.
nvidia/canary-qwen-2.5b
Automatic Speech Recognition • 3B • Updated • 41.2k • 465Note English ASR with punctuation and capitalization. Separate LLM mode can process the transcript; raw-audio understanding is not retained in that mode. CC-BY-4.0.
FunAudioLLM/SenseVoiceSmall
Automatic Speech Recognition • Updated • 19.3k • 486Note Non-autoregressive ASR with emotion recognition and audio-event tags. Useful for richer transcription than words alone. Custom model license.
FireRedTeam/FireRedASR2-AED
Automatic Speech Recognition • Updated • 816 • 33Note Mandarin, Chinese dialects/accents, English and code-switching; speech and singing transcription. ASR component of FireRedASR2S. Apache-2.0.
distil-whisper/distil-large-v3
Automatic Speech Recognition • 0.8B • Updated • 520k • 383Note Distilled Whisper large-v3 for English transcription, including long-form audio. Useful speed-oriented Whisper baseline. MIT.
-
Edge0/Audio8-ASR-Infinite
Automatic Speech Recognition • 4B • Updated • 23.7k • 1.53k