AI Audio & Voice
Every audio model we serve — text-to-speech, music composition, transcription, voice cloning, voice isolation, stem separation. From ElevenLabs and Suno to Cartesia and Deepgram. Real pricing, real capabilities.
Live data from 20 audio providers — updated daily
🗣 Voice
Async's legacy low-latency model with the broadest language coverage (15 languages including Arabic, Russian, Japanese, Hebrew, Armenian, Turkish, Hindi and Chinese); speed and stability controls, no text normalisation.
🗣 VoiceTry on Zubnet →
🗣 VoiceNew
Async's latency-optimised streaming model for real-time apps and voice agents: English, Spanish, French, German, Italian and Portuguese, with built-in normalisation of dates, currencies, numbers and abbreviations.
🗣 VoiceTry on Zubnet →
🗣 VoiceNew
Async's highest-quality English model for content production and audiobooks, with built-in text normalisation; twice the price of Flash.
🗣 VoiceTry on Zubnet →
🗣 Voice
Deepgram's original text-to-speech model, 12 English voices at half the cost of Aura 2. Fast and reliable.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Deepgram's latest text-to-speech model, 90+ natural voices across 8 languages (EN, FR, DE, ES, IT, NL, JA), sub-200ms latency. Greek mythology-themed voice names.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Sarvam's stable text-to-speech model for Indian languages and English. Returns base64-encoded WAV audio. Production-ready with consistent quality.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Sarvam's latest text-to-speech model, 30+ voices across Indian languages and English. Natural prosody with cultural intonation. Currently in beta.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
English TTS with emotion control and zero-shot voice cloning
🗣 VoiceTry on Zubnet →
🗣 Voice
Multilingual TTS supporting 23 languages with natural prosody
🗣 VoiceTry on Zubnet →
🗣 Voice
Most expressive TTS model with audio tags, dialogue mode, accent emulation, and 70+ languages. Best for long-form content.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Ultra-low latency TTS for real-time and conversational AI. ~75ms latency.
🗣 VoiceTry on Zubnet →
🗣 Voice
Google's low-latency text-to-speech with natural prosody and controllable style. Supports 24 languages.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Premium text-to-speech with enhanced expressivity, richer tone, and precision pacing. Best for high-quality output.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
Latest price-performant, low-latency controllable speech generation. 30 voices, 24 languages.
textaudio
🗣 VoiceTry on Zubnet →
🗣 Voice
xAI's expressive text-to-speech — 26 multilingual voices with inline speech tags for tone, pauses, whispers and laughter. 20+ languages with auto-detection.
🗣 VoiceTry on Zubnet →
📝 Transcription
Arabic-first speech-to-text with dialect-aware recognition and broad multilingual coverage.
$0.000292/track
audiotext
📝 TranscriptionTry on Zubnet →
🗣 Voice
Arabic-first text-to-speech, low-latency tier, with 50 voices spanning Arabic dialects and 20+ languages.
textaudio
🗣 VoiceTry on Zubnet →
🔊 Sound FX
Generate sound effects for longer videos (SFX 1.6), up to 60 seconds of audio from a video URL.
$0.1750/track
videoaudio
🔊 Sound FXTry on Zubnet →
🎵 Music
Google DeepMind's latest music generation — 30-second compositions from text prompts with SynthID watermarking.
$0.0700/track
textaudio
🎵 MusicTry on Zubnet →
🎵 Music
Full-song generation from Google DeepMind's Lyria 3 — complete compositions from text prompts with SynthID watermarking.
$0.1400/track
textaudio
🎵 MusicTry on Zubnet →
🎵 MusicNew
Google DeepMind's Lyria 3.5 — full-length songs with verses, choruses and bridges, vocals and timed lyrics, 44.1 kHz stereo, SynthID watermarking.
$0.1400/track
textaudio
🎵 MusicTry on Zubnet →
🗣 Voice
Most advanced emotionally-aware speech synthesis with rich expression across 29 languages
🗣 VoiceTry on Zubnet →
🗣 Voice
$0.000053/track
🗣 VoiceTry on Zubnet →
🗣 Voice
First-generation empathic speech synthesis with natural emotional expression
🗣 VoiceTry on Zubnet →
🗣 Voice
Latest empathic text-to-speech model supporting 11 languages with emotional awareness
🗣 VoiceTry on Zubnet →
🗣 Voice
Rime's flagship voice model, 269 voices across 9 languages with rich emotional range.
textaudiotts
🗣 VoiceTry on Zubnet →
🗣 Voice
Stylized voices in 4 categories (Professional / Formal / Casual / Energetic). 184 voices, 8 languages.
textaudiotts
🗣 VoiceTry on Zubnet →
🗣 Voice
Original Mist model, 117 English voices, fast.
textaudiotts
🗣 VoiceTry on Zubnet →
🗣 Voice
Mist generation 2, 141 voices across 4 languages.
textaudiotts
🗣 VoiceTry on Zubnet →
🗣 Voice
Latest Mist generation, 83 voices, 4 languages. Improved expressiveness.
textaudiotts
🗣 VoiceTry on Zubnet →
📝 Transcription
Batch speech recognition with word-level timestamps and language detection across 99 languages.
$0.000170/track
audiotext
📝 TranscriptionTry on Zubnet →
📝 Transcription
Ultra-low latency (<150ms) live speech recognition. 93.5% accuracy across 90+ languages. WebSocket streaming with VAD.
$0.000170/track
audiotext
📝 TranscriptionTry on Zubnet →
🗣 VoiceNew
Speechify's streaming-native multilingual model: English plus German, Spanish (Spain and Mexico), French, Italian and Brazilian Portuguese, and it accepts every Speechify voice whatever its locale.
🗣 VoiceTry on Zubnet →
🗣 VoiceNew
Speechify's recommended streaming-native English model: lowest time-to-first-byte in the Simba family, fine-grained emotional control and SSML prosody. English voices only.
🗣 VoiceTry on Zubnet →
🗣 VoiceNew
Cartesia's most natural streaming text-to-speech (2026-08-27): pacing and intonation from context, 60+ emotions, speed and volume controls, 44 languages including Odia and Urdu, sub-100 ms first audio.
🗣 VoiceTry on Zubnet →
🗣 Voice
Focuses on rhythm, stability, and high-quality voice replication
🗣 VoiceTry on Zubnet →
🗣 Voice
Enhanced multilingual capabilities with turbo speed
🗣 VoiceTry on Zubnet →
🗣 Voice
Latest HD variant emphasizing prosody and voice cloning quality
🗣 VoiceTry on Zubnet →
🗣 Voice
Turbo performance with 40 language support and low latency
🗣 VoiceTry on Zubnet →
🗣 Voice
Current HD voice model. Highest fidelity prosody and voice cloning.
🗣 VoiceTry on Zubnet →
🗣 Voice
Current Turbo voice model. Low latency, broad language coverage.
🗣 VoiceTry on Zubnet →
📝 Transcription
$0.000364/track
📝 TranscriptionTry on Zubnet →
📝 Transcription
$0.000117/track
📝 TranscriptionTry on Zubnet →
📝 Transcription
$0.000219/track
📝 TranscriptionTry on Zubnet →
🎵 Music
Generate music and sound effects up to 3 minutes from text prompts. Produces structured compositions with intros, development, and outros at 44.1kHz stereo.
$0.3500/track
🎵 MusicTry on Zubnet →
🎵 Music
Enterprise-grade music and sound generation. Produces structured compositions with intros, development, and outros at 44.1kHz stereo. 8-step inference for fast generation.
$0.3500/track
🎵 MusicTry on Zubnet →
🎚 Stems
Separate audio into individual stems (vocals, drums, bass, etc). 2-stem or 6-stem modes.
$0.0029/track
audioaudio
🎚 StemsTry on Zubnet →
🎵 Music
Better song organization with clear verse/chorus patterns
$0.1050/track
🎵 MusicTry on Zubnet →
🎵 Music
Enhanced vocal quality and refined audio processing for music generation
$0.1050/track
🎵 MusicTry on Zubnet →
🎵 Music
Excellent prompt understanding with faster generation speeds, supports up to 8 minute tracks
$0.1050/track
🎵 MusicTry on Zubnet →
🎵 Music
Advanced model with enhanced tonal variation and excellent prompt understanding
$0.1050/track
🎵 MusicTry on Zubnet →
🎵 Music
Cutting-edge model with enhanced quality and capabilities for AI music generation
$0.1050/track
🎵 MusicTry on Zubnet →
🗣 Voice
Low-latency speech generation in 32 languages, optimized for real-time conversational AI
🗣 VoiceTry on Zubnet →
🔊 Sound FX
Generate and edit sound effects from video (SFX 1.6). Provide a video URL and optional text prompt; adds seamless extension, looping ambiences, and AI inpainting to erase/replace moments.
$0.0875/track
videotextaudio
🔊 Sound FXTry on Zubnet →
🗣 Voice
Natural text-to-speech with adjustable speed, volume, and pitch.
🗣 VoiceTry on Zubnet →