跳到主要内容
Home / 模型 / 音频

诚实的 AI 观点,
来自构建者

没有新闻稿。没有赞助内容。没有炒作。只是每天用 AI 构建、并敢说真心话的人的观察。

来自18家音频提供商的实时数据——每日更新
显示 54 / 54 个音频模型
🗣 语音
Async Flash
Async
Ultra-fast real-time text-to-speech optimized for low-latency applications
🗣 语音在Zubnet上试用 →
🗣 语音
AsyncFlow Multilingual
Async
Multilingual speech synthesis supporting 15 languages including Arabic, Japanese, and Chinese
🗣 语音在Zubnet上试用 →
🗣 语音
AsyncFlow v2.0
Async
High-quality text-to-speech with natural prosody and emotional expression across 15 languages
🗣 语音在Zubnet上试用 →
🗣 语音
Aura
Deepgram
Deepgram's original text-to-speech model, 12 English voices at half the cost of Aura 2. Fast and reliable.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Aura 2
Deepgram
Deepgram's latest text-to-speech model, 90+ natural voices across 8 languages (EN, FR, DE, ES, IT, NL, JA), sub-200ms latency. Greek mythology-themed voice names.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Bulbul v2
Sarvam AI
Sarvam's stable text-to-speech model for Indian languages and English. Returns base64-encoded WAV audio. Production-ready with consistent quality.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Bulbul v3
Sarvam AI
Sarvam's latest text-to-speech model, 30+ voices across Indian languages and English. Natural prosody with cultural intonation. Currently in beta.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Eleven v3
ElevenLabs
Most expressive TTS model with audio tags, dialogue mode, accent emulation, and 70+ languages. Best for long-form content.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Flash v2.5
ElevenLabs
Ultra-low latency TTS for real-time and conversational AI. ~75ms latency.
🗣 语音在Zubnet上试用 →
🗣 语音
Gemini 2.5 Flash TTS
Google
Google's low-latency text-to-speech with natural prosody and controllable style. Supports 24 languages.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Gemini 2.5 Pro TTS
Google
Premium text-to-speech with enhanced expressivity, richer tone, and precision pacing. Best for high-quality output.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Gemini 3.1 Flash TTS
Google
Latest price-performant, low-latency controllable speech generation. 30 voices, 24 languages.
textaudio
🗣 语音在Zubnet上试用 →
🗣 语音
Grok TTS
xAI / Grok
xAI's expressive text-to-speech — 26 multilingual voices with inline speech tags for tone, pauses, whispers and laughter. 20+ languages with auto-detection.
🗣 语音在Zubnet上试用 →
📝 转录
Hakim Arabic v2
Hakim
Arabic-first speech-to-text with dialect-aware recognition and broad multilingual coverage.
$0.000292/track
audiotext
📝 转录在Zubnet上试用 →
🗣 语音
Hakim Fast v1
Hakim
Arabic-first text-to-speech, low-latency tier, with 50 voices spanning Arabic dialects and 20+ languages.
textaudio
🗣 语音在Zubnet上试用 →
🔊 音效
Long Video to SFX 1.6
Mirelo
Generate sound effects for longer videos (SFX 1.6), up to 60 seconds of audio from a video URL.
$0.1750/track
videoaudio
🔊 音效在Zubnet上试用 →
🎵 音乐
Lyria 3 Clip
Google
Google DeepMind's latest music generation — 30-second compositions from text prompts with SynthID watermarking.
$0.0700/track
textaudio
🎵 音乐在Zubnet上试用 →
🎵 音乐
Lyria 3 Pro
Google
Full-song generation from Google DeepMind's Lyria 3 — complete compositions from text prompts with SynthID watermarking.
$0.1400/track
textaudio
🎵 音乐在Zubnet上试用 →
🗣 语音
Multilingual v2
ElevenLabs
Most advanced emotionally-aware speech synthesis with rich expression across 29 languages
🗣 语音在Zubnet上试用 →
🗣 语音
Murf Gen2
Murf
$0.000053/track
🗣 语音在Zubnet上试用 →
🗣 语音
Octave 1
Hume
First-generation empathic speech synthesis with natural emotional expression
🗣 语音在Zubnet上试用 →
🗣 语音
Octave 2
Hume
Latest empathic text-to-speech model supporting 11 languages with emotional awareness
🗣 语音在Zubnet上试用 →
🗣 语音
Rime Arcana
Rime
Rime's flagship voice model, 269 voices across 9 languages with rich emotional range.
textaudiotts
🗣 语音在Zubnet上试用 →
🗣 语音
Rime Coda
Rime
Stylized voices in 4 categories (Professional / Formal / Casual / Energetic). 184 voices, 8 languages.
textaudiotts
🗣 语音在Zubnet上试用 →
🗣 语音
Rime Mist
Rime
Original Mist model, 117 English voices, fast.
textaudiotts
🗣 语音在Zubnet上试用 →
🗣 语音
Rime Mist v2
Rime
Mist generation 2, 141 voices across 4 languages.
textaudiotts
🗣 语音在Zubnet上试用 →
🗣 语音
Rime Mist v3
Rime
Latest Mist generation, 83 voices, 4 languages. Improved expressiveness.
textaudiotts
🗣 语音在Zubnet上试用 →
📝 转录
Scribe v1
ElevenLabs
Batch speech recognition with word-level timestamps and language detection across 99 languages.
$0.000170/track
audiotext
📝 转录在Zubnet上试用 →
📝 转录
Scribe v2 Realtime
ElevenLabs
Ultra-low latency (<150ms) live speech recognition. 93.5% accuracy across 90+ languages. WebSocket streaming with VAD.
$0.000170/track
audiotext
📝 转录在Zubnet上试用 →
🗣 语音
Simba English
Speechify
High-quality English-only voice synthesis with natural-sounding output
🗣 语音在Zubnet上试用 →
🗣 语音
Simba Multilingual
Speechify
Natural voice synthesis supporting 60+ languages with comprehensive multilingual capabilities
🗣 语音在Zubnet上试用 →
🗣 语音
Simba Turbo
Speechify
Ultra-fast English voice synthesis optimized for speed with minimal latency
🗣 语音在Zubnet上试用 →
🗣 语音
Sonic 3
Cartesia
World's fastest, most emotive ultra-realistic text-to-speech with 60+ emotions and 42 languages
🗣 语音在Zubnet上试用 →
🗣 语音
Sonic Turbo
Cartesia
Ultra-low latency (40ms) speech generation optimized for real-time applications
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 02 HD
MiniMax
Focuses on rhythm, stability, and high-quality voice replication
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 02 Turbo
MiniMax
Enhanced multilingual capabilities with turbo speed
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 2.6 HD
MiniMax
Latest HD variant emphasizing prosody and voice cloning quality
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 2.6 Turbo
MiniMax
Turbo performance with 40 language support and low latency
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 2.8 HD
MiniMax
Current HD voice model. Highest fidelity prosody and voice cloning.
🗣 语音在Zubnet上试用 →
🗣 语音
Speech 2.8 Turbo
MiniMax
Current Turbo voice model. Low latency, broad language coverage.
🗣 语音在Zubnet上试用 →
📝 转录
Speechmatics Enhanced
Speechmatics
$0.000364/track
📝 转录在Zubnet上试用 →
📝 转录
Speechmatics Melia
Speechmatics
$0.000117/track
📝 转录在Zubnet上试用 →
📝 转录
Speechmatics Standard
Speechmatics
$0.000219/track
📝 转录在Zubnet上试用 →
🎵 音乐
Stable Audio 2
StabilityAI
Generate music and sound effects up to 3 minutes from text prompts. Produces structured compositions with intros, development, and outros at 44.1kHz stereo.
$0.3500/track
🎵 音乐在Zubnet上试用 →
🎵 音乐
Stable Audio 2.5
StabilityAI
Enterprise-grade music and sound generation. Produces structured compositions with intros, development, and outros at 44.1kHz stereo. 8-step inference for fast generation.
$0.3500/track
🎵 音乐在Zubnet上试用 →
🎚 音轨分离
Stem Separation
ElevenLabs
Separate audio into individual stems (vocals, drums, bass, etc). 2-stem or 6-stem modes.
$0.0029/track
audioaudio
🎚 音轨分离在Zubnet上试用 →
🎵 音乐
Suno V3.5
Suno
Better song organization with clear verse/chorus patterns
$0.1050/track
🎵 音乐在Zubnet上试用 →
🎵 音乐
Suno V4
Suno
Enhanced vocal quality and refined audio processing for music generation
$0.1050/track
🎵 音乐在Zubnet上试用 →
🎵 音乐
Suno V4.5
Suno
Excellent prompt understanding with faster generation speeds, supports up to 8 minute tracks
$0.1050/track
🎵 音乐在Zubnet上试用 →
🎵 音乐
Suno V4.5 Plus
Suno
Advanced model with enhanced tonal variation and excellent prompt understanding
$0.1050/track
🎵 音乐在Zubnet上试用 →
🎵 音乐
Suno V5
Suno
Cutting-edge model with enhanced quality and capabilities for AI music generation
$0.1050/track
🎵 音乐在Zubnet上试用 →
🗣 语音
Turbo v2.5
ElevenLabs
Low-latency speech generation in 32 languages, optimized for real-time conversational AI
🗣 语音在Zubnet上试用 →
🔊 音效
Video to SFX 1.6
Mirelo
Generate and edit sound effects from video (SFX 1.6). Provide a video URL and optional text prompt; adds seamless extension, looping ambiences, and AI inpainting to erase/replace moments.
$0.0875/track
videotextaudio
🔊 音效在Zubnet上试用 →
🗣 语音
Vidu Text to Speech
Vidu
Natural text-to-speech with adjustable speed, volume, and pitch.
🗣 语音在Zubnet上试用 →
ESC