Audio y voz con IA
Cada modelo de audio que ofrecemos — texto a voz, composición musical, transcripción, clonación de voz, aislamiento vocal, separación de pistas. De ElevenLabs y Suno a Cartesia y Deepgram. Precios reales, capacidades reales.
Datos en vivo de 18 proveedores de audio — actualizado diariamente
🗣 Voz
Ultra-fast real-time text-to-speech optimized for low-latency applications
🗣 VozProbar en Zubnet →
🗣 Voz
Multilingual speech synthesis supporting 15 languages including Arabic, Japanese, and Chinese
🗣 VozProbar en Zubnet →
🗣 Voz
High-quality text-to-speech with natural prosody and emotional expression across 15 languages
🗣 VozProbar en Zubnet →
🗣 Voz
Deepgram's original text-to-speech model, 12 English voices at half the cost of Aura 2. Fast and reliable.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Deepgram's latest text-to-speech model, 90+ natural voices across 8 languages (EN, FR, DE, ES, IT, NL, JA), sub-200ms latency. Greek mythology-themed voice names.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Sarvam's stable text-to-speech model for Indian languages and English. Returns base64-encoded WAV audio. Production-ready with consistent quality.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Sarvam's latest text-to-speech model, 30+ voices across Indian languages and English. Natural prosody with cultural intonation. Currently in beta.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Most expressive TTS model with audio tags, dialogue mode, accent emulation, and 70+ languages. Best for long-form content.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Ultra-low latency TTS for real-time and conversational AI. ~75ms latency.
🗣 VozProbar en Zubnet →
🗣 Voz
Google's low-latency text-to-speech with natural prosody and controllable style. Supports 24 languages.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Premium text-to-speech with enhanced expressivity, richer tone, and precision pacing. Best for high-quality output.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
Latest price-performant, low-latency controllable speech generation. 30 voices, 24 languages.
textaudio
🗣 VozProbar en Zubnet →
🗣 Voz
xAI's expressive text-to-speech — 26 multilingual voices with inline speech tags for tone, pauses, whispers and laughter. 20+ languages with auto-detection.
🗣 VozProbar en Zubnet →
📝 Transcripción
Arabic-first speech-to-text with dialect-aware recognition and broad multilingual coverage.
$0.000292/track
audiotext
📝 TranscripciónProbar en Zubnet →
🗣 Voz
Arabic-first text-to-speech, low-latency tier, with 50 voices spanning Arabic dialects and 20+ languages.
textaudio
🗣 VozProbar en Zubnet →
🔊 Efectos de sonido
Generate sound effects for longer videos (SFX 1.6), up to 60 seconds of audio from a video URL.
$0.1750/track
videoaudio
🔊 Efectos de sonidoProbar en Zubnet →
🎵 Música
Google DeepMind's latest music generation — 30-second compositions from text prompts with SynthID watermarking.
$0.0700/track
textaudio
🎵 MúsicaProbar en Zubnet →
🎵 Música
Full-song generation from Google DeepMind's Lyria 3 — complete compositions from text prompts with SynthID watermarking.
$0.1400/track
textaudio
🎵 MúsicaProbar en Zubnet →
🗣 Voz
Most advanced emotionally-aware speech synthesis with rich expression across 29 languages
🗣 VozProbar en Zubnet →
🗣 Voz
$0.000053/track
🗣 VozProbar en Zubnet →
🗣 Voz
First-generation empathic speech synthesis with natural emotional expression
🗣 VozProbar en Zubnet →
🗣 Voz
Latest empathic text-to-speech model supporting 11 languages with emotional awareness
🗣 VozProbar en Zubnet →
🗣 Voz
Rime's flagship voice model, 269 voices across 9 languages with rich emotional range.
textaudiotts
🗣 VozProbar en Zubnet →
🗣 Voz
Stylized voices in 4 categories (Professional / Formal / Casual / Energetic). 184 voices, 8 languages.
textaudiotts
🗣 VozProbar en Zubnet →
🗣 Voz
Original Mist model, 117 English voices, fast.
textaudiotts
🗣 VozProbar en Zubnet →
🗣 Voz
Mist generation 2, 141 voices across 4 languages.
textaudiotts
🗣 VozProbar en Zubnet →
🗣 Voz
Latest Mist generation, 83 voices, 4 languages. Improved expressiveness.
textaudiotts
🗣 VozProbar en Zubnet →
📝 Transcripción
Batch speech recognition with word-level timestamps and language detection across 99 languages.
$0.000170/track
audiotext
📝 TranscripciónProbar en Zubnet →
📝 Transcripción
Ultra-low latency (<150ms) live speech recognition. 93.5% accuracy across 90+ languages. WebSocket streaming with VAD.
$0.000170/track
audiotext
📝 TranscripciónProbar en Zubnet →
🗣 Voz
High-quality English-only voice synthesis with natural-sounding output
🗣 VozProbar en Zubnet →
🗣 Voz
Natural voice synthesis supporting 60+ languages with comprehensive multilingual capabilities
🗣 VozProbar en Zubnet →
🗣 Voz
Ultra-fast English voice synthesis optimized for speed with minimal latency
🗣 VozProbar en Zubnet →
🗣 Voz
World's fastest, most emotive ultra-realistic text-to-speech with 60+ emotions and 42 languages
🗣 VozProbar en Zubnet →
🗣 Voz
Ultra-low latency (40ms) speech generation optimized for real-time applications
🗣 VozProbar en Zubnet →
🗣 Voz
Focuses on rhythm, stability, and high-quality voice replication
🗣 VozProbar en Zubnet →
🗣 Voz
Enhanced multilingual capabilities with turbo speed
🗣 VozProbar en Zubnet →
🗣 Voz
Latest HD variant emphasizing prosody and voice cloning quality
🗣 VozProbar en Zubnet →
🗣 Voz
Turbo performance with 40 language support and low latency
🗣 VozProbar en Zubnet →
🗣 Voznuevo
Current HD voice model. Highest fidelity prosody and voice cloning.
🗣 VozProbar en Zubnet →
🗣 Voznuevo
Current Turbo voice model. Low latency, broad language coverage.
🗣 VozProbar en Zubnet →
📝 Transcripción
$0.000364/track
📝 TranscripciónProbar en Zubnet →
📝 Transcripción
$0.000117/track
📝 TranscripciónProbar en Zubnet →
📝 Transcripción
$0.000219/track
📝 TranscripciónProbar en Zubnet →
🎵 Música
Generate music and sound effects up to 3 minutes from text prompts. Produces structured compositions with intros, development, and outros at 44.1kHz stereo.
$0.3500/track
🎵 MúsicaProbar en Zubnet →
🎵 Música
Enterprise-grade music and sound generation. Produces structured compositions with intros, development, and outros at 44.1kHz stereo. 8-step inference for fast generation.
$0.3500/track
🎵 MúsicaProbar en Zubnet →
🎚 Separación
Separate audio into individual stems (vocals, drums, bass, etc). 2-stem or 6-stem modes.
$0.0029/track
audioaudio
🎚 SeparaciónProbar en Zubnet →
🎵 Música
Better song organization with clear verse/chorus patterns
$0.1050/track
🎵 MúsicaProbar en Zubnet →
🎵 Música
Enhanced vocal quality and refined audio processing for music generation
$0.1050/track
🎵 MúsicaProbar en Zubnet →
🎵 Música
Excellent prompt understanding with faster generation speeds, supports up to 8 minute tracks
$0.1050/track
🎵 MúsicaProbar en Zubnet →
🎵 Música
Advanced model with enhanced tonal variation and excellent prompt understanding
$0.1050/track
🎵 MúsicaProbar en Zubnet →
🎵 Música
Cutting-edge model with enhanced quality and capabilities for AI music generation
$0.1050/track
🎵 MúsicaProbar en Zubnet →
🗣 Voz
Low-latency speech generation in 32 languages, optimized for real-time conversational AI
🗣 VozProbar en Zubnet →
🔊 Efectos de sonido
Generate and edit sound effects from video (SFX 1.6). Provide a video URL and optional text prompt; adds seamless extension, looping ambiences, and AI inpainting to erase/replace moments.
$0.0875/track
videotextaudio
🔊 Efectos de sonidoProbar en Zubnet →
🗣 Voz
Natural text-to-speech with adjustable speed, volume, and pitch.
🗣 VozProbar en Zubnet →