AI Model Directory
Every model. Every provider. Live data, real pricing, honest capabilities. Updated daily from our API monitor.
Live data — updated daily
AI performance capture. Upload a character image + your performance video to transfer your motion to the character.
$0.0875/unit
videoimage
🎬 VideoRead more →
Runway's current video-to-video model. Edit existing footage with text prompts; accepts 2-30 second inputs.
$0.4900/unit
video
🎬 VideoRead more →
Async's legacy low-latency model with the broadest language coverage (15 languages including Arabic, Russian, Japanese, Hebrew, Armenian, Turkish, Hindi and Chinese); speed and stability controls, no text normalisation.
In: $0.0192/1M
🗣 VoiceRead more →
New
Async's latency-optimised streaming model for real-time apps and voice agents: English, Spanish, French, German, Italian and Portuguese, with built-in normalisation of dates, currencies, numbers and abbreviations.
In: $0.0192/1M
🗣 VoiceRead more →
New
Async's highest-quality English model for content production and audiobooks, with built-in text normalisation; twice the price of Flash.
In: $0.0385/1M
🗣 VoiceRead more →
Deepgram's original text-to-speech model, 12 English voices at half the cost of Aura 2. Fast and reliable.
In: $0.0262/1M
Stream
🗣 VoiceRead more →
Deepgram's latest text-to-speech model, 90+ natural voices across 8 languages (EN, FR, DE, ES, IT, NL, JA), sub-200ms latency. Greek mythology-themed voice names.
In: $0.0525/1M
Stream
🗣 VoiceRead more →
Sarvam's stable text-to-speech model for Indian languages and English. Returns base64-encoded WAV audio. Production-ready with consistent quality.
In: $0.000315/1M
🗣 VoiceRead more →
Sarvam's latest text-to-speech model, 30+ voices across Indian languages and English. Natural prosody with cultural intonation. Currently in beta.
In: $0.000630/1M
🗣 VoiceRead more →
ByteDance Seed 1.8, multimodal agent model (text, image & video in → text), strong tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Code (preview), coding-specialized agent (text & image in → text), strong tool use + reasoning. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
ByteDance Seed 2.0 Lite, efficient multimodal agent (text, image & video in → text), tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Mini, fast lightweight multimodal agent (text, image & video in → text), tool use. Native via BytePlus (Ark).
131K ctx33K outIn: $0.000175/1MOut: $0.000700/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Pro, flagship multimodal agent (text, image & video in → text), strong tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.1 Turbo, fast multimodal reasoning agent (text & image in → text), strong tool use + thinking, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Specialized model for code generation, completion, and understanding.
262K ctx8K outIn: $0.000525/1MOut: $0.0016/1M
Streamchat
💬 LLMRead more →
SOTA open-source image model. Native Chinese text support, 6B parameters.
$0.0175/unit
🎨 ImageRead more →
High-quality neural machine translation supporting 33 languages
In: $0.0437/1M
🌐 TranslationRead more →
DeepSeek flagship V4. 512K context, advanced reasoning + tool use.
524K ctx66K outIn: $0.0012/1MOut: $0.0035/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
New
DeepSeek V4.1 Flash, the fast daily driver that replaced V4 Flash on 2026-09-10. 1M context, thinking (default) and non-thinking modes.
1M ctx393K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
Most expressive TTS model with audio tags, dialogue mode, accent emulation, and 70+ languages. Best for long-form content.
In: $0.3850/1M
🗣 VoiceRead more →
New
Low-latency v3 tuned for realtime conversation (~280 ms), high-quality expressive delivery with custom audio tags, 70+ languages.
In: $0.1925/1M
🗣 VoiceRead more →
New
ElevenLabs' most expressive model (September 2026): tone, pacing and emotion read from the text, speaker identity held across long-form, audio tags, 85 languages, up to 10,000 characters per request.
In: $0.3850/1M
🗣 VoiceRead more →
New
Low-latency Eleven v4 for realtime voice (~100 ms median inference), audio tags, 85 languages, up to 10,000 characters per request.
In: $0.1925/1M
🗣 VoiceRead more →
Fable 5 by Anthropic.
1M ctx128K outIn: $0.0175/1MOut: $0.0875/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Fable 5.1 by Anthropic: the most capable generally available Claude, for demanding reasoning and long-horizon agentic work. 1M context, 128K output, adaptive thinking always on.
1M ctx128K outIn: $0.0175/1MOut: $0.0875/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Ultra-low latency TTS for real-time and conversational AI. ~75ms latency.
In: $0.1925/1M
🗣 VoiceRead more →
The best of FLUX, offering state-of-the-art performance image generation at blazing fast speeds
$0.0700/unit
🎨 ImageRead more →
New
FLUX 3 image generation: text-to-image, multi-reference editing with up to ten references, and bounding-box placement.
🎨 ImageRead more →
FLUX 3 multimodal video: text-to-video, image-to-video with up to 10 keyframes, and video continuation — up to 20 s in HD or FHD with native synchronized audio. Draft mode gives a fast, cheaper preview.
videoimagevideo
🎬 VideoRead more →
A premium model brings maximum performance across all aspects – greatly improved quality, consistency, and speed
$0.1400/unit
🎨 ImageRead more →
A unified model delivering local editing, generative modifications, and text-to-image capabilities
$0.0700/unit
🎨 ImageRead more →
Directly distilled from FLUX.1 [pro], an open-weight, guidance-distilled model for non-commercial use
$0.0437/unit
🎨 ImageRead more →
Specialized for typography and text rendering with adjustable guidance and generation steps
$0.0875/unit
🎨 ImageRead more →
Ultra-fast open-source FLUX model optimized for real-time generation at the lowest cost. Apache 2.0 licensed.
$0.0245/unit
🎨 ImageRead more →
Fastest FLUX model with sub-second inference. 9B params with Qwen3 text embedder for superior prompt understanding.
$0.0262/unit
🎨 ImageRead more →
Highest quality FLUX 2.0 with grounding search capability for photorealistic images up to 4MP
$0.1400/unit
🎨 ImageRead more →
Production-grade FLUX 2.0 with multi-reference image editing and precise control over colors, poses, and composition
$0.0525/unit
🎨 ImageRead more →
Sakana's multi-agent conductor, orchestrates a pool of frontier models for complex reasoning. Billed per token, orchestration included.
272K ctx16K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Fast and efficient model with adaptive thinking for complex tasks
1M ctx66K outIn: $0.000525/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Google's low-latency text-to-speech with natural prosody and controllable style. Supports 24 languages.
In: $0.0350/1M
🗣 VoiceRead more →
Cost-efficient high-throughput model for budget-conscious applications
1M ctx66K outIn: $0.000175/1MOut: $0.000700/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Enhanced thinking and reasoning for complex problems
1M ctx66K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Premium text-to-speech with enhanced expressivity, richer tone, and precision pacing. Best for high-quality output.
In: $0.0700/1M
🗣 VoiceRead more →
Google's fastest frontier model. Beats 2.5 Pro at 1/4 the cost with 1M context
1M ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Latest price-performant, low-latency controllable speech generation. 30 voices, 24 languages.
In: $0.0700/1M
🗣 VoiceRead more →
Cost-efficient high-throughput model for budget-conscious applications.
1M ctx66K outIn: $0.000438/1MOut: $0.0026/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Google's most capable agentic model with 1M token context, 77.1% ARC-AGI-2 reasoning, and native tool use.
1M ctx66K outIn: $0.0035/1MOut: $0.0210/1M
ToolsStreamtoolsreasoningimage+2
💬 LLMRead more →
Fast flagship with adaptive thinking for complex tasks.
1M ctx66K outIn: $0.0026/1MOut: $0.0158/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Cost-efficient GA Flash model for high-throughput multimodal tasks.
1M ctx66K outIn: $0.000525/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Google's latest GA Flash model for fast multimodal reasoning and agentic workloads.
1M ctx66K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Google's most capable Flash model for agentic workflows and multimodal reasoning.
1.048576M ctx66K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
New
High-throughput, low-latency conversational speech generation. 30 curated voices, 24 languages.
In: $0.0210/1M
🗣 VoiceRead more →
New
Studio-grade speech generation with expressive acting and long-form stability. 30 prebuilt voices, 24 languages.
In: $0.0315/1M
🗣 VoiceRead more →
Cost-efficient text embeddings — 3072-dim vectors for search, clustering, and RAG.
In: $0.000262/1M
🔍 EmbeddingRead more →
Google's latest embedding model — 3072-dim vectors for retrieval and semantic search (multimodal upstream; text wired).
In: $0.000350/1M
🔍 EmbeddingRead more →
Google's open-weight mixture-of-experts model, 26B total, 4B active parameters for fast inference
262K ctx33K outIn: $0.000105/1MOut: $0.000577/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Google's open-weight dense model with 262K context and strong multilingual reasoning
262K ctx33K outIn: $0.000245/1MOut: $0.000700/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Transform reference images with text prompts. Preserves identity while changing pose, lighting, background.
$0.0875/unit
image
🎨 ImageRead more →
Fast reference-based image generation. 2.5x faster, preserves identity and style.
$0.0350/unit
image
🎨 ImageRead more →
Runway's flagship image-to-video model. Exceptional fidelity, stability, and controllability. Requires input image.
$0.0875/unit
image
🎬 VideoRead more →
Runway's most capable video model. #1 on Artificial Analysis text-to-video benchmark. Text-to-video or image-to-video with exceptional physics and human motion.
$0.2100/unit
image
🎬 VideoRead more →
Previous flagship with thinking mode. MoE architecture, 128K context.
128K ctx8K outIn: $0.0010/1MOut: $0.0039/1M
Streamreasoning
💬 LLMRead more →
Lightweight model optimized for efficiency. 128K context.
128K ctx8K outIn: $0.000350/1MOut: $0.0019/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Free tier model. Great for testing and light workloads.
128K ctx4K out
ToolsStreamtoolsreasoning
💬 LLMRead more →
Vision-language model for image understanding and analysis.
66K ctx16K outIn: $0.0010/1MOut: $0.0032/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Flagship model with reasoning, coding, and agentic capabilities. 128K context.
128K ctx8K outIn: $0.0010/1MOut: $0.0039/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Vision-capable model for image understanding and analysis.
128K ctx8K outIn: $0.000525/1MOut: $0.0016/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Lightweight vision model. Fast and cost-effective for image tasks.
128K ctx4K outIn: $0.000070/1MOut: $0.000700/1M
ToolsStreamreasoningtoolsimage
💬 LLMRead more →
Latest flagship. 358B params, 204K context, 131K output. #1 on LiveCodeBench.
205K ctx131K outIn: $0.0010/1MOut: $0.0039/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Lightweight, completely free GLM-4.7 variant with 200K context.
205K ctx131K out
ToolsStreamtoolsreasoning
💬 LLMRead more →
Lightweight, high-speed and affordable GLM-4.7 variant with 200K context.
205K ctx131K outIn: $0.000122/1MOut: $0.000700/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Most capable Z.ai model. 744B params (40B active MoE), 28.5T training tokens. Built for complex systems engineering and agentic tasks.
205K ctx131K outIn: $0.0018/1MOut: $0.0056/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Fast variant of GLM-5 with optimized speed and competitive quality.
In: $0.0021/1MOut: $0.0070/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's current flagship for agentic and coding tasks.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's latest flagship for agentic and coding tasks.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's newest flagship for agentic and coding tasks. Always reasons; text-only input.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's fast, low-cost GLM-5.3 variant: 1M context, tools, always reasons; text and image input.
1M ctx131K outIn: $0.000262/1MOut: $0.000875/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
Native multimodal coding model. 744B MoE (40B active), 203K context. Optimized for design-to-code, GUI automation, and vision-grounded agentic tasks.
205K ctx131K outIn: $0.0021/1MOut: $0.0070/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Z.ai's image generation model with strong prompt adherence and text rendering.
$0.0262/unit
🎨 ImageRead more →
Specialized OCR model for extracting text from images and documents.
8K ctx8K outIn: $0.000053/1MOut: $0.000053/1M
Streamimage
ocrRead more →
OpenAI's latest image generation model (high-quality tier).
$0.3693/unit
imageimage_editimage
🎨 ImageRead more →
New
OpenAI's ChatGPT Images 2.5, fast variant: everyday generation at about half the latency of GPT Image 2.
$0.0927/unit
imageimage_editimage
🎨 ImageRead more →
New
OpenAI's ChatGPT Images 2.5, precision variant: detailed creative work and editing fidelity, slower than Flare.
$0.0927/unit
imageimage_editimage
🎨 ImageRead more →
OpenAI GPT-3.5 Turbo (chat completions endpoint).
16K ctx4K outIn: $0.000875/1MOut: $0.0026/1M
tools
💬 LLMRead more →
OpenAI GPT-3.5 Turbo 16K (chat completions endpoint).
16K ctx4K outIn: $0.0053/1MOut: $0.0070/1M
tools
💬 LLMRead more →
OpenAI GPT-4 (chat completions endpoint).
8K ctx8K outIn: $0.0525/1MOut: $0.1050/1M
tools
💬 LLMRead more →
OpenAI GPT-4.1 (chat completions endpoint).
1.047576M ctx33K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4.1 Mini (chat completions endpoint).
1.047576M ctx33K outIn: $0.000700/1MOut: $0.0028/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4o (chat completions endpoint).
128K ctx16K outIn: $0.0044/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4o Mini (chat completions endpoint).
128K ctx16K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Mini (chat completions endpoint).
272K ctx128K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Nano (chat completions endpoint).
272K ctx128K outIn: $0.000087/1MOut: $0.000700/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0262/1MOut: $0.2100/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5 Search API (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
Streamweb_searchimage
💬 LLMRead more →
OpenAI GPT-5.1 (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.2 (chat completions endpoint).
272K ctx128K outIn: $0.0031/1MOut: $0.0245/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.2 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0367/1MOut: $0.2940/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5.3 Codex (chat completions endpoint).
272K ctx128K outIn: $0.0031/1MOut: $0.0245/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI mid-tier flagship (March 2026).
272K ctx128K outIn: $0.0044/1MOut: $0.0262/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI fast/cheap chat model.
272K ctx128K outIn: $0.0013/1MOut: $0.0079/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI cheapest tier, very fast.
272K ctx128K outIn: $0.000350/1MOut: $0.0022/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.4 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0525/1MOut: $0.3150/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI flagship chat model (April 2026 release).
272K ctx128K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI premium tier with deep reasoning. April 2026.
272K ctx128K outIn: $0.0525/1MOut: $0.3150/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5.6 Luna — high-volume streaming tasks, fast + economical 5.6 tier.
272K ctx128K outIn: $0.0018/1MOut: $0.0105/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
OpenAI GPT-5.6 Sol — frontier logic + deep reasoning, the flagship 5.6 tier.
272K ctx128K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
OpenAI GPT-5.6 Terra — balanced business logic, mid 5.6 tier.
272K ctx128K outIn: $0.0044/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
OpenAI GPT-6 Astra — frontier reasoning over a 1M-token context window.
1.05M ctx128K outIn: $0.0175/1MOut: $0.0875/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
New
OpenAI GPT-6 Luna: the fast, low-cost GPT-6 tier for high-volume chat and light agentic work, 1M-token context window.
1.05M ctx128K outIn: $0.000175/1MOut: $0.000875/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
New
OpenAI GPT-6 Sol: GPT-6 reasoning at a fifth of Astra's price, over a 1M-token context window.
1.05M ctx128K outIn: $0.0035/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
Non-reasoning variant of Grok 4.20. Fastest 4.20 path for chat and high-throughput tasks.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Reasoning-mode variant of Grok 4.20. Slower but stronger on multi-step problems.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchatreasoningtools+1
💬 LLMRead more →
xAI's flagship general-purpose model. 1M context, balanced reasoning and chat.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
xAI's newest flagship, launched July 2026. 500K context, strongest reasoning + coding, trained alongside Cursor.
500K ctx8K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
xAI's newest flagship. 500K context, accepts text and images, strongest reasoning + coding.
500K ctx8K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
New
xAI's newest flagship (2026-09-21). 500K context, accepts text and images, function calling, structured outputs, reasoning.
500K ctx8K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
Compact model tuned for code generation and structured output. 256K context.
256K ctx8K outIn: $0.0018/1MOut: $0.0035/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Fast text-to-image generation from xAI. Standard quality, lowest cost.
$0.0350/unit
🎨 ImageRead more →
xAI's current image model. Higher fidelity than Grok Imagine, with stronger prompt adherence.
$0.1050/unit
🎨 ImageRead more →
Higher-fidelity Grok Imagine. Better prompt adherence and detail, 2.5x the cost.
$0.0875/unit
🎨 ImageRead more →
xAI Grok Imagine 1.5 — image-to-video: animates a starting frame from your prompt. Up to 1080p, 1-15s clips, async render.
image
🎬 VideoRead more →
xAI's expressive text-to-speech — 26 multilingual voices with inline speech tags for tone, pauses, whispers and laughter. 20+ languages with auto-detection.
In: $0.0262/1M
🗣 VoiceRead more →
Claude's fastest and most intelligent Haiku model
200K ctx64K outIn: $0.0018/1MOut: $0.0088/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
New
Haiku 5.5 by Anthropic (2026-10-07): the fastest Claude, for high-volume, latency-sensitive work such as classification, extraction and routing. 1M context, 128K output, adaptive thinking.
1M ctx128K outIn: $0.000175/1MOut: $0.000875/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Higher resolution (1080p), longer duration (10s), stronger prompt adherence
$0.4900/unit
image
🎬 VideoRead more →
Breakthroughs in body movement, facial expressions, and physical realism
$0.4900/unit
image
🎬 VideoRead more →
Image-to-video model optimized for value and efficiency (requires image upload)
$0.3325/unit
image
🎬 VideoRead more →
Arabic-first speech-to-text with dialect-aware recognition and broad multilingual coverage.
$0.000292/unit
audio
📝 TranscriptionRead more →
Arabic-first text-to-speech, low-latency tier, with 50 voices spanning Arabic dialects and 20+ languages.
In: $0.0700/1M
🗣 VoiceRead more →
Alibaba HappyHorse image-to-video. Animates a still image into a coherent motion sequence guided by text.
$0.2450/unit
videoimage
🎬 VideoRead more →
Alibaba HappyHorse reference-to-video. Generates video conditioned on a reference image plus text prompt.
$0.2450/unit
videoimage
🎬 VideoRead more →
Alibaba HappyHorse text-to-video. Generates short cinematic clips from text prompts.
$0.2450/unit
video
🎬 VideoRead more →
Alibaba HappyHorse 1.1 image-to-video. Animates a still image into coherent motion guided by text (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba HappyHorse 1.1 reference-to-video. Generates video that preserves the subject and scene from one or more reference images (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba HappyHorse 1.1 text-to-video. Generates short cinematic clips from text prompts (native DashScope, US-Virginia).
video
🎬 VideoRead more →
Alibaba HappyHorse video editing. Edits a source video from natural-language instructions, optionally guided by reference images (native DashScope, US-Virginia).
videovideoimage
🎬 VideoRead more →
Previous generation model. Still excellent for text rendering at lower cost.
$0.1050/unit
🎨 ImageRead more →
State-of-the-art image generation with exceptional text rendering. Best balance of quality and speed.
$0.1050/unit
🎨 ImageRead more →
Highest quality Ideogram 3.0. Maximum detail and fidelity for final outputs.
$0.1575/unit
🎨 ImageRead more →
Faster Ideogram 3.0 variant. Quick generation for iterative workflows.
$0.0525/unit
🎨 ImageRead more →
New
Ideogram 4.0 text-to-image generation with the best text rendering. The rendering speed is the priced tier: Turbo, Default or Quality.
$0.1050/unit
image
🎨 ImageRead more →
New
Precise image editing on Ideogram 4.5: the first uploaded image is edited from the prompt (a mask marks what to change), up to four reference images steer the result, and unedited pixels are kept from the original.
$0.3850/unit
imageimage_editimage
🎨 ImageRead more →
Jina AI's CLIP model, cross-modal text-image embeddings supporting 89 languages with Matryoshka representations. 865M params.
8K ctxIn: $0.000087/1M
image
🔍 EmbeddingRead more →
Jina AI's code-specialized embedding model, 1.5B parameters optimized for code search and retrieval. 32K context.
33K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's ColBERT late-interaction model, multi-vector embeddings for fine-grained retrieval, 89 languages, user-controlled embedding sizes.
8K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's 570M parameter multilingual text embedding model, MTEB benchmark champion, 8K context, 89 languages.
8K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's latest multimodal embedding model, embeds text and images into a shared vector space with 32K context. State-of-the-art retrieval quality.
33K ctxIn: $0.000087/1M
image
🔍 EmbeddingRead more →
Jina AI's multimodal reranker, re-ranks results containing both text and images for cross-modal search pipelines. On Zubnet it ranks text documents; image documents are not supported yet.
10K ctxIn: $0.000087/1M
↕️ RerankRead more →
Jina AI's latest reranker, novel listwise architecture, SOTA multilingual retrieval, massive 131K context window. 0.6B params.
134K ctxIn: $0.000087/1M
↕️ RerankRead more →
Latest Kimi K2 reasoning model. Replaces K2 thinking variants (k2-thinking and k2-thinking-turbo) being retired 2026-05-25.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
Coding-specialized Kimi K2 model (k2.7). 256K context, agentic tool use and function calling, tuned for software engineering.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
High-speed variant of Kimi K2.7 Code for low-latency agentic software engineering.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
Moonshot's open-weight 2.8T-parameter multimodal reasoning model with a 1M-token context window.
1.048576M ctx131K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamchatreasoningfunction_calling+2
💬 LLMRead more →
Latest turbo model with fast text-to-video and image-to-video generation
$0.3675/unit
🎬 VideoRead more →
First model with simultaneous audio-visual generation. Creates video with native speech, dialogue, sound effects, and ambient audio in Chinese/English. No post-production dubbing needed.
$0.6125/unit
🎬 VideoRead more →
Cost-effective audio-visual generation with native speech, dialogue, and sound effects. Standard tier for balanced quality and cost.
$0.3675/unit
🎬 VideoRead more →
Kling 3 Omni unified image model, text-to-image, image-to-image, and image editing.
$0.0490/unit
🎨 ImageRead more →
Kling 3 Omni, audio-visual generation with native sound, multi-shot, start/end frame & reference video. Standard tier.
$0.7350/unit
🎬 VideoRead more →
New
Kling 3.0 Turbo, 720p, native audio included. Text-to-video and image-to-video, 3 to 15 seconds.
$0.9800/unit
🎬 VideoRead more →
New
Kling 3.0 Turbo Pro, 1080p, native audio included. Text-to-video and image-to-video, 3 to 15 seconds.
$1.23/unit
🎬 VideoRead more →
Unified image model for text-to-image, image-to-image, and detail editing. Accepts up to 10 reference images for guided generation. Part of the O1 multimodal family.
$0.0490/unit
🎨 ImageRead more →
Unified multimodal video model combining generation, editing, and transformation. Supports text, image, and video inputs with natural language control. Cost-effective tier.
$0.7350/unit
🎬 VideoRead more →
High-speed anime-focused model excelling at a range of anime styles.
$0.0280/unit
generation
🎨 ImageRead more →
Core Leonardo model with stunning outputs and broad style range.
$0.0280/unit
generation
🎨 ImageRead more →
Cinematic-focused model excelling at film-like compositions and lighting.
$0.0280/unit
generation
🎨 ImageRead more →
High-speed generalist image generation model for fast iteration.
$0.0140/unit
generation
🎨 ImageRead more →
Generate sound effects for longer videos (SFX 1.6), up to 60 seconds of audio from a video URL.
$0.1750/unit
video
🔊 Sound FXRead more →
Leonardo high-speed model designed for photorealistic outputs.
$0.0280/unit
generation
🎨 ImageRead more →
Google DeepMind's latest music generation — 30-second compositions from text prompts with SynthID watermarking.
$0.0700/unit
🎵 MusicRead more →
Full-song generation from Google DeepMind's Lyria 3 — complete compositions from text prompts with SynthID watermarking.
$0.1400/unit
🎵 MusicRead more →
Google DeepMind's Lyria 3.5 — full-length songs with verses, choruses and bridges, vocals and timed lyrics, 44.1 kHz stereo, SynthID watermarking.
$0.1400/unit
🎵 MusicRead more →
Agentic capabilities with advanced reasoning and thinking process visibility
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Optimized for high concurrency and commercial use with advanced reasoning
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Frontier reasoning model with SOTA coding, agentic tool use, and complex real-world task performance
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Fast variant of M2.5 optimized for speed at ~100 tokens per second
In: $0.0010/1MOut: $0.0042/1M
ToolsStream
💬 LLMRead more →
Autonomous real-world productivity with agentic collaboration, live debugging, and professional document generation
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Fast variant of M2.7 for latency-sensitive work. 200K context, text only.
200K ctx128K outIn: $0.0010/1MOut: $0.0042/1M
ToolsStreamtools
💬 LLMRead more →
Efficient reasoning model for everyday logic tasks.
131K ctx40K outIn: $0.000875/1MOut: $0.0026/1M
Streamchatreasoning
💬 LLMRead more →
Transform images into detailed 3D models with textures and PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
Generate detailed 3D models from text descriptions with PBR textures
$1.05/unit
3d-generation
📦 3DRead more →
Transform images into detailed 3D models with the highest image alignment Meshy has shipped, textures and PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
Combine two to four views of the same subject into one aligned, textured 3D model with PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
New
Transform images into detailed 3D models with Meshy's sharpest geometry yet, textures and PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
New
Combine two to four views of the same subject into one aligned, textured 3D model with Meshy's sharpest geometry yet
$1.05/unit
3d-generationimage
📦 3DRead more →
New
Generate detailed 3D models from text descriptions with Meshy's sharpest geometry yet, textured with PBR maps
$1.05/unit
3d-generation
📦 3DRead more →
New
Xiaomi's fast, low-cost MiMo model (V2.6, 2026-09). 1M context, accepts text and images, function calling and structured output, reasons before it answers.
1.048576M ctx131K outIn: $0.000245/1MOut: $0.000490/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
New
Xiaomi's flagship MiMo model (V2.6, 2026-09). 1M context, accepts text and images, function calling and structured output, reasons before it answers.
1.048576M ctx131K outIn: $0.000761/1MOut: $0.0015/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
New
MiMo V2.6 Pro on Xiaomi's high-speed tier, at a premium price. 1M context, accepts text and images, function calling and structured output.
1.048576M ctx131K outIn: $0.0076/1MOut: $0.0152/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
High-performance edge model with vision. Best Ministral for complex tasks.
131K ctx8K outIn: $0.000350/1MOut: $0.000350/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Ultra-lightweight edge model with vision. 3B params, runs on phones/laptops.
131K ctx8K outIn: $0.000175/1MOut: $0.000175/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Compact edge model with vision. 8B params for local deployment.
131K ctx8K outIn: $0.000262/1MOut: $0.000262/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Mistral's text embedding model for semantic search, retrieval, and classification.
In: $0.000175/1M
🔍 EmbeddingRead more →
Open-weight flagship. 675B MoE (41B active), multimodal with vision. Apache 2.0 licensed.
131K ctx8K outIn: $0.000875/1MOut: $0.0026/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Premier frontier-class multimodal model. Great balance between Large and Small.
131K ctx8K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamchattoolsreasoning
💬 LLMRead more →
Mistral's content moderation model for detecting harmful, unsafe, or policy-violating content.
In: $0.000175/1M
⚡ CapabilityRead more →
Document AI model for high-accuracy OCR. Extracts text from images, PDFs, handwritten content, forms, and complex tables.
In: $0.0018/1MOut: $0.0053/1M
image
ocrRead more →
Fast and efficient model for simple tasks. Great balance of performance and cost.
131K ctx8K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamchattools
💬 LLMRead more →
Mistral's model optimized for command-line and terminal-based coding workflows.
128K ctx33K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamtools
💬 LLMRead more →
Most advanced emotionally-aware speech synthesis with rich expression across 29 languages
In: $0.3850/1M
🗣 VoiceRead more →
New
ElevenLabs' value-tier music generation: songs up to 10 minutes from a prompt, instrumental on request. Billed per generated minute upstream.
$0.0044/unit
🎵 MusicRead more →
New
ElevenLabs' studio-grade music generation: songs up to 10 minutes from a prompt, instrumental on request. Billed per generated minute upstream.
$0.0044/unit
🎵 MusicRead more →
Gemini 2.5 Flash Image, optimized for speed and efficiency, high-volume low-latency image generation.
66K ctx33K out$0.0683/unit
🎨 ImageRead more →
High-efficiency image generation, optimized for speed and high-volume use.
66K ctx33K out$0.1172/unit
image
🎨 ImageRead more →
Efficiency specialist of the Nano Banana family — ultra-low-latency, cost-effective image generation and editing.
66K ctx33K out$0.0588/unit
image
🎨 ImageRead more →
Google's native image generation with text rendering and multimodal understanding.
66K ctx33K out$0.2345/unit
image
🎨 ImageRead more →
OpenAI o3 (chat completions endpoint).
200K ctx100K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Premium model combining maximum intelligence with practical performance
200K ctx64K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Most capable Claude model with enhanced reasoning and coding
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Highly capable Claude Opus for complex reasoning and coding (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's most capable model for complex reasoning and agentic coding (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's flagship Opus — deep reasoning, agentic coding and long-horizon work (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
New
Opus 5.5 by Anthropic (2026-09-21): the next Opus, for long-running agentic coding and knowledge work. 1M context, 128K output, adaptive thinking always on.
1M ctx128K outIn: $0.0070/1MOut: $0.0350/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Large-context generalist with 128K window. Excellent for complex generation and analysis. Augure sovereign Canadian AI.
256K ctx33K outIn: $0.0013/1MOut: $0.0038/1M
ToolsStreamchattools
💬 LLMRead more →
Premium multimodal model with vision and tool use, strong bilingual EN/FR. Augure sovereign Canadian AI.
256K ctx33K outIn: $0.0032/1MOut: $0.0076/1M
ToolsStreamchattoolsreasoning+2
💬 LLMRead more →
New
Newest reasoning model: the fastest of the Augure menu, strong at code, 256K context. Augure sovereign Canadian AI.
256K ctx33K outIn: $0.0019/1MOut: $0.0038/1M
Streamchatreasoning
💬 LLMRead more →
Leonardo foundational model preview with extreme prompt adherence and versatility.
$0.0280/unit
generation
🎨 ImageRead more →
Leonardo flagship foundational model with exceptional prompt adherence and text rendering.
$0.0280/unit
generation
🎨 ImageRead more →
PixVerse C1, cinematic-quality video generation tuned for action scenes and reference-based work, clips from 1 to 15 seconds with optional generated audio.
video
🎬 VideoRead more →
PixVerse C1 image-to-video: cinematic animation of a source image into a 1–15 second clip with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse C1 fusion: cinematic video from 1–3 reference images (subjects or scenes) guided by a prompt, with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse V6, general-purpose video generation with clips from 1 to 15 seconds, optional generated audio, and multi-clip support.
video
🎬 VideoRead more →
PixVerse V6 image-to-video: animate a source image into a 1–15 second clip with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse V6 fusion: generate a video from 1–3 reference images (subjects or scenes) guided by a prompt, with optional generated audio.
videoimage
🎬 VideoRead more →
Perplexity's lightweight embedding model. 1024-dimensional INT8 vectors, 32K context, Matryoshka dimension reduction.
33K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Perplexity's most capable embedding model. 2560-dimensional INT8 vectors, 32K context, Matryoshka dimension reduction.
33K ctxIn: $0.000175/1M
🔍 EmbeddingRead more →
Pruna's optimized image editing: transform an input image with a prompt.
$0.0175/unit
image
🎨 ImageRead more →
Pruna's premium video generation: text-to-video and image-to-video, up to 1080p. Draft mode gives fast, cheap previews.
videoimage
🎬 VideoRead more →
Ultra-fast model with 1M context. Best latency and cost efficiency for simple tasks.
1M ctx8K outIn: $0.000039/1MOut: $0.000378/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Alibaba's flagship text-to-image model — successor to the whole Qwen-Image family, with strong prompt following and text rendering (native DashScope, Singapore).
$0.0612/unit
🎨 ImageRead more →
The higher-fidelity "pro" tier of Alibaba's Qwen Image 2.0 — stronger prompt following and text rendering (native DashScope, Singapore).
$0.1313/unit
🎨 ImageRead more →
New
Alibaba's Qwen Image 3.0 — the next generation of the Qwen text-to-image family: denser layouts, images within images, sharper small text, native rendering in 12 languages (native DashScope, Singapore).
$0.0525/unit
🎨 ImageRead more →
New
The higher-fidelity pro tier of Alibaba's Qwen Image 3.0 — prompts up to 4.5k tokens, dense information layouts, precise text down to 10 px (native DashScope, Singapore).
$0.0700/unit
🎨 ImageRead more →
Balanced flagship model with 1M context. Best performance/cost ratio for most tasks.
1M ctx8K outIn: $0.000700/1MOut: $0.0021/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Cost-effective coding model optimized for speed. Fast code generation and completion at lower cost.
131K ctx8K outIn: $0.000252/1MOut: $0.0010/1M
ToolsStreamtools
💬 LLMRead more →
Advanced vision-language model with 262K context. Excels at visual coding, spatial perception, and multimodal reasoning.
262K ctx8K outIn: $0.000250/1MOut: $0.0025/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Flagship 1T+ parameter model. State-of-the-art reasoning, coding, and agent capabilities.
262K ctx33K outIn: $0.000051/1MOut: $0.000502/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Alibaba Qwen3.5 Omni Flash, fast multimodal (text & image input, text output). Singapore deployment; native voice/audio coming.
66K ctx8K outIn: $0.000700/1MOut: $0.0039/1M
Streamimage
💬 LLMRead more →
Alibaba Qwen3.5 Omni, multimodal (text & image input, text output). Singapore deployment; native voice/audio coming.
66K ctx8K outIn: $0.0024/1MOut: $0.0145/1M
Streamimage
💬 LLMRead more →
397B MoE model with 17B active parameters, 1M token context, multimodal (text/image/video)
1M ctx8K outIn: $0.000201/1MOut: $0.0012/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Flagship 1T+ parameter model. State-of-the-art reasoning, coding, and agent capabilities.
262K ctx33K outIn: $0.000289/1MOut: $0.0017/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Qwen3.6 Plus, next-gen frontier LLM (Agentic Coding, vision/OCR, fine-grained localization). Replaces Qwen3.5 Plus.
1M ctx8K outIn: $0.000483/1MOut: $0.0029/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Alibaba's Agent Frontier flagship.
262K ctx33K outIn: $0.0029/1MOut: $0.0087/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Multimodal frontier (vision/video); cheaper than Max.
1M ctx8K outIn: $0.000483/1MOut: $0.0019/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Alibaba's 2.4T-parameter multimodal MoE flagship. 1M context, accepts text and images.
1M ctx33K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
New
Fast multimodal Qwen (text & image in, text out) on the Singapore tenant. Thinking on by default. Context measured served: 1,006,592 tokens.
1.006592M ctx131K outIn: $0.000262/1MOut: $0.000822/1M
ToolsStreamreasoningtoolsimage
💬 LLMRead more →
Classic pixel art generation with 15 styles. Great for game assets, characters, textures, and retro-styled visuals. 16-512px.
$0.0350/unit
🎨 ImageRead more →
Advanced pixel art with Flux model. Same 15 styles as Classic with enhanced quality and detail. 16-512px.
$0.0350/unit
🎨 ImageRead more →
Recraft V3 (Red Panda), supports text-to-image, image-to-image, inpainting, and background replacement. 1024x1024.
$0.0700/unit
image
🎨 ImageRead more →
Recraft V3 for vector graphics, SVG output with image-to-image, inpainting, and background replacement support.
$0.1400/unit
image
🎨 ImageRead more →
Recraft's latest image generation model, photorealistic and illustration styles, text rendering, 1024x1024, ~10s generation.
$0.0700/unit
🎨 ImageRead more →
Recraft's highest quality image model, 2048x2048 output, superior anatomy and detail, ideal for print-ready assets. ~30s generation.
$0.4375/unit
🎨 ImageRead more →
Recraft's highest quality vector model, high-resolution SVG output with fine detail, ideal for logos and brand assets.
$0.5250/unit
🎨 ImageRead more →
Recraft V4 for scalable vector graphics, generates production-quality SVG with discrete color regions and clean geometry.
$0.1400/unit
🎨 ImageRead more →
Voyage AI's latest reranking model, re-scores search results for maximum relevance. Best quality. 200M free tokens.
In: $0.000087/1M
↕️ RerankRead more →
Voyage AI's fast reranking model, cost-efficient re-scoring for high-volume retrieval pipelines. 200M free tokens.
In: $0.000035/1M
↕️ RerankRead more →
New
Voyage AI's next-generation reranker, strongest on long documents and code. 32K-token context.
In: $0.000087/1M
↕️ RerankRead more →
New
Voyage AI's fast next-generation reranker, cost-efficient re-scoring for high-volume retrieval. 32K-token context.
In: $0.000035/1M
↕️ RerankRead more →
Rime's flagship voice model, 269 voices across 9 languages with rich emotional range.
In: $0.0700/1M
tts
🗣 VoiceRead more →
Stylized voices in 4 categories (Professional / Formal / Casual / Energetic). 184 voices, 8 languages.
In: $0.0525/1M
tts
🗣 VoiceRead more →
Latest Mist generation, 83 voices, 4 languages. Improved expressiveness.
In: $0.0437/1M
tts
🗣 VoiceRead more →
Premium model for complex, multi-step work. Augure sovereign Canadian AI.
1.048576M ctx33K outIn: $0.0032/1MOut: $0.0076/1M
ToolsStreamchattools
💬 LLMRead more →
Sarvam's 24B multilingual LLM, fluent in all 22 official Indian languages + English. OpenAI-compatible API. Free tier with no per-token charges. Supports wiki grounding, reasoning effort control, and tool calls.
In: $0.000081/1MOut: $0.000325/1M
Toolstoolsreasoning
💬 LLMRead more →
Batch speech recognition with word-level timestamps and language detection across 99 languages.
$0.000170/unit
audio
📝 TranscriptionRead more →
New
Batch state-of-the-art speech recognition: 98%+ accuracy, keyterm prompting, 90+ languages.
$0.000107/unit
audio
📝 TranscriptionRead more →
New
Speech recognition fine-tuned for clinical audio: 35% fewer clinical errors than Scribe v2.
$0.000107/unit
audio
📝 TranscriptionRead more →
Ultra-low latency (<150ms) live speech recognition. 93.5% accuracy across 90+ languages. WebSocket streaming with VAD.
$0.000190/unit
audio
📝 TranscriptionRead more →
Highest-quality photo-realistic image generation perfect for professional print media
$0.1400/unit
🎨 ImageRead more →
8B parameter MMDiT model. Superior quality, typography, and prompt adherence. Most powerful in SD family.
$0.1138/unit
🎨 ImageRead more →
Distilled SD3.5 Large optimized for faster inference with fewer steps.
$0.0700/unit
🎨 ImageRead more →
Balanced SD3.5 model. Great quality with lower resource requirements than Large.
$0.0612/unit
🎨 ImageRead more →
⚡ CapabilityRead more →
ByteDance Seedance 1.5 Pro, text & image to video with native synced audio (voice, SFX, music), via BytePlus (Ark). Billed per output token; audio videos bill a higher rate.
Out: $0.0021/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0, text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0135/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0 Fast — speed-optimized text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0098/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0 Mini — cost-optimized text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0061/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.5, text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0187/1M
image
🎬 VideoRead more →
ByteDance Seedream 5.0 Lite, text-to-image (2K+), native via BytePlus (Ark).
$0.0612/unit
🎨 ImageRead more →
ByteDance Seedream 5.0 Pro, flagship text-to-image (2K+), highest quality + consistency. Native via BytePlus (Ark).
$0.1575/unit
🎨 ImageRead more →
New
Speechify's streaming-native multilingual model: English plus German, Spanish (Spain and Mexico), French, Italian and Brazilian Portuguese, and it accepts every Speechify voice whatever its locale.
In: $0.0175/1M
🗣 VoiceRead more →
New
Speechify's recommended streaming-native English model: lowest time-to-first-byte in the Simba family, fine-grained emotional control and SSML prosody. English voices only.
In: $0.0175/1M
🗣 VoiceRead more →
New
Upstage compact agentic MoE (35B total, 3B active). 524,288-token context (API-confirmed), tool calling and opt-in reasoning, English/Korean/Japanese.
524K ctx131K outIn: $0.000175/1MOut: $0.000700/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
Upstage nightly build of Solar Pro 2 with latest improvements.
66K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamchattools
💬 LLMRead more →
New
Upstage flagship agentic model. 512K context (API-confirmed family), tool calling, structured outputs and opt-in reasoning. English/Korean/Japanese.
524K ctx131K outIn: $0.000525/1MOut: $0.0021/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
Perplexity's lightweight search-augmented model, real-time web search with citations, fast responses, 128K context.
128K ctx16K outIn: $0.0018/1MOut: $0.0018/1M
Streamsearch
💬 LLMRead more →
Perplexity's most thorough research model, multi-step deep web investigation with comprehensive citations and reasoning.
128K ctx33K outIn: $0.0035/1MOut: $0.0140/1M
Streamsearchreasoning
💬 LLMRead more →
Perplexity's advanced search model, deeper web research, multi-step reasoning with citations, 200K context.
200K ctx16K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolssearch
💬 LLMRead more →
Perplexity's reasoning model, extended thinking with real-time web search, ideal for complex research and analysis.
128K ctx16K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolssearchreasoning
💬 LLMRead more →
New
Cartesia's most natural streaming text-to-speech (2026-08-27): pacing and intonation from context, 60+ emotions, speed and volume controls, 44 languages including Odia and Urdu, sub-100 ms first audio.
In: $0.0875/1M
🗣 VoiceRead more →
Claude's best model for complex agents and coding
200K ctx64K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Best Claude Sonnet for complex agents and coding (1M context)
1M ctx128K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's most capable Sonnet — closes the gap with Opus for complex agents and coding (1M context)
1M ctx128K outIn: $0.0035/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
New
Sonnet 5.5 by Anthropic (2026-09-28): the best combination of speed and intelligence, for everyday coding, documents and agents. 1M context, 128K output, adaptive thinking.
1M ctx128K outIn: $0.0035/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
New
Sound effects from a text description: automatic length or a set length up to 22 seconds, with adjustable prompt influence.
$0.0875/unit
🔊 Sound FXRead more →
Generate music and sound effects up to 3 minutes from text prompts. Produces structured compositions with intros, development, and outros at 44.1kHz stereo.
$0.3500/unit
🎵 MusicRead more →
Enterprise-grade music and sound generation. Produces structured compositions with intros, development, and outros at 44.1kHz stereo. 8-step inference for fast generation.
$0.3500/unit
🎵 MusicRead more →
Separate audio into individual stems (vocals, drums, bass, etc). 2-stem or 6-stem modes.
$0.0029/unit
audio
🎚 StemsRead more →
StepFun's latest reasoning model, 196B MoE (11B active), 256K context, open-source Apache 2.0. Fast agentic intelligence with strong code and math.
256K ctx16K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
StepFun's multimodal flagship reasoning model.
256K ctx16K outIn: $0.000350/1MOut: $0.0020/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Enhanced vocal quality and refined audio processing for music generation
$0.1050/unit
🎵 MusicRead more →
Excellent prompt understanding with faster generation speeds, supports up to 8 minute tracks
$0.1050/unit
🎵 MusicRead more →
Advanced model with enhanced tonal variation and excellent prompt understanding
$0.1050/unit
🎵 MusicRead more →
Cutting-edge model with enhanced quality and capabilities for AI music generation
$0.1050/unit
🎵 MusicRead more →
Upstage advanced synthetic reasoning model with 64K context.
66K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
Streamchat
💬 LLMRead more →
Studio-grade lip sync. Syncs video lip movements to match any audio input with natural speaker style preservation.
$0.0875/unit
videoaudio
🎬 VideoRead more →
Premium lip sync with enhanced detail preservation for beards, teeth, and fine facial features using diffusion-based super resolution.
$0.1452/unit
videoaudio
🎬 VideoRead more →
New
Light tier: fast, thinking off, 256K context, served in Canada with no failover outside Canada. Replaces Tofino 2.5, retired by Augure on 2026-09-19.
256K ctx33K outIn: $0.000756/1MOut: $0.0025/1M
Streamchat
💬 LLMRead more →
Tripo's P1 model for image-to-3D, converts a single image into a game-ready low-poly mesh (50 to 20k faces) with clean topology and PBR textures.
$0.8750/unit
image
📦 3DRead more →
Tripo's P1 model for multiview reconstruction, builds a game-ready low-poly mesh with clean topology from up to 4 view images (front required).
$0.8750/unit
image
📦 3DRead more →
Tripo's P1 model for text-to-3D, generates game-ready low-poly meshes (50 to 20k faces) with clean topology and PBR textures. Exports GLB, FBX, OBJ, USD, STL.
$0.7000/unit
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for image-to-3D, converts a single image into a native quad mesh (48 to 25k four-sided faces) ready for editing and subdivision, with PBR textures.
$2.19/unit
image
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for image-to-3D, converts a single image into a low-poly mesh (48 to 50k faces) with cleaner topology and PBR textures.
$2.19/unit
image
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for multiview reconstruction, builds a native quad mesh (48 to 25k four-sided faces) ready for editing and subdivision from up to 4 view images (front required).
$2.19/unit
image
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for multiview reconstruction, builds a low-poly mesh (48 to 50k faces) with cleaner topology from up to 4 view images (front required).
$2.19/unit
image
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for text-to-3D, generates native quad meshes (48 to 25k four-sided faces) ready for editing and subdivision, with PBR textures. Exports GLB, FBX, OBJ, USD, STL.
$2.33/unit
📦 3DRead more →
New
Tripo's next-gen P2.0 Preview model for text-to-3D, generates low-poly meshes (48 to 50k faces) with cleaner topology and PBR textures. Exports GLB, FBX, OBJ, USD, STL.
$2.33/unit
📦 3DRead more →
Splits an existing 3D model into separately editable parts, textures and PBR materials preserved. A practical level of separation for most assets.
$0.7000/unit
3d
📦 3DRead more →
Splits an existing 3D model into many separately editable parts, textures and PBR materials preserved. Maximum component separation for deep editing and production.
$0.7000/unit
3d
📦 3DRead more →
Splits an existing 3D model into a few separately editable parts, textures and PBR materials preserved. Quick separation for review and lightweight prep.
$0.7000/unit
3d
📦 3DRead more →
Tripo's balanced v2.5 model for converting a single image into a 3D mesh, extracts geometry, texture, and materials from a photo or illustration.
$0.5250/unit
image
📦 3DRead more →
Tripo v2.5 multiview generation, reconstructs detailed meshes from multiple image perspectives for improved accuracy.
$0.5250/unit
image
📦 3DRead more →
Tripo's balanced v2.5 model for generating 3D meshes from text descriptions, produces detailed geometry with PBR materials in seconds. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's stable v3.0 model for image-to-3D, reconstructs clean geometry, textures, and PBR materials from a single image, up to 2M polygons in Ultra mode.
$0.5250/unit
image
📦 3DRead more →
Tripo v3.0 multiview generation, reconstructs detailed meshes from up to 4 view images (front required) with stable, production-proven quality.
$0.5250/unit
image
📦 3DRead more →
Tripo's stable v3.0 model for text-to-3D, crisp edges and coherent structure with PBR materials, up to 2M polygons in Ultra mode. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's flagship v3.1 model for image-to-3D, reconstructs high-fidelity geometry, textures, and PBR materials from a single image, up to 2M polygons in Ultra mode.
$0.5250/unit
image
📦 3DRead more →
Tripo's flagship v3.1 image-to-3D with 8K Ultra textures, reconstructs maximum-fidelity geometry and materials from a single image for hero assets.
$0.8750/unit
image
📦 3DRead more →
Tripo v3.1 image-to-3D in part-segmentation mode, converts a single image into an untextured mesh split into separately editable parts.
$0.7000/unit
image
📦 3DRead more →
Tripo's highest-fidelity 3D generation, v3.1 reconstructs detailed meshes from up to 4 view images (front required) for maximum accuracy.
$0.5250/unit
image
📦 3DRead more →
Tripo's highest-fidelity pipeline, v3.1 multiview reconstruction with 8K Ultra textures from up to 4 view images (front required).
$0.8750/unit
image
📦 3DRead more →
Tripo v3.1 multiview reconstruction in part-segmentation mode, builds an untextured segmented mesh from up to 4 view images (front required).
$0.7000/unit
image
📦 3DRead more →
Tripo's flagship v3.1 model for text-to-3D, sculpture-level geometry with crisp edges and PBR materials, up to 2M polygons in Ultra mode. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's flagship v3.1 text-to-3D with 8K Ultra textures, maximum-fidelity materials and fine surface detail for hero assets and close-ups. Exports GLB, FBX, OBJ, USD, STL.
$0.7000/unit
📦 3DRead more →
Tripo v3.1 text-to-3D in part-segmentation mode, generates an untextured mesh split into separately editable parts for editing, rigging prep, and 3D printing. Exports GLB, FBX, OBJ, USD, STL.
$0.5250/unit
📦 3DRead more →
Turns a single image into a photorealistic 3D Gaussian Splat in under a minute. Ideal for AR/VR, web scenes and visualization — a view-ready format, not an editable or printable mesh. Exports SPLAT.
$0.5250/unit
image
📦 3DRead more →
Low-latency speech generation in 32 languages, optimized for real-time conversational AI
In: $0.1925/1M
🗣 VoiceRead more →
The most creative, intelligent and personalizable image generation models built on a new groundbreaking architecture that delivers ultra high quality and 10x higher cost efficiency.
$0.0707/unit
🎨 ImageRead more →
The most creative, intelligent and personalizable image generation models built on a new groundbreaking architecture that delivers ultra high quality and 10x higher cost efficiency.
$0.1750/unit
🎨 ImageRead more →
Latest Veo with video extending capability. Create and extend AI-generated videos with improved consistency.
$0.7000/unit
imagevideo
🎬 VideoRead more →
Fast variant of Veo 3.1. Quick video generation and extending with good quality.
$0.1750/unit
imagevideo
🎬 VideoRead more →
Cost-efficient Veo 3.1 tier — video with audio at a fraction of the price (720p/1080p, no 4K).
$0.0875/unit
image
🎬 VideoRead more →
Generate and edit sound effects from video (SFX 1.6). Provide a video URL and optional text prompt; adds seamless extension, looping ambiences, and AI inpainting to erase/replace moments.
$0.0875/unit
video
🔊 Sound FXRead more →
Reference-based video generation using multiple images for visual consistency. 4-second 720p output.
$0.7000/unit
image
🎬 VideoRead more →
Create a talking avatar from an image or video with AI-generated speech from text
$0.8750/unit
🎬 VideoRead more →
Synchronize video lips to an audio file for realistic dubbing and voice replacement
$0.7000/unit
🎬 VideoRead more →
High-quality text-to-video with anime style support. 5-second 1080p output.
$0.7000/unit
🎬 VideoRead more →
Generate images from 1-7 reference images. Upload reference images to get started.
$0.0700/unit
🎨 ImageRead more →
Generate images from text or reference images. Supports 1080p, 2K, and 4K resolution.
$0.0525/unit
🎨 ImageRead more →
Ultra-fast video from text or image. Upload an image for image-to-video mode.
$0.3325/unit
🎬 VideoRead more →
New
Transform a recording into another voice, keeping the words, timing and emotion. English only.
$0.0035/unit
audio
voice_changeRead more →
New
Transform a recording into another voice, keeping the words, timing and emotion. Multilingual.
$0.0035/unit
audio
voice_changeRead more →
Voyage AI's balanced general-purpose embedding model, strong quality at moderate cost. 200M free tokens.
In: $0.000105/1M
🔍 EmbeddingRead more →
Voyage AI's most capable general-purpose embedding model, highest quality retrieval and semantic search. 200M free tokens.
In: $0.000210/1M
🔍 EmbeddingRead more →
Voyage AI's fastest and cheapest embedding model, ideal for high-volume, low-latency use cases. 200M free tokens.
In: $0.000035/1M
🔍 EmbeddingRead more →
Voyage AI's code-optimized embedding model, best-in-class for code search, retrieval, and similarity. 200M free tokens.
In: $0.000315/1M
🔍 EmbeddingRead more →
Voyage AI's long-context embedding model, optimized for large documents and RAG with extended context windows. 200M free tokens.
In: $0.000315/1M
🔍 EmbeddingRead more →
Voyage AI's multimodal embedding model, embeds both text and images into a shared vector space for cross-modal search.
In: $0.000210/1M
image
🔍 EmbeddingRead more →
Alibaba Wan 2.2 Animate Mix, composite a character into a reference video (native DashScope, Singapore).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.2 Animate Move, animate a character image with a reference video's motion (native DashScope, Singapore).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.6 image editing. Edits 1-4 input images from a natural-language instruction (native DashScope, US-Virginia).
$0.0525/unit
image
🎨 ImageRead more →
Alibaba Wan 2.6 image-to-video. Animates a still image into video while preserving subject, style and detail (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba Wan 2.6 text-to-image. Photorealistic generation with accurate text rendering and flexible artistic styles (native DashScope, US-Virginia).
$0.0502/unit
🎨 ImageRead more →
Alibaba Wan 2.6 reference-to-video. Generates video preserving the look (and voice) of subjects from a reference video and/or reference images (native DashScope, US-Virginia).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.6 text-to-video. Cinematic motion generation from text with strong instruction following (native DashScope, US-Virginia).
video
🎬 VideoRead more →
Alibaba Wan 2.7 image-to-video. Animates a still image into video with first-frame fidelity, clips up to 15 seconds (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba Wan 2.7 text-to-image and editing. Served from Singapore — the US tenant returns AccessDenied.
$0.0481/unit
🎨 ImageRead more →
Alibaba Wan 2.7 pro tier, up to 4K output for print and large-format work. Served from Singapore.
$0.1203/unit
🎨 ImageRead more →
Alibaba Wan 2.7 text-to-video. First and last frame control, clips up to 15 seconds, sharper motion and instruction following (native DashScope, US-Virginia).
video
🎬 VideoRead more →