AI Model Directory
Every model. Every provider. Live data, real pricing, honest capabilities. Updated daily from our API monitor.
Live data — updated daily
AI performance capture. Upload a character image + your performance video to transfer your motion to the character.
$0.0875/unit
videoimage
🎬 VideoRead more →
New
Runway's current video-to-video model. Edit existing footage with text prompts; accepts 2-30 second inputs.
$0.4900/unit
video
🎬 VideoRead more →
Ultra-fast real-time text-to-speech optimized for low-latency applications
In: $0.0192/1M
🗣 VoiceRead more →
Multilingual speech synthesis supporting 15 languages including Arabic, Japanese, and Chinese
In: $0.0192/1M
🗣 VoiceRead more →
High-quality text-to-speech with natural prosody and emotional expression across 15 languages
In: $0.0192/1M
🗣 VoiceRead more →
Deepgram's original text-to-speech model, 12 English voices at half the cost of Aura 2. Fast and reliable.
In: $0.0262/1M
Stream
🗣 VoiceRead more →
Deepgram's latest text-to-speech model, 90+ natural voices across 8 languages (EN, FR, DE, ES, IT, NL, JA), sub-200ms latency. Greek mythology-themed voice names.
In: $0.0525/1M
Stream
🗣 VoiceRead more →
Sarvam's stable text-to-speech model for Indian languages and English. Returns base64-encoded WAV audio. Production-ready with consistent quality.
In: $0.000315/1M
🗣 VoiceRead more →
Sarvam's latest text-to-speech model, 30+ voices across Indian languages and English. Natural prosody with cultural intonation. Currently in beta.
In: $0.000630/1M
🗣 VoiceRead more →
ByteDance Seed 1.8, multimodal agent model (text, image & video in → text), strong tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Code (preview), coding-specialized agent (text & image in → text), strong tool use + reasoning. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
ByteDance Seed 2.0 Lite, efficient multimodal agent (text, image & video in → text), tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Mini, fast lightweight multimodal agent (text, image & video in → text), tool use. Native via BytePlus (Ark).
131K ctx33K outIn: $0.000175/1MOut: $0.000700/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.0 Pro, flagship multimodal agent (text, image & video in → text), strong tool use + reasoning, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
ByteDance Seed 2.1 Turbo, fast multimodal reasoning agent (text & image in → text), strong tool use + thinking, 256K context. Native via BytePlus (Ark).
262K ctx66K outIn: $0.000875/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Specialized model for code generation, completion, and understanding.
262K ctx8K outIn: $0.000525/1MOut: $0.0016/1M
Streamchat
💬 LLMRead more →
SOTA open-source image model. Native Chinese text support, 6B parameters.
$0.0175/unit
🎨 ImageRead more →
High-quality neural machine translation supporting 33 languages
In: $0.0437/1M
🌐 TranslationRead more →
DeepSeek V4 Flash, fast, cost-effective daily driver. 1M context, thinking + non-thinking modes.
1M ctx66K outIn: $0.000385/1MOut: $0.0012/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
DeepSeek flagship V4. 512K context, advanced reasoning + tool use.
524K ctx66K outIn: $0.0012/1MOut: $0.0035/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
Most expressive TTS model with audio tags, dialogue mode, accent emulation, and 70+ languages. Best for long-form content.
In: $0.3850/1M
🗣 VoiceRead more →
Fable 5 by Anthropic.
1M ctx128K outIn: $0.0175/1MOut: $0.0875/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Ultra-low latency TTS for real-time and conversational AI. ~75ms latency.
In: $0.1925/1M
🗣 VoiceRead more →
The best of FLUX, offering state-of-the-art performance image generation at blazing fast speeds
$0.0700/unit
🎨 ImageRead more →
A premium model brings maximum performance across all aspects – greatly improved quality, consistency, and speed
$0.1400/unit
🎨 ImageRead more →
A unified model delivering local editing, generative modifications, and text-to-image capabilities
$0.0700/unit
🎨 ImageRead more →
Directly distilled from FLUX.1 [pro], an open-weight, guidance-distilled model for non-commercial use
$0.0437/unit
🎨 ImageRead more →
State-of-the-art performance image generation with top of the line prompt following and visual quality
$0.0875/unit
🎨 ImageRead more →
Specialized for typography and text rendering with adjustable guidance and generation steps
$0.0875/unit
🎨 ImageRead more →
Ultra-fast open-source FLUX model optimized for real-time generation at the lowest cost. Apache 2.0 licensed.
$0.0245/unit
🎨 ImageRead more →
Fastest FLUX model with sub-second inference. 9B params with Qwen3 text embedder for superior prompt understanding.
$0.0262/unit
🎨 ImageRead more →
Highest quality FLUX 2.0 with grounding search capability for photorealistic images up to 4MP
$0.1400/unit
🎨 ImageRead more →
Production-grade FLUX 2.0 with multi-reference image editing and precise control over colors, poses, and composition
$0.0525/unit
🎨 ImageRead more →
Sakana's multi-agent conductor, orchestrates a pool of frontier models for complex reasoning. Billed per token, orchestration included.
272K ctx16K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Fast and efficient model with adaptive thinking for complex tasks
1M ctx66K outIn: $0.000525/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Google's low-latency text-to-speech with natural prosody and controllable style. Supports 24 languages.
In: $0.0350/1M
🗣 VoiceRead more →
Cost-efficient high-throughput model for budget-conscious applications
1M ctx66K outIn: $0.000175/1MOut: $0.000700/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Enhanced thinking and reasoning for complex problems
1M ctx66K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Premium text-to-speech with enhanced expressivity, richer tone, and precision pacing. Best for high-quality output.
In: $0.0700/1M
🗣 VoiceRead more →
Google's fastest frontier model. Beats 2.5 Pro at 1/4 the cost with 1M context
1M ctx66K outIn: $0.000875/1MOut: $0.0053/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Latest price-performant, low-latency controllable speech generation. 30 voices, 24 languages.
In: $0.0700/1M
🗣 VoiceRead more →
Cost-efficient high-throughput model for budget-conscious applications.
1M ctx66K outIn: $0.000438/1MOut: $0.0026/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Google's most capable agentic model with 1M token context, 77.1% ARC-AGI-2 reasoning, and native tool use.
1M ctx66K outIn: $0.0035/1MOut: $0.0210/1M
ToolsStreamtoolsreasoningimage+2
💬 LLMRead more →
Fast flagship with adaptive thinking for complex tasks.
1M ctx66K outIn: $0.0026/1MOut: $0.0158/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Cost-efficient GA Flash model for high-throughput multimodal tasks.
1M ctx66K outIn: $0.000525/1MOut: $0.0044/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Google's latest GA Flash model for fast multimodal reasoning and agentic workloads.
1M ctx66K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
New
Google's most capable Flash model for agentic workflows and multimodal reasoning.
1.048576M ctx66K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Cost-efficient text embeddings — 3072-dim vectors for search, clustering, and RAG.
In: $0.000262/1M
🔍 EmbeddingRead more →
Google's latest embedding model — 3072-dim vectors for retrieval and semantic search (multimodal upstream; text wired).
In: $0.000350/1M
🔍 EmbeddingRead more →
Google's open-weight mixture-of-experts model, 26B total, 4B active parameters for fast inference
262K ctx33K outIn: $0.000105/1MOut: $0.000577/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Google's open-weight dense model with 262K context and strong multilingual reasoning
262K ctx33K outIn: $0.000245/1MOut: $0.000700/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Transform reference images with text prompts. Preserves identity while changing pose, lighting, background.
$0.0875/unit
image
🎨 ImageRead more →
Fast reference-based image generation. 2.5x faster, preserves identity and style.
$0.0350/unit
image
🎨 ImageRead more →
Runway's flagship image-to-video model. Exceptional fidelity, stability, and controllability. Requires input image.
$0.0875/unit
image
🎬 VideoRead more →
Runway's most capable video model. #1 on Artificial Analysis text-to-video benchmark. Text-to-video or image-to-video with exceptional physics and human motion.
$0.2100/unit
image
🎬 VideoRead more →
Previous flagship with thinking mode. MoE architecture, 128K context.
128K ctx8K outIn: $0.0010/1MOut: $0.0039/1M
Streamreasoning
💬 LLMRead more →
Lightweight model optimized for efficiency. 128K context.
128K ctx8K outIn: $0.000350/1MOut: $0.0019/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Free tier model. Great for testing and light workloads.
128K ctx4K out
ToolsStreamtoolsreasoning
💬 LLMRead more →
Vision-language model for image understanding and analysis.
66K ctx16K outIn: $0.0010/1MOut: $0.0032/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Flagship model with reasoning, coding, and agentic capabilities. 128K context.
128K ctx8K outIn: $0.0010/1MOut: $0.0039/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Vision-capable model for image understanding and analysis.
128K ctx8K outIn: $0.000525/1MOut: $0.0016/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Lightweight vision model. Fast and cost-effective for image tasks.
128K ctx4K outIn: $0.000070/1MOut: $0.000700/1M
ToolsStreamreasoningtoolsimage
💬 LLMRead more →
Latest flagship. 358B params, 204K context, 131K output. #1 on LiveCodeBench.
205K ctx131K outIn: $0.0010/1MOut: $0.0039/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Lightweight, completely free GLM-4.7 variant with 200K context.
205K ctx131K out
ToolsStreamtoolsreasoning
💬 LLMRead more →
Lightweight, high-speed and affordable GLM-4.7 variant with 200K context.
205K ctx131K outIn: $0.000122/1MOut: $0.000700/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Most capable Z.ai model. 744B params (40B active MoE), 28.5T training tokens. Built for complex systems engineering and agentic tasks.
205K ctx131K outIn: $0.0018/1MOut: $0.0056/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Fast variant of GLM-5 with optimized speed and competitive quality.
In: $0.0021/1MOut: $0.0070/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's current flagship for agentic and coding tasks.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Z.ai's latest flagship for agentic and coding tasks.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
New
Z.ai's newest flagship for agentic and coding tasks. Always reasons; text-only input.
205K ctx131K outIn: $0.0024/1MOut: $0.0077/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
New
Z.ai's fast, low-cost GLM-5.3 variant: 1M context, tools, always reasons; text and image input.
1M ctx131K outIn: $0.000262/1MOut: $0.000875/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
Native multimodal coding model. 744B MoE (40B active), 203K context. Optimized for design-to-code, GUI automation, and vision-grounded agentic tasks.
205K ctx131K outIn: $0.0021/1MOut: $0.0070/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Z.ai's image generation model with strong prompt adherence and text rendering.
$0.0262/unit
🎨 ImageRead more →
Specialized OCR model for extracting text from images and documents.
8K ctx8K outIn: $0.000053/1MOut: $0.000053/1M
Streamimage
ocrRead more →
OpenAI's latest image generation model (high-quality tier).
$0.3693/unit
imageimage_editimage
🎨 ImageRead more →
OpenAI GPT-3.5 Turbo (chat completions endpoint).
In: $0.000875/1MOut: $0.0026/1M
tools
💬 LLMRead more →
OpenAI GPT-3.5 Turbo 16K (chat completions endpoint).
In: $0.0053/1MOut: $0.0070/1M
tools
💬 LLMRead more →
OpenAI GPT-4 Turbo (chat completions endpoint).
In: $0.0175/1MOut: $0.0525/1M
toolsvisionimage
💬 LLMRead more →
OpenAI GPT-4.1 (chat completions endpoint).
1.047576M ctx33K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4.1 Mini (chat completions endpoint).
1.047576M ctx33K outIn: $0.000700/1MOut: $0.0028/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4o (chat completions endpoint).
128K ctx16K outIn: $0.0044/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-4o Mini (chat completions endpoint).
128K ctx16K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Mini (chat completions endpoint).
272K ctx128K outIn: $0.000438/1MOut: $0.0035/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Nano (chat completions endpoint).
272K ctx128K outIn: $0.000087/1MOut: $0.000700/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0262/1MOut: $0.2100/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5 Search API (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
Streamweb_searchimage
💬 LLMRead more →
OpenAI GPT-5.1 (chat completions endpoint).
272K ctx128K outIn: $0.0022/1MOut: $0.0175/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.2 (chat completions endpoint).
272K ctx128K outIn: $0.0031/1MOut: $0.0245/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.2 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0367/1MOut: $0.2940/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5.3 Codex (chat completions endpoint).
272K ctx128K outIn: $0.0031/1MOut: $0.0245/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI mid-tier flagship (March 2026).
272K ctx128K outIn: $0.0044/1MOut: $0.0262/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI fast/cheap chat model.
272K ctx128K outIn: $0.0013/1MOut: $0.0079/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI cheapest tier, very fast.
272K ctx128K outIn: $0.000350/1MOut: $0.0022/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI GPT-5.4 Pro (chat completions endpoint).
272K ctx128K outIn: $0.0525/1MOut: $0.3150/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI flagship chat model (April 2026 release).
272K ctx128K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsimage
💬 LLMRead more →
OpenAI premium tier with deep reasoning. April 2026.
272K ctx128K outIn: $0.0525/1MOut: $0.3150/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI GPT-5.6 Luna — high-volume streaming tasks, fast + economical 5.6 tier.
272K ctx128K outIn: $0.0018/1MOut: $0.0105/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
OpenAI GPT-5.6 Sol — frontier logic + deep reasoning, the flagship 5.6 tier.
272K ctx128K outIn: $0.0088/1MOut: $0.0525/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
OpenAI GPT-5.6 Terra — balanced business logic, mid 5.6 tier.
272K ctx128K outIn: $0.0044/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningvision+1
💬 LLMRead more →
Non-reasoning variant of Grok 4.20. Fastest 4.20 path for chat and high-throughput tasks.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Reasoning-mode variant of Grok 4.20. Slower but stronger on multi-step problems.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchatreasoningtools+1
💬 LLMRead more →
xAI's flagship general-purpose model. 1M context, balanced reasoning and chat.
1M ctx8K outIn: $0.0022/1MOut: $0.0044/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
xAI's newest flagship, launched July 2026. 500K context, strongest reasoning + coding, trained alongside Cursor.
500K ctx8K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
New
xAI's newest flagship. 500K context, accepts text and images, strongest reasoning + coding.
500K ctx8K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamchattoolsreasoning+1
💬 LLMRead more →
Compact model tuned for code generation and structured output. 256K context.
256K ctx8K outIn: $0.0018/1MOut: $0.0035/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Fast text-to-image generation from xAI. Standard quality, lowest cost.
$0.0350/unit
🎨 ImageRead more →
New
xAI's current image model. Higher fidelity than Grok Imagine, with stronger prompt adherence.
$0.1050/unit
🎨 ImageRead more →
Higher-fidelity Grok Imagine. Better prompt adherence and detail, 2.5x the cost.
$0.0875/unit
🎨 ImageRead more →
xAI Grok Imagine 1.5 — image-to-video: animates a starting frame from your prompt. Up to 1080p, 1-15s clips, async render.
image
🎬 VideoRead more →
xAI's expressive text-to-speech — 26 multilingual voices with inline speech tags for tone, pauses, whispers and laughter. 20+ languages with auto-detection.
In: $0.0262/1M
🗣 VoiceRead more →
Claude's fastest and most intelligent Haiku model
200K ctx64K outIn: $0.0018/1MOut: $0.0088/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Higher resolution (1080p), longer duration (10s), stronger prompt adherence
$0.4900/unit
image
🎬 VideoRead more →
Breakthroughs in body movement, facial expressions, and physical realism
$0.4900/unit
image
🎬 VideoRead more →
Image-to-video model optimized for value and efficiency (requires image upload)
$0.3325/unit
image
🎬 VideoRead more →
Arabic-first speech-to-text with dialect-aware recognition and broad multilingual coverage.
$0.000292/unit
audio
📝 TranscriptionRead more →
Arabic-first text-to-speech, low-latency tier, with 50 voices spanning Arabic dialects and 20+ languages.
In: $0.0700/1M
🗣 VoiceRead more →
Alibaba HappyHorse image-to-video. Animates a still image into a coherent motion sequence guided by text.
$0.2450/unit
videoimage
🎬 VideoRead more →
Alibaba HappyHorse reference-to-video. Generates video conditioned on a reference image plus text prompt.
$0.2450/unit
videoimage
🎬 VideoRead more →
Alibaba HappyHorse text-to-video. Generates short cinematic clips from text prompts.
$0.2450/unit
video
🎬 VideoRead more →
Alibaba HappyHorse 1.1 image-to-video. Animates a still image into coherent motion guided by text (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba HappyHorse 1.1 reference-to-video. Generates video that preserves the subject and scene from one or more reference images (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba HappyHorse 1.1 text-to-video. Generates short cinematic clips from text prompts (native DashScope, US-Virginia).
video
🎬 VideoRead more →
Alibaba HappyHorse video editing. Edits a source video from natural-language instructions, optionally guided by reference images (native DashScope, US-Virginia).
videovideoimage
🎬 VideoRead more →
Previous generation model. Still excellent for text rendering at lower cost.
$0.1050/unit
🎨 ImageRead more →
State-of-the-art image generation with exceptional text rendering. Best balance of quality and speed.
$0.1225/unit
🎨 ImageRead more →
Highest quality Ideogram 3.0. Maximum detail and fidelity for final outputs.
$0.1750/unit
🎨 ImageRead more →
Faster Ideogram 3.0 variant. Quick generation for iterative workflows.
$0.0700/unit
🎨 ImageRead more →
Jina AI's CLIP model, cross-modal text-image embeddings supporting 89 languages with Matryoshka representations. 865M params.
8K ctxIn: $0.000087/1M
image
🔍 EmbeddingRead more →
Jina AI's code-specialized embedding model, 1.5B parameters optimized for code search and retrieval. 32K context.
33K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's ColBERT late-interaction model, multi-vector embeddings for fine-grained retrieval, 89 languages, user-controlled embedding sizes.
8K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's 570M parameter multilingual text embedding model, MTEB benchmark champion, 8K context, 89 languages.
8K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Jina AI's latest multimodal embedding model, embeds text and images into a shared vector space with 32K context. State-of-the-art retrieval quality.
33K ctxIn: $0.000087/1M
image
🔍 EmbeddingRead more →
Jina AI's multimodal reranker, re-ranks results containing both text and images for cross-modal search pipelines.
10K ctxIn: $0.000087/1M
image
↕️ RerankRead more →
Jina AI's latest reranker, novel listwise architecture, SOTA multilingual retrieval, massive 131K context window. 0.6B params.
134K ctxIn: $0.000087/1M
↕️ RerankRead more →
Latest Kimi K2 reasoning model. Replaces K2 thinking variants (k2-thinking and k2-thinking-turbo) being retired 2026-05-25.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
Coding-specialized Kimi K2 model (k2.7). 256K context, agentic tool use and function calling, tuned for software engineering.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
High-speed variant of Kimi K2.7 Code for low-latency agentic software engineering.
262K ctx131K outIn: $0.0017/1MOut: $0.0070/1M
ToolsStreamchatreasoningfunction_calling+1
💬 LLMRead more →
Moonshot's open-weight 2.8T-parameter multimodal reasoning model with a 1M-token context window.
1.048576M ctx131K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamchatreasoningfunction_calling+2
💬 LLMRead more →
High-speed anime-focused model excelling at a range of anime styles.
$0.0280/unit
generation
🎨 ImageRead more →
Core Leonardo model with stunning outputs and broad style range.
$0.0280/unit
generation
🎨 ImageRead more →
Cinematic-focused model excelling at film-like compositions and lighting.
$0.0280/unit
generation
🎨 ImageRead more →
High-speed generalist image generation model for fast iteration.
$0.0140/unit
generation
🎨 ImageRead more →
Generate sound effects for longer videos (SFX 1.6), up to 60 seconds of audio from a video URL.
$0.1750/unit
video
🔊 Sound FXRead more →
Leonardo high-speed model designed for photorealistic outputs.
$0.0280/unit
generation
🎨 ImageRead more →
Google DeepMind's latest music generation — 30-second compositions from text prompts with SynthID watermarking.
$0.0700/unit
🎵 MusicRead more →
Full-song generation from Google DeepMind's Lyria 3 — complete compositions from text prompts with SynthID watermarking.
$0.1400/unit
🎵 MusicRead more →
Agentic capabilities with advanced reasoning and thinking process visibility
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Optimized for high concurrency and commercial use with advanced reasoning
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Frontier reasoning model with SOTA coding, agentic tool use, and complex real-world task performance
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
Fast variant of M2.5 optimized for speed at ~100 tokens per second
In: $0.0010/1MOut: $0.0042/1M
ToolsStream
💬 LLMRead more →
Autonomous real-world productivity with agentic collaboration, live debugging, and professional document generation
In: $0.000525/1MOut: $0.0021/1M
ToolsStream
💬 LLMRead more →
New
Fast variant of M2.7 for latency-sensitive work. 200K context, text only.
200K ctx128K outIn: $0.0010/1MOut: $0.0042/1M
ToolsStreamtools
💬 LLMRead more →
Efficient reasoning model for everyday logic tasks.
131K ctx40K outIn: $0.000875/1MOut: $0.0026/1M
Streamchatreasoning
💬 LLMRead more →
Twelve Labs multimodal video embedding model. Converts video, audio, and text into a shared vector space for semantic search across 36 languages.
$0.000963/unit
embeddingsearchvideoaudioimage
🔍 EmbeddingRead more →
Transform images into detailed 3D models with textures and PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
Generate detailed 3D models from text descriptions with PBR textures
$1.05/unit
3d-generation
📦 3DRead more →
New
Transform images into detailed 3D models with the highest image alignment Meshy has shipped, textures and PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
New
Combine two to four views of the same subject into one aligned, textured 3D model with PBR maps
$1.05/unit
3d-generationimage
📦 3DRead more →
High-performance edge model with vision. Best Ministral for complex tasks.
131K ctx8K outIn: $0.000350/1MOut: $0.000350/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Ultra-lightweight edge model with vision. 3B params, runs on phones/laptops.
131K ctx8K outIn: $0.000175/1MOut: $0.000175/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Compact edge model with vision. 8B params for local deployment.
131K ctx8K outIn: $0.000262/1MOut: $0.000262/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Mistral's text embedding model for semantic search, retrieval, and classification.
In: $0.000175/1M
🔍 EmbeddingRead more →
Open-weight flagship. 675B MoE (41B active), multimodal with vision. Apache 2.0 licensed.
131K ctx8K outIn: $0.000875/1MOut: $0.0026/1M
ToolsStreamchattoolsimage
💬 LLMRead more →
Premier frontier-class multimodal model. Great balance between Large and Small.
131K ctx8K outIn: $0.0026/1MOut: $0.0131/1M
ToolsStreamchattoolsreasoning
💬 LLMRead more →
Mistral's content moderation model for detecting harmful, unsafe, or policy-violating content.
In: $0.000175/1M
⚡ CapabilityRead more →
Document AI model for high-accuracy OCR. Extracts text from images, PDFs, handwritten content, forms, and complex tables.
In: $0.0018/1MOut: $0.0053/1M
image
ocrRead more →
Fast and efficient model for simple tasks. Great balance of performance and cost.
131K ctx8K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamchattools
💬 LLMRead more →
Mistral's model optimized for command-line and terminal-based coding workflows.
128K ctx33K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamtools
💬 LLMRead more →
Most advanced emotionally-aware speech synthesis with rich expression across 29 languages
In: $0.3850/1M
🗣 VoiceRead more →
Gemini 2.5 Flash Image, optimized for speed and efficiency, high-volume low-latency image generation.
66K ctx33K out$0.0683/unit
🎨 ImageRead more →
High-efficiency image generation, optimized for speed and high-volume use.
66K ctx33K out$0.1172/unit
image
🎨 ImageRead more →
Efficiency specialist of the Nano Banana family — ultra-low-latency, cost-effective image generation and editing.
66K ctx33K out$0.0588/unit
image
🎨 ImageRead more →
Google's native image generation with text rendering and multimodal understanding.
66K ctx33K out$0.2345/unit
image
🎨 ImageRead more →
OpenAI o3 (chat completions endpoint).
200K ctx100K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
First-generation empathic speech synthesis with natural emotional expression
In: $0.2625/1M
🗣 VoiceRead more →
Latest empathic text-to-speech model supporting 11 languages with emotional awareness
In: $0.1313/1M
🗣 VoiceRead more →
Premium model combining maximum intelligence with practical performance
200K ctx64K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Most capable Claude model with enhanced reasoning and coding
1M ctx64K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Highly capable Claude Opus for complex reasoning and coding (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's most capable model for complex reasoning and agentic coding (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's flagship Opus — deep reasoning, agentic coding and long-horizon work (1M context)
1M ctx128K outIn: $0.0088/1MOut: $0.0437/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Large-context generalist with 128K window. Excellent for complex generation and analysis. Augure sovereign Canadian AI.
131K ctx33K outIn: $0.0013/1MOut: $0.0038/1M
ToolsStreamchattools
💬 LLMRead more →
Premium multimodal model with vision and tool use, strong bilingual EN/FR. Augure sovereign Canadian AI.
33K ctx33K outIn: $0.0032/1MOut: $0.0101/1M
ToolsStreamchattoolsvision+1
💬 LLMRead more →
Leonardo foundational model preview with extreme prompt adherence and versatility.
$0.0280/unit
generation
🎨 ImageRead more →
Leonardo flagship foundational model with exceptional prompt adherence and text rendering.
$0.0280/unit
generation
🎨 ImageRead more →
PixVerse C1, cinematic-quality video generation tuned for action scenes and reference-based work, clips from 1 to 15 seconds with optional generated audio.
video
🎬 VideoRead more →
PixVerse C1 image-to-video: cinematic animation of a source image into a 1–15 second clip with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse C1 fusion: cinematic video from 1–3 reference images (subjects or scenes) guided by a prompt, with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse V6, general-purpose video generation with clips from 1 to 15 seconds, optional generated audio, and multi-clip support.
video
🎬 VideoRead more →
PixVerse V6 image-to-video: animate a source image into a 1–15 second clip with optional generated audio.
videoimage
🎬 VideoRead more →
PixVerse V6 fusion: generate a video from 1–3 reference images (subjects or scenes) guided by a prompt, with optional generated audio.
videoimage
🎬 VideoRead more →
Perplexity's lightweight embedding model. 1024-dimensional INT8 vectors, 32K context, Matryoshka dimension reduction.
33K ctxIn: $0.000087/1M
🔍 EmbeddingRead more →
Perplexity's most capable embedding model. 2560-dimensional INT8 vectors, 32K context, Matryoshka dimension reduction.
33K ctxIn: $0.000175/1M
🔍 EmbeddingRead more →
Pruna's optimized image editing: transform an input image with a prompt.
$0.0175/unit
image
🎨 ImageRead more →
Pruna's premium video generation: text-to-video and image-to-video, up to 1080p. Draft mode gives fast, cheap previews.
videoimage
🎬 VideoRead more →
Ultra-fast model with 1M context. Best latency and cost efficiency for simple tasks.
1M ctx8K outIn: $0.000039/1MOut: $0.000378/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Alibaba's flagship text-to-image model — successor to the whole Qwen-Image family, with strong prompt following and text rendering (native DashScope, Singapore).
$0.0612/unit
🎨 ImageRead more →
The higher-fidelity "pro" tier of Alibaba's Qwen Image 2.0 — stronger prompt following and text rendering (native DashScope, Singapore).
$0.1313/unit
🎨 ImageRead more →
Balanced flagship model with 1M context. Best performance/cost ratio for most tasks.
1M ctx8K outIn: $0.000700/1MOut: $0.0021/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Cost-effective coding model optimized for speed. Fast code generation and completion at lower cost.
131K ctx8K outIn: $0.000252/1MOut: $0.0010/1M
ToolsStreamtools
💬 LLMRead more →
Advanced vision-language model with 262K context. Excels at visual coding, spatial perception, and multimodal reasoning.
262K ctx8K outIn: $0.000250/1MOut: $0.0025/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Flagship 1T+ parameter model. State-of-the-art reasoning, coding, and agent capabilities.
262K ctx33K outIn: $0.000051/1MOut: $0.000502/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Alibaba Qwen3.5 Omni Flash, fast multimodal (text & image input, text output). Singapore deployment; native voice/audio coming.
66K ctx8K outIn: $0.000700/1MOut: $0.0039/1M
Streamimage
💬 LLMRead more →
Alibaba Qwen3.5 Omni, multimodal (text & image input, text output). Singapore deployment; native voice/audio coming.
66K ctx8K outIn: $0.0024/1MOut: $0.0145/1M
Streamimage
💬 LLMRead more →
397B MoE model with 17B active parameters, 1M token context, multimodal (text/image/video)
1M ctx8K outIn: $0.000201/1MOut: $0.0012/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Flagship 1T+ parameter model. State-of-the-art reasoning, coding, and agent capabilities.
262K ctx33K outIn: $0.000289/1MOut: $0.0017/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Qwen3.6 Plus, next-gen frontier LLM (Agentic Coding, vision/OCR, fine-grained localization). Replaces Qwen3.5 Plus.
1M ctx8K outIn: $0.000483/1MOut: $0.0029/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
Alibaba's Agent Frontier flagship.
262K ctx33K outIn: $0.0029/1MOut: $0.0087/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Multimodal frontier (vision/video); cheaper than Max.
1M ctx8K outIn: $0.000483/1MOut: $0.0019/1M
ToolsStreamtoolsreasoningimage+1
💬 LLMRead more →
New
Alibaba's 2.4T-parameter multimodal MoE flagship. 1M context, accepts text and images.
1M ctx33K outIn: $0.0035/1MOut: $0.0105/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Classic pixel art generation with 15 styles. Great for game assets, characters, textures, and retro-styled visuals. 16-512px.
$0.0350/unit
🎨 ImageRead more →
Advanced pixel art with Flux model. Same 15 styles as Classic with enhanced quality and detail. 16-512px.
$0.0350/unit
🎨 ImageRead more →
Recraft V3 (Red Panda), supports text-to-image, image-to-image, inpainting, and background replacement. 1024x1024.
$0.0700/unit
image
🎨 ImageRead more →
Recraft V3 for vector graphics, SVG output with image-to-image, inpainting, and background replacement support.
$0.1400/unit
image
🎨 ImageRead more →
Recraft's latest image generation model, photorealistic and illustration styles, text rendering, 1024x1024, ~10s generation.
$0.0700/unit
🎨 ImageRead more →
Recraft's highest quality image model, 2048x2048 output, superior anatomy and detail, ideal for print-ready assets. ~30s generation.
$0.4375/unit
🎨 ImageRead more →
Recraft's highest quality vector model, high-resolution SVG output with fine detail, ideal for logos and brand assets.
$0.5250/unit
🎨 ImageRead more →
Recraft V4 for scalable vector graphics, generates production-quality SVG with discrete color regions and clean geometry.
$0.1400/unit
🎨 ImageRead more →
Voyage AI's latest reranking model, re-scores search results for maximum relevance. Best quality. 200M free tokens.
In: $0.000087/1M
↕️ RerankRead more →
Voyage AI's fast reranking model, cost-efficient re-scoring for high-volume retrieval pipelines. 200M free tokens.
In: $0.000035/1M
↕️ RerankRead more →
Rime's flagship voice model, 269 voices across 9 languages with rich emotional range.
In: $0.0700/1M
tts
🗣 VoiceRead more →
Stylized voices in 4 categories (Professional / Formal / Casual / Energetic). 184 voices, 8 languages.
In: $0.0525/1M
tts
🗣 VoiceRead more →
Latest Mist generation, 83 voices, 4 languages. Improved expressiveness.
In: $0.0437/1M
tts
🗣 VoiceRead more →
Premium model for complex, multi-step work. Augure sovereign Canadian AI.
33K ctx33K outIn: $0.0044/1MOut: $0.0139/1M
ToolsStreamchattools
💬 LLMRead more →
Sarvam's 24B multilingual LLM, fluent in all 22 official Indian languages + English. OpenAI-compatible API. Free tier with no per-token charges. Supports wiki grounding, reasoning effort control, and tool calls.
In: $0.000081/1MOut: $0.000325/1M
Toolstoolsreasoning
💬 LLMRead more →
Batch speech recognition with word-level timestamps and language detection across 99 languages.
$0.000170/unit
audio
📝 TranscriptionRead more →
Ultra-low latency (<150ms) live speech recognition. 93.5% accuracy across 90+ languages. WebSocket streaming with VAD.
$0.000170/unit
audio
📝 TranscriptionRead more →
Highest-quality photo-realistic image generation perfect for professional print media
$0.1400/unit
🎨 ImageRead more →
8B parameter MMDiT model. Superior quality, typography, and prompt adherence. Most powerful in SD family.
$0.1138/unit
🎨 ImageRead more →
Distilled SD3.5 Large optimized for faster inference with fewer steps.
$0.0700/unit
🎨 ImageRead more →
Balanced SD3.5 model. Great quality with lower resource requirements than Large.
$0.0612/unit
🎨 ImageRead more →
⚡ CapabilityRead more →
ByteDance Seedance 1.5 Pro, text & image to video with native synced audio (voice, SFX, music), via BytePlus (Ark). Billed per output token; audio videos bill a higher rate.
Out: $0.0021/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0, text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0135/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0 Fast — speed-optimized text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0098/1M
image
🎬 VideoRead more →
ByteDance Seedance 2.0 Mini — cost-optimized text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0061/1M
image
🎬 VideoRead more →
New
ByteDance Seedance 2.5, text & image to video, native via BytePlus (Ark). Billed per output token.
Out: $0.0187/1M
image
🎬 VideoRead more →
ByteDance Seedream 5.0 Lite, text-to-image (2K+), native via BytePlus (Ark).
$0.0612/unit
🎨 ImageRead more →
ByteDance Seedream 5.0 Pro, flagship text-to-image (2K+), highest quality + consistency. Native via BytePlus (Ark).
$0.1575/unit
🎨 ImageRead more →
High-quality English-only voice synthesis with natural-sounding output
In: $0.0175/1M
🗣 VoiceRead more →
Natural voice synthesis supporting 60+ languages with comprehensive multilingual capabilities
In: $0.0175/1M
🗣 VoiceRead more →
Ultra-fast English voice synthesis optimized for speed with minimal latency
In: $0.0175/1M
🗣 VoiceRead more →
Upstage lightweight model for fast, cost-effective inference.
33K ctx4K outIn: $0.000262/1MOut: $0.0010/1M
Streamchat
💬 LLMRead more →
Upstage previous-generation pro model with 64K context.
66K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamchattools
💬 LLMRead more →
Upstage nightly build of Solar Pro 2 with latest improvements.
66K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamchattools
💬 LLMRead more →
Upstage flagship 100B MoE model with reasoning capabilities and strong multilingual performance.
131K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
ToolsStreamchatreasoningtools
💬 LLMRead more →
Perplexity's lightweight search-augmented model, real-time web search with citations, fast responses, 128K context.
128K ctx16K outIn: $0.0018/1MOut: $0.0018/1M
Streamsearch
💬 LLMRead more →
Perplexity's most thorough research model, multi-step deep web investigation with comprehensive citations and reasoning.
128K ctx33K outIn: $0.0035/1MOut: $0.0140/1M
Streamsearchreasoning
💬 LLMRead more →
Perplexity's advanced search model, deeper web research, multi-step reasoning with citations, 200K context.
200K ctx16K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolssearch
💬 LLMRead more →
Perplexity's reasoning model, extended thinking with real-time web search, ideal for complex research and analysis.
128K ctx16K outIn: $0.0035/1MOut: $0.0140/1M
ToolsStreamtoolssearchreasoning
💬 LLMRead more →
World's fastest, most emotive ultra-realistic text-to-speech with 60+ emotions and 42 languages
In: $0.0875/1M
🗣 VoiceRead more →
Ultra-low latency (40ms) speech generation optimized for real-time applications
In: $0.0875/1M
🗣 VoiceRead more →
Claude's best model for complex agents and coding
200K ctx64K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Best Claude Sonnet for complex agents and coding (1M context)
1M ctx64K outIn: $0.0053/1MOut: $0.0262/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
Anthropic's most capable Sonnet — closes the gap with Opus for complex agents and coding (1M context)
1M ctx64K outIn: $0.0035/1MOut: $0.0175/1M
ToolsStreamtoolsreasoningimage
💬 LLMRead more →
OpenAI Sora 2 video generation. Async jobs polled by app:openai:poll-videos.
$0.7000/unit
🎬 VideoRead more →
OpenAI Sora 2 Pro, higher resolution (1080p cinematic 1792x1024 / 1024x1792) at same aspect ratios as Sora 2.
$3.50/unit
🎬 VideoRead more →
New
Current HD voice model. Highest fidelity prosody and voice cloning.
In: $0.1750/1M
🗣 VoiceRead more →
Generate music and sound effects up to 3 minutes from text prompts. Produces structured compositions with intros, development, and outros at 44.1kHz stereo.
$0.3500/unit
🎵 MusicRead more →
Enterprise-grade music and sound generation. Produces structured compositions with intros, development, and outros at 44.1kHz stereo. 8-step inference for fast generation.
$0.3500/unit
🎵 MusicRead more →
Separate audio into individual stems (vocals, drums, bass, etc). 2-stem or 6-stem modes.
$0.0029/unit
audio
🎚 StemsRead more →
StepFun's latest reasoning model, 196B MoE (11B active), 256K context, open-source Apache 2.0. Fast agentic intelligence with strong code and math.
256K ctx16K outIn: $0.000175/1MOut: $0.000525/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
StepFun's multimodal flagship reasoning model.
256K ctx16K outIn: $0.000350/1MOut: $0.0020/1M
ToolsStreamtoolsreasoning
💬 LLMRead more →
Enhanced vocal quality and refined audio processing for music generation
$0.1050/unit
🎵 MusicRead more →
Excellent prompt understanding with faster generation speeds, supports up to 8 minute tracks
$0.1050/unit
🎵 MusicRead more →
Advanced model with enhanced tonal variation and excellent prompt understanding
$0.1050/unit
🎵 MusicRead more →
Cutting-edge model with enhanced quality and capabilities for AI music generation
$0.1050/unit
🎵 MusicRead more →
Upstage advanced synthetic reasoning model with 64K context.
66K ctx8K outIn: $0.000262/1MOut: $0.0010/1M
Streamchat
💬 LLMRead more →
Studio-grade lip sync. Syncs video lip movements to match any audio input with natural speaker style preservation.
$0.0875/unit
videoaudio
🎬 VideoRead more →
Premium lip sync with enhanced detail preservation for beards, teeth, and fine facial features using diffusion-based super resolution.
$0.1452/unit
videoaudio
🎬 VideoRead more →
Fast bilingual workhorse with 128K context for general inference, document processing, and EN/FR tasks. Augure sovereign Canadian AI.
131K ctx33K outIn: $0.000630/1MOut: $0.0019/1M
ToolsStreamchattools
💬 LLMRead more →
Tripo's P1 model for image-to-3D, converts a single image into a game-ready low-poly mesh (50 to 20k faces) with clean topology and PBR textures.
$0.8750/unit
image
📦 3DRead more →
Tripo's P1 model for multiview reconstruction, builds a game-ready low-poly mesh with clean topology from up to 4 view images (front required).
$0.8750/unit
image
📦 3DRead more →
Tripo's P1 model for text-to-3D, generates game-ready low-poly meshes (50 to 20k faces) with clean topology and PBR textures. Exports GLB, FBX, OBJ, USD, STL.
$0.7000/unit
📦 3DRead more →
Splits an existing 3D model into separately editable parts, textures and PBR materials preserved. A practical level of separation for most assets.
$0.7000/unit
3d
📦 3DRead more →
Splits an existing 3D model into many separately editable parts, textures and PBR materials preserved. Maximum component separation for deep editing and production.
$0.7000/unit
3d
📦 3DRead more →
Splits an existing 3D model into a few separately editable parts, textures and PBR materials preserved. Quick separation for review and lightweight prep.
$0.7000/unit
3d
📦 3DRead more →
Tripo's balanced v2.5 model for converting a single image into a 3D mesh, extracts geometry, texture, and materials from a photo or illustration.
$0.5250/unit
image
📦 3DRead more →
Tripo v2.5 multiview generation, reconstructs detailed meshes from multiple image perspectives for improved accuracy.
$0.5250/unit
image
📦 3DRead more →
Tripo's balanced v2.5 model for generating 3D meshes from text descriptions, produces detailed geometry with PBR materials in seconds. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's stable v3.0 model for image-to-3D, reconstructs clean geometry, textures, and PBR materials from a single image, up to 2M polygons in Ultra mode.
$0.5250/unit
image
📦 3DRead more →
Tripo v3.0 multiview generation, reconstructs detailed meshes from up to 4 view images (front required) with stable, production-proven quality.
$0.5250/unit
image
📦 3DRead more →
Tripo's stable v3.0 model for text-to-3D, crisp edges and coherent structure with PBR materials, up to 2M polygons in Ultra mode. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's flagship v3.1 model for image-to-3D, reconstructs high-fidelity geometry, textures, and PBR materials from a single image, up to 2M polygons in Ultra mode.
$0.5250/unit
image
📦 3DRead more →
Tripo's flagship v3.1 image-to-3D with 8K Ultra textures, reconstructs maximum-fidelity geometry and materials from a single image for hero assets.
$0.8750/unit
image
📦 3DRead more →
Tripo v3.1 image-to-3D in part-segmentation mode, converts a single image into an untextured mesh split into separately editable parts.
$0.7000/unit
image
📦 3DRead more →
Tripo's highest-fidelity 3D generation, v3.1 reconstructs detailed meshes from up to 4 view images (front required) for maximum accuracy.
$0.5250/unit
image
📦 3DRead more →
Tripo's highest-fidelity pipeline, v3.1 multiview reconstruction with 8K Ultra textures from up to 4 view images (front required).
$0.8750/unit
image
📦 3DRead more →
Tripo v3.1 multiview reconstruction in part-segmentation mode, builds an untextured segmented mesh from up to 4 view images (front required).
$0.7000/unit
image
📦 3DRead more →
Tripo's flagship v3.1 model for text-to-3D, sculpture-level geometry with crisp edges and PBR materials, up to 2M polygons in Ultra mode. Exports GLB, FBX, OBJ, USD, STL.
$0.3500/unit
📦 3DRead more →
Tripo's flagship v3.1 text-to-3D with 8K Ultra textures, maximum-fidelity materials and fine surface detail for hero assets and close-ups. Exports GLB, FBX, OBJ, USD, STL.
$0.7000/unit
📦 3DRead more →
Tripo v3.1 text-to-3D in part-segmentation mode, generates an untextured mesh split into separately editable parts for editing, rigging prep, and 3D printing. Exports GLB, FBX, OBJ, USD, STL.
$0.5250/unit
📦 3DRead more →
Turns a single image into a photorealistic 3D Gaussian Splat in under a minute. Ideal for AR/VR, web scenes and visualization — a view-ready format, not an editable or printable mesh. Exports SPLAT.
$0.5250/unit
image
📦 3DRead more →
Low-latency speech generation in 32 languages, optimized for real-time conversational AI
In: $0.1925/1M
🗣 VoiceRead more →
The most creative, intelligent and personalizable image generation models built on a new groundbreaking architecture that delivers ultra high quality and 10x higher cost efficiency.
$0.0707/unit
🎨 ImageRead more →
The most creative, intelligent and personalizable image generation models built on a new groundbreaking architecture that delivers ultra high quality and 10x higher cost efficiency.
$0.1750/unit
🎨 ImageRead more →
Latest Veo with video extending capability. Create and extend AI-generated videos with improved consistency.
$0.7000/unit
imagevideo
🎬 VideoRead more →
Fast variant of Veo 3.1. Quick video generation and extending with good quality.
$0.1750/unit
imagevideo
🎬 VideoRead more →
Cost-efficient Veo 3.1 tier — video with audio at a fraction of the price (720p/1080p, no 4K).
$0.0875/unit
image
🎬 VideoRead more →
Generate and edit sound effects from video (SFX 1.6). Provide a video URL and optional text prompt; adds seamless extension, looping ambiences, and AI inpainting to erase/replace moments.
$0.0875/unit
video
🔊 Sound FXRead more →
Reference-based video generation using multiple images for visual consistency. 4-second 720p output.
$0.7000/unit
image
🎬 VideoRead more →
Create a talking avatar from an image or video with AI-generated speech from text
$0.8750/unit
🎬 VideoRead more →
Synchronize video lips to an audio file for realistic dubbing and voice replacement
$0.7000/unit
🎬 VideoRead more →
High-quality text-to-video with anime style support. 5-second 1080p output.
$0.7000/unit
🎬 VideoRead more →
Generate images from 1-7 reference images. Upload reference images to get started.
$0.0700/unit
🎨 ImageRead more →
Generate images from text or reference images. Supports 1080p, 2K, and 4K resolution.
$0.0525/unit
🎨 ImageRead more →
Ultra-fast video from text or image. Upload an image for image-to-video mode.
$0.3325/unit
🎬 VideoRead more →
Voyage AI's balanced general-purpose embedding model, strong quality at moderate cost. 200M free tokens.
In: $0.000105/1M
🔍 EmbeddingRead more →
Voyage AI's most capable general-purpose embedding model, highest quality retrieval and semantic search. 200M free tokens.
In: $0.000210/1M
🔍 EmbeddingRead more →
Voyage AI's fastest and cheapest embedding model, ideal for high-volume, low-latency use cases. 200M free tokens.
In: $0.000035/1M
🔍 EmbeddingRead more →
Voyage AI's code-optimized embedding model, best-in-class for code search, retrieval, and similarity. 200M free tokens.
In: $0.000315/1M
🔍 EmbeddingRead more →
Voyage AI's long-context embedding model, optimized for large documents and RAG with extended context windows. 200M free tokens.
In: $0.000315/1M
🔍 EmbeddingRead more →
Voyage AI's multimodal embedding model, embeds both text and images into a shared vector space for cross-modal search.
In: $0.000210/1M
image
🔍 EmbeddingRead more →
Alibaba Wan 2.2 Animate Mix, composite a character into a reference video (native DashScope, Singapore).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.2 Animate Move, animate a character image with a reference video's motion (native DashScope, Singapore).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.6 image editing. Edits 1-4 input images from a natural-language instruction (native DashScope, US-Virginia).
$0.0525/unit
image
🎨 ImageRead more →
Alibaba Wan 2.6 image-to-video. Animates a still image into video while preserving subject, style and detail (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
Alibaba Wan 2.6 text-to-image. Photorealistic generation with accurate text rendering and flexible artistic styles (native DashScope, US-Virginia).
$0.0502/unit
🎨 ImageRead more →
Alibaba Wan 2.6 reference-to-video. Generates video preserving the look (and voice) of subjects from a reference video and/or reference images (native DashScope, US-Virginia).
videoimagevideo
🎬 VideoRead more →
Alibaba Wan 2.6 text-to-video. Cinematic motion generation from text with strong instruction following (native DashScope, US-Virginia).
video
🎬 VideoRead more →
Alibaba Wan 2.7 image-to-video. Animates a still image into video with first-frame fidelity, clips up to 15 seconds (native DashScope, US-Virginia).
videoimage
🎬 VideoRead more →
New
Alibaba Wan 2.7 text-to-image and editing. Served from Singapore — the US tenant returns AccessDenied.
$0.0481/unit
🎨 ImageRead more →
New
Alibaba Wan 2.7 pro tier, up to 4K output for print and large-format work. Served from Singapore.
$0.1203/unit
🎨 ImageRead more →
Alibaba Wan 2.7 text-to-video. First and last frame control, clips up to 15 seconds, sharper motion and instruction following (native DashScope, US-Virginia).
video
🎬 VideoRead more →