Meet the underlying models behind the tools: vendors, modalities, and which tools use them. Browse by vendor or modality, then verify against official docs on the tool pages.
OpenAI
OpenAI's latest frontier model, released in September 2026 with an Astra Pro tier and served in ChatGPT and Codex as the current default flagship.
OpenAI
OpenAI's current API line since July 2026, split into Sol, Terra, and Luna tiers by capability and cost, carrying over GPT-5's reasoning and tool calling.
OpenAI
OpenAI model released on 2025-08-07 that replaced GPT-4o as ChatGPT's flagship on day one and routes each prompt to the reasoning effort it needs.
OpenAI
OpenAI's fourth-generation model, known for reasoning and writing quality, used as the default engine by many writing, coding, and knowledge tools.
OpenAI
OpenAI's first open-weight models, published on 2025-08-05 as 120b and 20b variants under Apache 2.0, so they can be self-hosted and fine-tuned freely.
OpenAI
OpenAI's image model dated 2026-09-08 that replaced DALL·E 3 as ChatGPT's built-in generator, with stronger long-prompt adherence and in-image text.
OpenAI
OpenAI's third text-to-image model, permanently removed from the API on 2026-05-12; ChatGPT images now run on GPT Image 2.5, and Microsoft has de-branded Bing Image Creator to Image Creator from Designer.
OpenAI
OpenAI's second text-to-image model with editing and variations, which commercialized diffusion-based generation; retired alongside DALL·E 3 on 2026-05-12.
OpenAI
The OpenAI text-to-image family from DALL·E 2 to DALL·E 3, both SKUs removed on 2026-05-12 in favour of GPT Image 2.5.
OpenAI
Sora 2 and Sora 2 Pro (2025-09-30) raised physics realism and audio-visual sync, but the standalone app closed on 2026-04-26 and the API is being retired from 2026-09-24.
Anthropic
Anthropic's Claude family, known for long context, writing, and coding; current members are Claude Opus 5, Claude Sonnet 5, and the 4.5 tier, widely adopted by developer and writing tools.
Anthropic
Anthropic's flagship since 2026-07-24, the strongest of the line on long-horizon agents and hard reasoning, and the most expensive and slowest to call.
Anthropic
The main Claude 5 tier released in June 2026, close to Opus 5 in capability at lower cost and latency, and a common default for coding and agent workflows.
Anthropic
The previous flagship from November 2025, a step change on long-horizon coding and multi-step agents and the coding benchmark before the Claude 5 generation.
Anthropic
The previous mainstream model (September 2025), balanced across reasoning, writing, and coding, which took over Cursor-style defaults from Claude 3.5 Sonnet.
Anthropic
The light tier of 2025-10-15, near-instant responses with capability close to the previous flagship, suited to high-frequency chat and delegated subtasks.
Anthropic
The Claude 3.5 generation, strong at long documents and code comprehension; the whole line left service with the 2025-10-28 retirement window.
Anthropic
The Claude 3 generation that introduced vision across Haiku, Sonnet, and Opus; Sonnet retired on 2025-07-21 and the line no longer serves as a flagship tier.
Google's natively multimodal family; the text line today is Gemini 3 Pro with 2.5 and Omni tiers, while images go to Nano Banana Pro, video to Veo, and music to Lyria.
Google's flagship since 2025-11-18, clearly ahead of the 2.5 generation on reasoning and multimodality; the Ultra tier was dropped in favour of the Pro naming.
Google's 2026 omni-class fast tier, unifying text, image, audio, and video input for low-latency realtime multimodal interaction.
The previous flagship from 2025-03-25, a long-context reasoning model with thinking budgets that still powers many third-party tools.
The fast tier of the 2.5 generation from June 2025, handling high-concurrency reasoning and multimodal work at low cost and latency.
Google's image model (model id gemini-3.1-flash-image), popularised under the Nano Banana name, which replaced Imagen 4 as Google's image SKU with strong multi-image consistency and conversational editing.
Google DeepMind
Google's current video generation release, keeping Veo 3's joint audio-visual output and camera-language control with better subject consistency and reference and frame controls.
Google DeepMind
Google DeepMind's first audio-visual video model, top-tier on quality and consistency; veo-3.0 was closed on 2026-06-30 and Veo 3.1 is the current version.
Google DeepMind
Google DeepMind's music generation model, producing full vocal tracks from text prompts for creative tools such as Flow and the Gemini app.
Meta
Meta's open-weight model family and the de facto base for local and self-hosted deployments; the newest generation is Llama 4.
Meta
Meta's open-weight generation released 2025-04-05: Scout and Maverick were downloadable for self-hosting on day one, while Behemoth was explicitly left in training.
xAI
xAI's flagship in its current documentation (2026), extending 4-series reasoning and realtime information across Grok and the X product line.
xAI
xAI's fourth-generation model with stronger reasoning and realtime information, deeply tied to X data but now two point releases behind Grok 4.6.
xAI
xAI's third-generation model for realtime Q&A in a direct style.
xAI
xAI's image, video, and voice generation line, a consumer-facing product for turning text into quick clips with audio.
Mistral
Mistral's current open flagship (version v25.12), strong at multilingual work, reasoning, and function calling, with open weights for self-hosting.
Mistral
The generic label for Mistral's flagship tier, popular in European enterprise deployments; in practice this string now resolves to Mistral Large 3.
Mistral
Mistral's efficient lightweight model for local and edge deployment.
Mistral
Mistral's code-specialised model (25.08 release) for inline completion and repository-scale context, a common backend for IDE plugins.
Mistral
Mistral's audio model, feeding speech straight into a multimodal model for transcription and spoken question answering across languages.
DeepSeek
DeepSeek's open mainline from December 2025, taking over from V3.1 on long-context reasoning and tool calling while keeping aggressive pricing.
DeepSeek
DeepSeek's official V4-era SKU (snapshot V4-Pro-0813) for complex reasoning and long-code work, called as deepseek-v4-pro in the API.
DeepSeek
The fast tier of DeepSeek's V4 generation, published as deepseek-flash with V4.1-Flash current, serving high-volume and batch calls cheaply.
DeepSeek
DeepSeek's open model family; today that is the V3.1/V3.2 mainline plus the two V4 tiers, and plain 'DeepSeek V4' is not an official SKU name.
Alibaba
Alibaba's open-weight Qwen generation released 2025-04-28, spanning edge to server sizes and one of the main self-hosted bases in China.
Alibaba
Alibaba's flagship announced 2025-09-24 with over 1T parameters, aimed at complex reasoning and agent tasks and served mainly through the API.
Alibaba
Alibaba's next Qwen release, weights published 2026-08-12, continuing the open-weights line for self-hosting and fine-tuning.
Alibaba
Qwen's flagship API tier with strong Chinese understanding and multimodal ability, widely deployed in enterprises; the matching open weights today are the Qwen3 series.
Alibaba
Qwen's balanced tier for most Chinese chat and writing scenarios.
Alibaba
Qwen's high-throughput tier for large-scale text processing at low cost.
Alibaba
Alibaba's open video generation family, top-tier among open models for text- and image-to-video, widely used in Chinese creator communities.
Zhipu AI
Zhipu AI's previous flagship from September 2025 with open weights and strong coding and long-context work, widely cited by self-hosted stacks.
Zhipu AI
Zhipu AI's newer mainline generation, replacing GLM-4.6 as the promoted tier with stronger agent and coding capability and longer context.
Zhipu AI
Zhipu AI's current higher tier, shipping open weights with million-token context for long documents and whole-repository code work.
ByteDance
ByteDance's image generation model of 2025-09-09, delivered via Volcano Engine and the Doubao and Jimeng entry points; a mainstream Chinese text-to-image option.
Kuaishou
Kuaishou Kling's mature mainline release, with clearer subject consistency, motion range, and Chinese-prompt adherence than the 1.x line.
Kuaishou
Kuaishou's latest Kling generation (2026), generating up to 15 seconds per run with 4K output and lip-sync, a leading Chinese video model.
MiniMax
Umbrella for MiniMax Hailuo video, covering Hailuo-02, 2.3, and the open omni H3, with natural motion, realistic physics, and good Chinese-prompt understanding.
MiniMax
Hailuo's previous mainline video release, ahead of 02 on instruction adherence, physics, and character consistency, common for social clips and ad storyboards.
MiniMax
MiniMax's newest generation (July 2026), delivered as an open-weight omni model so video generation and understanding can share one stack.
MiniMax
MiniMax's general-purpose model from May 2026, focused on long context and agent tasks as the text counterpart to the Hailuo video line.
Moonshot AI
Moonshot AI's Kimi family, evolved from early long-context models to K2, K2.5, and the 2.8T-parameter K3, with ultra-long document reading and file parsing still its signature.
Moonshot AI
Moonshot AI's open flagship of July 2025: a 1T-total, 32B-activated MoE released under a modified MIT licence, standout at agentic and coding work.
Moonshot AI
The January 2026 iteration of K2, strengthening multimodal understanding and reasoning chains while keeping long-document parsing ahead.
Moonshot AI
Moonshot AI's newest flagship (July 2026) at 2.8T parameters, built for long-horizon agents and complex reasoning.
Stability AI
Stability AI's open flagship text-to-image model with a huge local-deployment and fine-tuning ecosystem; the base model for Civitai and ComfyUI communities.
Stability AI
The classic Stable Diffusion version — compact with a mature ControlNet ecosystem, the default entry point for local deployment.
Stability AI
Stability AI's realtime text-to-image model with one-step generation for interactive canvases.
Stability AI
Stable Video Diffusion, the reference open model for image-to-video with broad community fine-tuning.
Stability AI
Stability AI's music and sound model generating commercially usable music, samples, and effects.
Black Forest Labs
Black Forest Labs' image models with leading quality and text rendering; the current line is FLUX.2 [pro], [dev], and [klein], with the FLUX.1 trio still in circulation.
Black Forest Labs
Black Forest Labs' image generation line of December 2025: [pro] via API, [dev] with open weights for fine-tuning, and [klein] released open in January 2026 as the light, fast tier.
Open-source community
The open-source image super-resolution model running fully local and free — the engine inside upscalers like Upscayl.
Midjourney
Midjourney's newest generation, offered as an alpha from 2026-03-17 and running alongside V7 while it stabilises.
Midjourney
Midjourney's generation released on 2025-04-04, ahead of V6 on image quality and prompt adherence and long its main image product.
Recraft
Recraft V3 leads in vector generation and brand-style control for designer workflows; its successor generation is not catalogued here, so check the vendor's current docs.
Ideogram
Ideogram's March 2025 generation, another step up in text rendering and layout composition, common for posters and brand visuals.
Ideogram
Ideogram 2.0 with reliable text rendering for posters, logo concepts, and typographic images, superseded by Ideogram 3 in March 2025.
Ideogram
Ideogram 1.0 pioneered dependable text rendering in text-to-image.
Leonardo AI
Leonardo AI's next-gen foundation model with much better quality and prompt adherence for game and creative assets.
Leonardo AI
Leonardo AI's earlier diffusion line with rich style presets and a large creative community.
Adobe
Adobe's third Firefly model, deeply integrated into Photoshop and Illustrator with clear commercial licensing.
Adobe
Adobe's second Firefly model with mature generative fill and text effects inside design workflows.
Freepik
Freepik's high-fidelity image model producing richly detailed output, bundled with its asset library and licensing.
Magnific AI
Magnific's upscaler known for hallucinated detail reconstruction, turning low-res images into high-detail artwork.
Playground
Playground's canvas model combining generation, inpainting, and layer editing.
Krea
Krea's realtime model rendering sketches live on canvas for extremely fast concept iteration.
Runway
Runway's higher tier released 2025-12-11, pushing physical realism and instruction following further as its new video flagship.
Runway
Runway's video model of 2025-04-01, built around reference-driven character and scene consistency plus camera control, taking the mainline from Gen-3 Alpha.
Runway
Runway's video editing model of 2025-07-31, treating already generated footage as editable material so elements or camera moves can change while the rest holds.
Runway
Runway's performance transfer model of 2025-08-20, mapping an actor's motion and facial performance onto a target character while keeping the original nuance.
Pika
Pika 2.0 with effect templates and character-consistent videos, common for social shorts.
Pika
Pika 1.0, a lightweight and approachable model popular for creative social video.
Luma AI
Luma's video model known for natural physics and cinematic shots, with keyframe control and loops.
Luma AI
Luma's capture and reconstruction tech turning video and photos into 3D scenes and Gaussian splats.
Genmo
Genmo's open video model deployable locally — a milestone for open video generation.
LTX Studio
LTX Studio's video model integrating storyboards, character consistency, and shot control for film narrative.
D-ID
D-ID's photo-driven avatar model animating a single face photo, with an API for support and marketing integrations.
HeyGen
HeyGen's avatar model turning scripts into realistic talking-head video with strong translation and lip-sync.
ByteDance
CapCut's built-in AI models covering smart editing, auto captions, AI effects, and avatars — the richest Chinese template ecosystem.
VEED
VEED's in-browser AI covering auto captions, AI voiceover, noise removal, and avatars.
Suno
Suno's current release from September 2026, whose training set is the first cleared with record-label licences, making compliance the headline claim.
Suno
Suno's music model of 2025-09-23, a clear step over V4 on vocals and arrangement and the mainline until v6 arrived.
Udio
Udio V2 excels at vocal songs and full tracks across genres, with regional regeneration and fine editing.
ElevenLabs
ElevenLabs' TTS family umbrella — industry benchmark for realism and emotional expression, with Eleven v3 as the current mainline.
ElevenLabs
ElevenLabs' voice model released in June 2025, covering 70+ languages with finer emotional and stylistic control, longer continuous generation, and multi-speaker dialogue.
Murf
Murf's voiceover library with 120+ voices in 20+ languages, common in corporate training and explainers.
LOVO
LOVO's Genny voice model with 500+ voices and expressive emotion, paired with video editing.
Play.ht
PlayHT 2.0's low-latency streaming voice model with a multilingual voice library.
Speechify
Speechify's reading model turning documents, web pages, and books into natural speech for learning and accessibility.
Kits.AI
Kits.AI's voice-model training for cloning your own voice or using licensed artist voices in demos.
Adobe
Adobe Podcast's speech enhancement model turning noisy recordings into studio-quality audio in one click, free to use.
Meshy
Meshy-2, the second generation with better text-to-3D and image-to-3D quality for game and e-commerce assets.
Meshy
Meshy-1, the first generation making text-to-3D approachable.
Tripo AI
Tripo AI's 3D model with fast image- and text-to-3D and steadily improving topology.
Deemos
Deemos' high-fidelity 3D model focused on photorealistic character and portrait assets for film and virtual humans.
Spline
Spline's in-browser 3D design AI generating scenes and motion from prompts, embeddable on web pages.
Grammarly
Grammarly's writing assistant model covering grammar, tone, and plagiarism across browser and mobile keyboard.
Rytr
Rytr's lightweight writing model with 40+ use cases and tone templates at an affordable price.
Consensus
Consensus' academic Q&A model answering from peer-reviewed papers with findings and citation-quality markers.
Character.AI
Character.AI's roleplay model with a huge community-built character ecosystem.
Replit
Replit's coding models building runnable, deployable apps from prompts.
ClickUp
ClickUp Brain's project-management models auto-writing tasks, summarizing docs, and answering status questions.
Gamma
Gamma's presentation model generating polished decks and sites from outlines or documents.
Beautiful.ai
Beautiful.ai's smart-layout model auto-aligning and reflowing slides as content changes.
Descript
Descript's editing model that pioneered transcript-driven audio and video editing.
User-trained
Models fine-tuned or trained by users on their own material — common in LoRA communities like Civitai and Tensor.Art.
Multiple vendors
Vendor-built, unpublished models; capabilities follow each vendor's official documentation.
Fliki
Fliki's voiceover library with 2,000+ AI voices in 80+ languages for text-to-video.
NovelAI
NovelAI's story-continuation model for long-form fiction with strong style and character consistency.
NovelAI
NovelAI's anime-style image model with good character consistency and a strong fan-fiction community reputation.
The model directory only lists models that have a public name and are explicitly noted as used by a tool in this directory. Tools that use an unnamed in-house model keep that model string on the tool page instead of becoming a model entry; multiple versions from one vendor are grouped into the corresponding family entry. Capabilities, availability, and licensing terms are governed by the vendor's own documentation.
Tool pages describe product capabilities; model pages answer what engine sits underneath: how versions from the same vendor differ, and whether your shortlisted tools share the same base. Checking the base first narrows the field faster than comparing tool-level features alone.