M
MJK.Supplies
Home / AI Models / AI Models Directory: Text, Vision, Audio, Image,…
AI Models

AI Models Directory: Text, Vision, Audio, Image, and Video Models

The AI model market is no longer one list of chatbots. Builders now choose between text, reasoning, coding, vision, audio, image, and video models, often inside the same product. This directory gives you a practical map of the major model families available today and what each category is best used for.

M
MJK Supplies · Jun 20, 2026 · 12 min read
ShareXinf↗
AI Models Directory: Text, Vision, Audio, Image, and Video Models

How To Read This Directory

This is a working directory, not a permanent leaderboard. Model names, availability, pricing, context windows, and deprecation dates change quickly. Before you build production logic around any model, check the provider's official model page and record the exact model ID you are calling.

Use this as a decision map:

  • Text models handle writing, reasoning, extraction, support, research, and planning.
  • Coding models specialize in code generation, migration, debugging, tests, and repository work.
  • Vision models read images, charts, screenshots, documents, and UI states.
  • Audio models handle speech-to-text, text-to-speech, real-time voice, music, and sound.
  • Image models generate or edit still images.
  • Video models generate, edit, extend, or reframe moving footage.

Text, Reasoning, and Coding Models

OpenAI. The OpenAI model directory includes the GPT-5 family, Codex-oriented models, realtime models, image models, and legacy models. Use OpenAI when you need broad API coverage, tool use, structured output, coding support, and strong multimodal app infrastructure.

Anthropic Claude. Anthropic's Claude model overview covers current Claude models and the model lifecycle. Claude is often used for long-form reasoning, writing, code review, agent workflows, and high-context business tasks. Fable 5 and Mythos 5 require special attention because access was suspended after launch.

Google Gemini. Google's Gemini model documentation lists Gemini models for text, multimodal reasoning, image, video, embedding, and specialized workflows. Gemini is a strong fit when you want multimodal input, Google ecosystem integration, and media-generation access in the same platform family.

Mistral AI. The Mistral models overview lists open and commercial models such as Medium, Small, Large, Ministral, Codestral, Devstral, OCR, moderation, and Voxtral audio models. Mistral is useful when open-weight strategy, European deployment options, and specialist coding or document models matter.

xAI Grok. The xAI model docs cover Grok text models plus Imagine and Voice APIs. Grok is relevant for chat, reasoning, search-connected experiences, voice, image, and video features inside xAI's API surface.

Meta Llama. Meta's Llama 4 announcement introduced open-weight multimodal models such as Scout and Maverick, with Behemoth previewed as a larger model. Llama is most attractive when you need open models, self-hosting, customization, or model control outside a closed API.

Vision and Multimodal Models

Vision-capable models read visual inputs and produce text, structured data, or tool calls. They are used for document understanding, UI testing, screenshot QA, receipts, charts, compliance review, and customer-uploaded media.

  • OpenAI's latest API models support text and image input for many assistant and agent workflows.
  • Claude models support document, screenshot, and image understanding across business and coding use cases.
  • Gemini is built around multimodal input and is a natural choice for mixed text, image, audio, and video pipelines.
  • Llama 4 moved Meta's open models into native multimodal territory.
  • Mistral includes multimodal generalist and OCR models for document-heavy workloads.

For production, test vision models on your real image quality. A model that performs well on clean benchmarks may fail on blurry receipts, dense dashboards, low-light photos, handwriting, or screenshots with tiny text.

Audio and Speech Models

Audio models fall into several groups: transcription, speech generation, real-time voice agents, music, sound effects, dubbing, voice isolation, and speech-to-speech conversion.

  • OpenAI provides realtime, speech-to-text, text-to-speech, and audio-capable model options through its API model family.
  • Gemini supports audio and multimodal workflows through Google's model ecosystem.
  • xAI exposes Voice APIs alongside Grok and Imagine.
  • Mistral lists Voxtral models for transcription, realtime transcription, text-to-speech, and audio input tasks.
  • Stability AI includes Stable Audio models for music and sound generation.
  • Runway exposes audio generation and transformation options alongside video and image models.

Audio products need extra quality checks. Measure transcription accuracy by accent, background noise, call quality, domain vocabulary, and speaker overlap. For generated speech, evaluate latency, pronunciation, emotional tone, licensing terms, and voice-consent requirements.

Image Generation and Editing Models

Image models are now split between pure generation, editing, inpainting, outpainting, product photography, design iteration, character consistency, and reference-based generation.

  • OpenAI includes GPT Image models for image generation and editing.
  • Google offers Gemini image capabilities and dedicated image-generation models in its AI developer stack.
  • Stability AI's developer platform focuses heavily on image generation and editing endpoints.
  • Runway lists image models such as Gen-4 Image and other hosted image-generation options.
  • Luma pairs Uni-1 image generation with Ray video workflows through its API positioning.
  • xAI Imagine covers image generation and editing features in the Grok ecosystem.

Choose image models by control needs, not only beauty. Product shots, ads, diagrams, thumbnails, and app assets all need different guarantees around text rendering, aspect ratio, reference adherence, brand consistency, and editability.

Video Generation Models

Video models are the fastest-moving and most expensive category. They differ by whether they support text-to-video, image-to-video, video-to-video, reference-to-video, extension, reframing, editing, audio, duration, resolution, and camera control.

  • Google Veo models cover video generation in the Gemini ecosystem.
  • Runway's available models page lists video models such as Seedance, Aleph, Gen, Veo-hosted options, and avatar or realtime outputs.
  • Luma's Ray API focuses on cinematic video generation, video-to-video, reframing, HDR, and production pipeline control.
  • OpenAI's Sora video API guide is important to check because video model availability and retirement timing can change.
  • xAI Imagine includes video generation, image-to-video, video editing, reference-to-video, and video extension categories.

For video, benchmarks matter less than workflow fit. Test prompt control, motion coherence, character consistency, scene edits, duration, output rights, cost per usable clip, moderation rules, and queue reliability.

Choosing By Modality

Use this quick map when deciding where to start:

  • Text and reasoning: OpenAI, Claude, Gemini, Grok, Mistral, Llama.
  • Coding and repository work: OpenAI Codex models, Claude, Mistral Codestral and Devstral, Gemini, Grok Build, Llama for self-hosted setups.
  • Vision and documents: OpenAI, Claude, Gemini, Mistral OCR, Llama multimodal.
  • Voice and audio: OpenAI, Gemini, xAI Voice, Mistral Voxtral, Stability Audio, Runway audio endpoints.
  • Image generation: OpenAI image models, Google image models, Stability AI, Runway, Luma Uni-1, xAI Imagine.
  • Video generation: Google Veo, Runway, Luma Ray, OpenAI Sora where available, xAI Imagine video.

The best stack is usually multi-model. Use one model for research, another for extraction, another for image generation, and another for video. The winning product is rarely the one with the fanciest model name. It is the one with the best routing, fallbacks, evals, and user experience.

Keep A Live Model Inventory

Every serious AI product should maintain a live model inventory. Track provider, model ID, modality, context window, pricing, region availability, data retention, safety policy, latency, quality score, fallback model, and deprecation date.

Update that inventory every time you ship a model change. When a provider retires a model, changes access policy, or introduces a stronger replacement, your product should degrade gracefully instead of breaking in front of users.

#ai-models#llm#vision#audio#video

Related articles

MJK Supplies · Automation Services

Want this built for you?

We design and ship custom AI agents and automation systems for teams that want results, not a backlog. Book a free 30-minute consult — no commitment, no pitch deck.