AI Models
The models behind Visxy
What each model is good at, where it differs from the others, and what it costs to run.
AI Models
What each model is good at, where it differs from the others, and what it costs to run.
AI Models
What each model is good at, where it differs from the others, and what it costs to run.
Google's general-purpose Gemini image model. Open it when a Visxy library prompt already describes the picture and you want a finished frame, not a thumbnail.
The fast lane of Google's Nano Banana family. Built for trying Visxy's long 9:16 prompts — not for print-size files or long edit chains.
The high-fidelity member of the family. Use it when a Visxy prompt has to survive a close look — type, hands, labels — not when you are still shopping looks.
OpenAI's current image model. Built for a long Visxy library prompt — layout, type, and what to keep — not a five-word vibe.
ByteDance's current Seedream stills model. For a Visxy library prompt that has to look designed — type, layout, a 9:16 poster — not a five-word vibe.
ByteDance's cheap Seedream 5 stills SKU. Try a Visxy 9:16 poster here — not Seedream 5 Pro, and not Seedance.
Black Forest Labs' production image model. Photoreal product and materials — not the Flux 3 video row.
Black Forest Labs' cheaper Flux 2 stills SKU — Flex under a short catalog name. Same frames as Flux 2 Pro, not the Pro row and not Flux 3.
Ideogram's type-first image model. Use it when the Visxy prompt is a poster, a label, or a word that has to read.
xAI's current Grok stills model. A Visxy 9:16 library prompt — not the unversioned Imagine row, and not a video model.
xAI's unversioned Grok stills model. Fast 9:16 tries — not Grok Imagine 2.0, and not a video row.
Kuaishou's cinematic clip model. One subject and one camera move that stay coherent — not the five-second Turbo try.
Kuaishou's Omni Kling beside Kling 3.0. Same coherent-move job, with a size picker — not Turbo, not 2.6, and not a multi-file studio.
The short Kling lane next to Kling 3.0. A five-second try — not a setting on the 3.0 page.
Kuaishou's previous Kling. Five or ten seconds, optional sound — not 3.0, not Omni, and not the five-second Turbo row.
Google's current Veo quality tier. Short cinematic clips with sound, 9:16 from the video library — not Fast, not Lite.
ByteDance's long-take video model. Use it when a Visxy video prompt has a sequence to hold — not when you only need a few seconds of motion.
The short draft lane next to Seedance 2.5. Try the shot here; move a keeper to 2.5 only when the clip has to run past a quarter-minute.
MiniMax's current Hailuo video model. For a Visxy 9:16 clip with one clear action — not a stills generator.
Black Forest Labs' video model. A clip with sound in the same pass — not the Flux 2 / Flux 2 Pro stills rows.
Alibaba's current Wan video model. Smooth motion from a Visxy 9:16 prompt, or a still with an optional last frame.
PixVerse V6 under a short catalog name. Short-form motion, audio off until you turn it on — not V5.
xAI's unversioned Grok video model. A longer clip than 1.5 — not the stills rows, and not a five-second sting.
xAI's versioned Grok video model. A short clip from a Visxy prompt or one still — not the unversioned Video row.
Luma's current Ray. A short directed clip — five or ten seconds — not a 16-keyframe studio and not Ray 2.