AI Models
The models behind Visxy
What each model is good at, where it differs from the others, and what it costs to run.
AI Models
What each model is good at, where it differs from the others, and what it costs to run.
MiniMax's current Hailuo video model. For a Visxy 9:16 clip with one clear action — not a stills generator.
Credits come with a Generator plan. The exact cost is shown on the Generate button before anything is spent.
MiniMax H3 is the current Hailuo video model. MiniMax built it as a general-purpose clip engine: a sentence of action becomes a short video, with optional first and last stills when you already know the frame. On Visxy it is a video-library prompt that already describes one take — 9:16, one subject, one beat.
It is not a photograph model. If the brief is a still, you are in the wrong place. It is also not Hailuo 02 with a new badge: H3 is the flagship MiniMax ships now. We do not list the older Hailuo rows here, so there is nothing to "upgrade from" in the picker.
A single, followable action. The Visxy video library is mostly a vertical shot with a sentence of motion. H3 is built to take that sentence — a walk, a turn, a look into camera, a small piece of business with the hands. MiniMax sells instruction-following and in-clip text/brand work as the reason H3 exists. Treat that as the category, not as a promise we have measured on Visxy: we have not published a run yet.
9:16 is the native library frame and it is on this model. Use it when the prompt came from the video library. A wide frame is for a landscape brief, not for "making it cinematic."
What it is not: a twelve-file omni studio. MiniMax's own Hailuo tool accepts stacks of images, clips, and audio. Visxy does not. Here you get a prompt, and at most two stills on the image-to-video side.
There is no other MiniMax row in this catalog. The comparison that matters is H3 against itself used wrongly.
Text-to-video is for a prompt that already is the shot. Image-to-video is for a frame you already like. Mixing those — attaching a still and then re-describing the entire set in words — is how the clip leaves the picture. Older Hailuo 02 habits (short, soft, "just make it move") undersell H3; stuffing it with five events oversells it. One take.
MiniMax talks about native stereo sound on H3. Visxy does not give this model a separate audio chip. Do not hunt for an on/off control that is not on the page. If the clip arrives with room tone, that is the model; if a line of dialogue has to be perfect, plan to replace the track.
Paste the library prompt, then delete the second event. "She turns and then walks away" is two shots. H3 will try both and usually land neither.
Name the camera only if the move is the point. Name the wardrobe only if there is no still. A Visxy prompt that already has lens, light, and clothes is fine as text-to-video; it is too much as an image-to-video prompt.
Keep the clip inside a length you can watch twice. Extra seconds with no new action are drift. Do not max the slider on a one-beat library prompt.
Try the smaller size first, then re-run a keeper larger. A bad crop at the top size is still a bad crop.
No Visxy wall-clock yet. Generate fewer clips. Change one thing.
There is no public MiniMax H3 Edit page. Image-to-video is this model with a picture attached. The catalog names the partner "MiniMax H3 Edit"; it is not a separate article.
One still is the opening frame. A second, if you attach it, is the closing frame. Crop to 9:16 before you attach if that is the delivery — this shape has no frame picker of its own.
Describe only the motion. The still has already said who and where.
No. H3 is MiniMax's current Hailuo flagship. Hailuo 02 is an older generation and is not a row in this catalog. Do not copy 02 settings onto this page.
Yes, when they describe one action. A 9:16 library sentence is the intended input. A prompt that is really a still belongs on an image model, not here.
No. On Visxy, H3 takes a prompt, and at most two stills when you animate. There is no public page for a multi-file omni workflow.
No. Attach a reference to MiniMax H3 and write the motion. A second still is the last frame. We do not publish a separate article for that shape.