How to Use MiniMax H3: A Step by Step Guide
A walkthrough of generating your first MiniMax H3 clip on VIDEO AI ME, from picking the model and settings to writing a prompt that survives the first render.

To use MiniMax H3 on VIDEO AI ME, open the editor, select MiniMax H3 in the model dropdown, pick a duration and resolution, write your brief, and generate. Attaching reference images changes the mode automatically. This guide walks the whole loop, including how to fix a draft that comes back wrong.
Step 1: Open the editor and select the model
Open the editor and find the model dropdown in the generation bar. Select MiniMax H3. In the model picker it is listed as MiniMax H3 Max.
Selecting the model updates the settings around it, because each model in the picker supports different durations, resolutions and aspect ratios. If a duration you used with another model disappears, that is why.
Step 2: Choose your settings before you write
Three settings, and getting them right first saves you a rerun.
Duration. Five, eight, ten, twelve or fifteen seconds. Start at five. A five second clip is enough to judge whether a hook works, and there is no point paying for fifteen seconds of an idea you are going to throw away.
Resolution. 480p or 768p. Draft at 480p. Once a version earns its place, rerun the exact same prompt at 768p for the file you publish. Almost all of your generations should be drafts.
Aspect ratio. 9:16, 16:9, 1:1, 4:3, 3:4 or 21:9. Pick the placement you are actually shipping to. If you are making a vertical ad, generate vertical rather than cropping a horizontal clip afterwards, because cropping throws away the composition the model built for you.
One exception worth remembering: if you attach a single image, the output follows that image's aspect ratio and the picker no longer applies.
Step 3: Write the brief as a shot, not a keyword list
This is where most first attempts go wrong. H3 reads the whole brief as one context, so it responds to writing that reads like a director's note and gets confused by writing that reads like a tag cloud.
A brief that works has five parts: who is in frame, what they do, how the camera behaves, how it is lit and graded, and what you hear.
A woman in her late twenties sits cross legged on a bed in a sunlit bedroom, holding a small pink skincare tube. She looks straight into the camera and says, "Nobody warned me about the purge week." Handheld medium close up with a small amount of natural sway, soft window light from the left, clean natural colour. Room tone only, no music. Vertical framing.
Compare that to "skincare ugc ad, girl, bedroom, natural light, vertical, trending", which gives the model nothing to sequence and no idea what should be heard.
A few rules that hold up:
- One subject, one action, one camera move. Stacking three camera moves into five seconds is the most common way to waste a generation.
- Put dialogue in quotation marks. Roughly two short lines fit comfortably in five seconds.
- Say what you want to hear, including silence. "No music" is an instruction worth giving, because otherwise you may get a score you have to strip out.
- Describe the grade in film language. "Warm natural light, shallow depth of field" lands better than "cinematic".
Step 4: Attach reference images if you need them
Attaching images changes the mode, and you should know which mode you are asking for.
No images gives you text to video. The model invents everything.
One image gives you image to video. Your image becomes the literal first frame. This is the mode for a product photo you want to put into motion, and it is the most reliable way to get a real product to look like itself. Crop the still to your placement before uploading.
Two to four images gives you reference to video. The model composes a fresh shot and carries the subjects, product and style from your references through it. Address them by their order in the prompt:
The woman from Image 1 sits at an outdoor cafe table holding the ceramic mug from Image 2. She takes a sip, sets it down, and says, "This is the third one today." Warm late afternoon light, shallow depth of field, static camera at eye level. Street ambience, distant traffic, no music.
Binding each image explicitly is the whole trick. "Use these references" gets you a blend. Naming which image supplies the person and which supplies the product gets you the shot you asked for.
Reference to video always renders at 768p, so the resolution choice is made for you in that mode.
Step 5: Generate, watch, and change one thing
Generate and watch the result the whole way through, with sound on, at the size your audience will see it. A clip that looks weak on a laptop preview often plays fine on a phone, and the reverse happens too.
When something is wrong, resist the urge to rewrite the prompt from scratch. Change one thing:
- Delivery too flat? Rewrite only the dialogue line.
- Framing wrong? Change only the shot size.
- Too much motion? Remove the camera move and leave everything else.
- Wrong mood? Change only the lighting and grade sentence.
This is single variable testing, and it is the reason a fast model is worth using. You learn what the model responds to, and that knowledge makes every later prompt cheaper. Our guide to drafting an ad in minutes turns this loop into a repeatable routine.
Step 6: Finish the clip in the editor
The generated file is a starting point, not a deliverable. Finish it properly:
- Add captions. Most social video is watched muted first, and captions are the single highest impact addition to an AI generated ad. Our guide on adding captions covers the workflow.
- Trim the dead frames. Generated clips often have a soft first and last beat. Cutting them tightens the hook.
- Add your end card, logo and offer. Do not ask the model to draw text, because rendered lettering is still unreliable across every model in this class.
What does it cost to work this way?
Generation is billed in credits included with your plan, from 1,400 a month on Starter at $29 through to 12,000 on Premium at $199. Because drafting at 480p costs meaningfully less than finishing at 768p, a disciplined draft then finish habit stretches an allowance a long way. The pricing page carries the current numbers, and the editor shows the cost of a clip before you commit to it.
Frequently Asked Questions
Do I need to install anything to use MiniMax H3?
No. It runs in the browser on any paid plan. There is no download, no GPU requirement and no local setup.
How do I switch between text to video, image to video and reference to video?
You do not switch manually. The mode follows how many reference images you attach: none is text to video, one is image to video, and two to four is reference to video.
Why did my clip come out in the wrong aspect ratio?
You almost certainly attached one image. Image to video always follows the aspect ratio of the image you upload, so the aspect picker does not apply in that mode. Crop the still before uploading.
How do I get dialogue into the clip?
Write the line inside quotation marks in the prompt. Sound is generated with the picture, so a quoted line is spoken by the subject rather than added afterwards.
What should I do when a generation comes back wrong?
Change one thing and rerun. Rewriting the whole prompt makes it impossible to learn which instruction caused the problem, and this model is quick enough that single variable testing is practical.
Can I keep the same person across several clips?
Attach the same reference image in each generation and describe them the same way. For a spokesperson who must be identical across a whole campaign, the actor tools are the more reliable route.
Where to go next
Once the basic loop is comfortable, the prompt guide goes deeper on writing, and the reference image walkthrough covers multi image shots properly. For platform specific formats, TikTok for Business and Meta's ads guidance both publish current placement specs worth checking before you generate.
Frequently Asked Questions
Share
AI Summary

Paul Grisel
Paul Grisel is the founder of VIDEOAI.ME, dedicated to empowering creators and entrepreneurs with innovative AI-powered video solutions.
@grsl_frReady to Create Professional AI Videos?
Join thousands of entrepreneurs and creators who use VIDEO AI ME to produce stunning videos in minutes, not hours.
- Create professional videos in under 5 minutes
- No video skills experience required, No camera needed
- Hyper-realistic actors that look and sound like real people
Get your first video in minutes
Related Articles

How to Run MiniMax H3 Without a GPU
MiniMax H3 has public weights, so people try to self host it. Here is what that actually costs in hardware and time, and the browser route if you just want the clips.

Writing Sound and Dialogue in MiniMax H3 Prompts
Audio is generated with the picture, so it is yours to direct. How to write dialogue, ambience and silence, and how much speech actually fits in a short clip.

MiniMax H3 Prompt Guide: How to Write a Shot
The five part structure that gets usable clips out of MiniMax H3, why keyword lists fail, and how to write the sound as deliberately as the picture.