Writing Sound and Dialogue in MiniMax H3 Prompts
Audio is generated with the picture, so it is yours to direct. How to write dialogue, ambience and silence, and how much speech actually fits in a short clip.

MiniMax H3 generates audio in the same pass as the picture, which means the soundtrack is something you direct rather than something you fix afterwards. Dialogue goes in quotation marks, ambience gets its own sentence, and silence has to be requested explicitly or the model will fill it.
Why does writing sound matter?
Because the model decides it whether you write it or not.
Leave audio unspecified and you get whatever seems to fit, which regularly includes a music bed. For UGC style creative that is actively harmful, because a score signals advertising within the first half second and undoes the unproduced look you were going for.
The fix is one short sentence at the end of every prompt. It costs nothing and it removes the most common reason a clip has to be regenerated.
How do you write dialogue?
Put the exact words in quotation marks, inside the sentence that describes the action.
She looks straight into the camera and says, "Nobody warned me about the purge week."
Not "she talks about the product", which leaves the model to invent words. Not a separate line saying "dialogue: ...", which reads as metadata rather than as part of the shot.
Because the line and the mouth are produced together, you get matching lip movement without a separate lip sync step. That single fact is the biggest practical improvement over the split pipelines of a year ago, and it is why a talking head draft is now watchable straight out of the generation.
How much dialogue actually fits?
Less than people write. This is the most common cause of a clip coming back rushed or with a clipped ending.
As a rough guide at a natural speaking pace:
| Clip length | Comfortable dialogue |
|---|---|
| 5 seconds | One or two short lines, around 12 to 15 words |
| 8 seconds | Two to three lines, around 20 to 25 words |
| 10 to 15 seconds | A short paragraph, but consider breaking it into beats |
If a line feels tight when you say it out loud at a normal pace, it is too long. Say it aloud before you generate; it takes three seconds and saves a render.
For ad hooks this constraint is a feature. A line that has to fit in five seconds is forced to be specific, and specific lines outperform general ones.
Can you direct the delivery?
Partly, and the method matters.
Asking for an emotion by name is unreliable. "She says it excitedly" produces a wide range of results.
Describing the manner works better, because it gives the model something physical to reproduce:
She leans slightly toward the camera and says quietly and quickly, as if telling a secret, "I paid full price and I would do it again."
He pauses, exhales, and says flatly, "Three months. That is how long I ignored this."
She laughs once, shakes her head slightly, and says, "I did not think this would work."
Physical direction plus the line beats an adverb. And a small action before the line, a pause, an exhale, a glance away, buys a beat of natural rhythm that makes the delivery sound less recited.
How do you write ambience?
One sentence, naming the room and one or two specific sounds.
Kitchen room tone, the low hum of a fridge, no music.
Street ambience, distant traffic, no music.
Quiet room tone, faint keyboard sounds, no music.
Ambient car interior sound, voice close and clear, no music.
Wind in the trees, distant birds, no music.
Ambience does more for believability than most people expect. A clip with clean silent audio feels like a render. A clip with a fridge humming feels like a kitchen. It is a cheap sentence with a large effect.
How do you write action sounds?
Name the sound the action makes. This is especially useful for product beats where there is no dialogue.
Satisfying paper crinkling and gentle tapping sounds, no music.
The quiet click of the pump, no music.
The sound of pouring liquid and quiet kitchen room tone, no music.
A soft connector click, quiet room tone, no music.
Unboxing and demonstration shots live or die on this. The sound of the action is a large part of why that format holds attention.
When should you want music?
Rarely in UGC style creative, sometimes in lifestyle and brand frames.
If you do want it, describe it rather than naming a genre and hoping:
A sparse, slow piano line under the ambience, low in the mix, no vocals.
Be aware that a music bed makes a clip read as produced, which is the opposite of what most creator style ads are trying to achieve. For a lifestyle brand frame that is fine and often right. For a hook meant to look like someone's phone footage, it undermines the whole thing.
A useful default is to generate without music and add a track in the editor if the finished cut wants one. That way you keep the choice rather than inheriting it.
What if the voice is wrong for your brand?
Keep the clip and replace the audio. The picture is what took the work.
Generated delivery is variable in a way that generated lip sync is not. You will get readings that are too flat or too keen for your brand, and the fastest fix is often not another generation but a voice swap in the editor, followed by captions.
This is a normal finishing step rather than a failure, and it is one reason to think of the generation as a draft. The reasoning is in drafts against final cuts.
How does sound differ across the three modes?
The audio instruction works the same way in every mode, but what you should write changes with what the model can see.
Text to video. Full control and full responsibility. The model is inventing the room, so name the ambience or it will choose one that may not match the setting you described.
Image to video. Your still fixes the location, so write ambience that matches what is visibly in the frame. A kitchen photo with street noise under it reads as a mistake. This is a small detail that separates convincing clips from odd ones, and it is covered alongside the mode in how to use image to video.
Reference to video. The model composes a new scene from your references, so the ambience should describe the scene you asked for rather than the rooms your reference photos were taken in.
Across all three, keep the audio sentence last in the prompt. It is not a rule the model enforces, it is a habit that makes prompts easier to scan and edit when you are producing a batch and changing one element at a time.
Does sound matter if people watch muted?
Yes, for two reasons.
First, muted is only the first contact. Viewers who stay often turn sound on, and that is the moment the clip either holds up or falls apart.
Second, captions come from somewhere. A clip with a real spoken line gives you a caption track with natural timing. A silent clip means writing and timing captions from scratch. Our guide to adding captions covers the workflow.
Frequently Asked Questions
How do I write dialogue in a MiniMax H3 prompt?
Put the exact line in quotation marks inside the action sentence. Sound is generated with the picture, so a quoted line gets spoken rather than described.
How much dialogue fits in a five second clip?
Roughly two short lines, or about twelve to fifteen words spoken at a natural pace. Writing more is the main cause of rushed delivery.
How do I stop the model adding music?
Write no music explicitly. If you say nothing about audio the model decides, and a score often appears where you did not want one.
Can I control the tone of delivery?
To a degree. Describing the manner, such as speaking quietly and quickly as if telling a secret, shifts delivery more reliably than asking for an emotion by name.
What is ambience good for?
Making a scene feel like a real place. Room tone, a fridge hum or distant traffic does more for believability than another adjective about lighting.
What if the generated voice is wrong for the brand?
Keep the clip and replace the audio with the voice tools. The generated performance is a draft like everything else, and swapping it is a normal finishing step.
Try one experiment
Generate the same shot three times changing only the audio sentence: once with no music, once with ambience named, once with a music bed. Watch all three. The difference in how produced each one feels will change how you write every prompt afterwards. The full prompt structure is in the prompt guide, and MiniMax publishes the technical background on the model, including its Hugging Face model card.
Frequently Asked Questions
Share
AI Summary

Paul Grisel
Paul Grisel is the founder of VIDEOAI.ME, dedicated to empowering creators and entrepreneurs with innovative AI-powered video solutions.
@grsl_frReady to Create Professional AI Videos?
Join thousands of entrepreneurs and creators who use VIDEO AI ME to produce stunning videos in minutes, not hours.
- Create professional videos in under 5 minutes
- No video skills experience required, No camera needed
- Hyper-realistic actors that look and sound like real people
Get your first video in minutes
Related Articles

How to Run MiniMax H3 Without a GPU
MiniMax H3 has public weights, so people try to self host it. Here is what that actually costs in hardware and time, and the browser route if you just want the clips.

MiniMax H3 Prompt Guide: How to Write a Shot
The five part structure that gets usable clips out of MiniMax H3, why keyword lists fail, and how to write the sound as deliberately as the picture.

12 MiniMax H3 Prompt Examples You Can Copy
Twelve complete, ready to paste prompts for MiniMax H3 covering UGC hooks, product demos, unboxing, app screens and more, with notes on what to change.