MiniMax H3 Prompt Guide: Reference Images, Voice Referencing, and Storyboards

Learn how to prompt MiniMax H3 with reference images and video, character sheets, storyboards, voice referencing, consistent characters, camera direction, and sound.

MiniMax H3 Prompt Guide on a video-editing monitor with six cinematic scene examples.

MiniMax H3 can generate video and synchronized sound from text, images, and reference video. You can start with a written idea, animate a first frame, connect a first and last frame, or combine visual references to guide characters, movement, camera work, voice, sound effects and background music.

The most important prompting rule is simple: give every reference one clear job.

Quick Answer

MiniMax H3 Ref. to Video on RunDiffusion accepts reference images and reference videos. Images can define characters, clothing, products, environments, storyboards, or visual style. A reference video can guide motion, camera behavior, cuts, pacing, and timing. When that video includes clear audio, it can also guide voice, delivery, rhythm, and sound. We call using those vocal qualities voice referencing.

H3 can also work with a storyboard. There is not a separate storyboard file type. You upload the storyboard as one or more reference images and explain which shots, framing, subject placement, or shot order those images should guide.

Simple rule for beginners:

  1. Start with one clear subject.
  2. Give the subject one clear action.
  3. Use one main camera movement.
  4. Keep dialogue short.
  5. Add one new reference or instruction at a time.

Start with MiniMax H3 on RunDiffusion. When an image or video should guide the character, product, scene, motion, camera work, voice, or overall direction, use MiniMax H3 Ref. to Video on RunDiffusion.

MiniMax H3 at a Glance

The official MiniMax H3 video-generation documentation lists the model's broader capabilities and limits. The current RunDiffusion tools may expose a more focused set of choices, so use the controls shown on your Board as the final authority:

  • Video length: H3 supports 4 to 15 seconds; the current RunDiffusion tool cards show 5 to 15 seconds
  • Prompt field: Up to 2,000 characters in the current RunDiffusion tool
  • Resolution: 768P or 2K through supported generation and regeneration workflows
  • Reference images: The current MiniMax H3 Ref. to Video tool shows up to 5 images
  • Reference videos: Available through the separate reference-video upload control
  • Video with audio: A reference video can carry voice, rhythm, and other sound cues along with its visual guidance

In the Prompt field, identify uploaded assets clearly as Image 1, Image 2, Video 1, and so on. Give each reference one job so H3 knows whether it should guide identity, styling, movement, camera work, timing, voice, or sound.

Use the duration, resolution, aspect ratio, and input choices shown in the current RunDiffusion workflow. Those platform controls are the final authority for your generation.

The Simple MiniMax H3 Prompt Formula

You do not need to begin with a long technical prompt. For many generations, a clear natural-language brief is enough.

First, select the available duration, aspect ratio, and resolution in RunDiffusion. Those values belong in the platform controls, not in the Prompt field. MiniMax's official examples sometimes repeat duration or aspect ratio in prompt text, but this RunDiffusion guide keeps them separate so readers follow the controls provided by the platform.

Then build the prompt in this order:

  1. Visual direction: State the overall visual style, lighting, and mood.
  2. Reference roles: Explain what each image or reference video should control.
  3. Opening scene: Describe where the clip begins and what is visible.
  4. Action: Explain what the subject does in chronological order.
  5. Camera: Choose one main camera behavior for each shot.
  6. Dialogue and sound: Write the exact dialogue, scene sounds, and music direction.
  7. Ending: Describe the final pose, composition, or action.

For this simple example, select 8 seconds and 16:9 in RunDiffusion, then enter:

Cinematic live-action product video with warm sunrise light and
restrained color contrast.

Image 1 defines the woman's face, copper hair, and green jacket.
Image 2 defines the rooftop garden and sunrise lighting. Video 1
provides the woman's measured walking pace, slow camera movement,
and the warm vocal timbre and calm delivery heard in its audio. Use
Image 1 as the identity and clothing source and Image 2 as the
location source.

The woman walks along the rooftop path while holding a black travel
cup. The camera follows her from the front-left with a slow,
controlled tracking movement. After three steps, she raises the cup,
looks toward the camera, and says, "Take the morning with you." End
on a stable close view of her face and the cup.

Sound includes light wind, soft footsteps, and clear dialogue.

This works because every reference has a job, the action is easy to follow, and the clip has a clear ending.

Choose the Right Generation Mode

Before writing the prompt, decide how your images and reference videos should be used.

  • Text-to-video: You want to create the full scene from a written description
  • First-frame image-to-video: Your image must be the exact opening frame
  • Last-frame image-to-video: Your image must be the exact ending frame
  • First-and-last-frame video: You want H3 to create a transition between two known frames
  • Reference generation: Your images or reference video should guide identity, style, motion, camera work, voice, sound, or pacing

If an image must appear as the exact first moment, use MiniMax H3 and add it as the Start Frame. If an image or video should guide identity, wardrobe, objects, visual direction, a storyboard, motion, camera behavior, timing, voice, or sound, use MiniMax H3 Ref. to Video on RunDiffusion and add it through the matching reference upload control.

How to Use Reference Images

H3 can combine several images, but more references do not automatically produce a better video. Clear, compatible references are more useful than a large, mixed collection.

Give Every Image One Main Job

An image might contain a person, clothing, a product, a background, and a visual style. Tell H3 which part matters.

Image 1 defines the character's face and hairstyle.
Image 2 defines the green jacket and cream shirt.
Image 3 defines the rooftop garden and sunrise lighting.

Build a Small Character Reference Pack

For an important character, start with two or three compatible images:

  1. Identity image: A sharp front or three-quarter view with a clearly visible face.
  2. Second angle or full-body image: A view that shows hairstyle, body proportions, wardrobe, and silhouette.
  3. Optional detail image: A close view of an accessory, makeup detail, tattoo, costume feature, or product that needs to remain consistent.

Choose images with clear lighting without blur or facial obstruction. Avoid references that disagree on age, facial shape, hair length, clothing, or visual style unless you explain exactly which feature to take from each image.

Create a Storyboard with ChatGPT Image 2.0 on RunDiffusion

If you do not already have a storyboard, you can create one with ChatGPT Image 2.0 on RunDiffusion before moving to H3. Add ChatGPT Image 2.0 to your RunDiffusion Board and describe the characters, setting, visual style, shot order, camera views, and actions. Ask for one clean reference image with large, clearly numbered panels.

When character consistency matters, combine the storyboard with a character reference section. Place the recurring characters, clothing, expressions, and important props above the storyboard, then place the ordered sequence below it. This gives H3 one organized image that defines both who appears and what happens.

You can adapt this image-generation prompt:

Create a clean character reference sheet and cinematic storyboard
for a [duration]-second video. In the upper section, show the
recurring characters with clear facial features, full outfits,
expressions, and important props. In the lower section, create
[number] large, numbered storyboard panels in chronological order.
Each panel should show one main action and a clearly labeled camera
view. Maintain the same character identities, wardrobe, setting, and
visual style throughout the sheet. Keep the layout uncluttered and
easy to read.

Match the storyboard to the duration you plan to select in MiniMax H3. This timing instruction belongs in the ChatGPT Image 2.0 storyboard prompt because it helps plan the sequence. Select the actual video duration, aspect ratio, and resolution using the MiniMax H3 controls on RunDiffusion rather than repeating those settings in the H3 video prompt.

RunDiffusion architecture poster with a designer, concept sketches, floor plans, and finished building renderings.

Use a Character Sheet, Product Sheet, or Building Sheet as a Reference Image

Sheets are a practical way to organize character identities, outfits, props, expressions, or multiple views inside one reference image. You can use Nano Banana Pro or ChatGPT Image 2.0 to create them.

On RunDiffusion, upload the sheet through MiniMax H3 Ref. to Video. Use the Start Frame field when the uploaded image should be the literal opening composition; use the reference workflow when the sheet should guide the generated scenes.

A useful character sheet can include:

  • A labeled cast row with one clear full-body view per character
  • Front, three-quarter, profile, or full-body views of one important character
  • Wardrobe, equipment, product, or prop details
  • Expression references
  • A clearly separated storyboard with shot order, actions, and camera framing

Keep the layout easy to read. Large, sharp subjects and clearly separated sections give H3 stronger visual information than a crowded page of tiny panels.

Let a Detailed Sheet Carry the Detail

When a reference sheet already contains labels, character designs, shot order, actions, camera suggestions, and timing, the prompt does not need to rewrite every panel. Give each region one clear job and let the image carry the specifics.

This concise prompt was used with a labeled cast-and-storyboard sheet in the current RunDiffusion reference workflow:

The upper cast row defines four recurring professionals: the
architect in black, construction engineer in safety gear, real
estate agent in a cream suit, and site manager in workwear with a
white hard hat. Maintain their facial identity, clothing, and
professional appearance.

Use the lower storyboard as visual guidance for the complete
architectural journey.
Four-person architecture team reference sheet above an eight-panel storyboard from empty lot to completed modern home.

The resulting 15-second video followed the progression from site survey and planning through construction, walkthrough, and final reveal while maintaining the professional roles. The image supplied the detailed sequence; the prompt simply explained how to interpret its two main regions. The reference sheet remained visible for roughly the first half-second before dissolving into the opening scene, so a small opening trim may be useful when a completely clean start is required.

Describe the Features That Must Stay Consistent

Instead of writing use the same woman, name a few recognizable details:

Preserve her oval face, shoulder-length copper curls, dark brown
eyes, small mole below the left eye, athletic build, and
forest-green jacket with brass buttons.

Focus on the details that make the character easy to recognize. You do not need to repeat every feature in every shot.

Understand Pictures and Subjects

The official advanced prompt format separates a source file from the content inside it:

  • <Picture 1> means the uploaded image itself, often used as a first frame, last frame, keyframe, composition guide, or storyboard.
  • <Subject 1> means a reusable person, object, outfit, location, action, pose, or style taken from a reference.

In simple terms: the picture is the file; the subject is the thing you want to reuse.

For example:

<Subject 1> is the woman shown in <Picture 1>. Preserve her facial
identity, copper curls, dark brown eyes, and green jacket.

You only need this label-based format when a project has enough references that plain language becomes difficult to track.

Can MiniMax H3 Use a Storyboard?

Yes. MiniMax's official full-reference guide says a reference image can be used as a storyboard or shot-planning guide.

Upload the storyboard as an image, then state what it should control:

Image 4 is a storyboard reference for Shots 1, 2, and 3. Use it only
to guide the shot order, camera viewpoint, subject placement, and
approximate framing. Render the final video in the visual style
described in the prompt.

If the storyboard uses several separate images, give each image a shot number:

Image 4 guides Shot 1, the wide opening view.
Image 5 guides Shot 2, the close-up of the product.
Image 6 guides Shot 3, the final character reaction.

The storyboard is guidance, not a guarantee that every line or panel will be reproduced exactly. Keep each shot achievable within the selected clip length. Two or three strong panels are usually easiest to control, but a clearly labeled multi-panel sheet can also guide a fast 15-second montage when the prompt assigns the storyboard one clear role.

Start with the concise region-mapping approach when the storyboard is already detailed. Add shot-by-shot text only when a panel is ambiguous or a specific action needs more control.

Storyboard Reference Examples from RunDiffusion

These two 15-second MiniMax H3 Ref. to Video results show how one organized image can define both the recurring character or cast and the order of events. Treat the storyboard as strong visual guidance rather than a promise that every panel will be reproduced frame for frame.

Example 1: Aric Vale Character and Adventure Storyboard

The upper section defines Aric Vale's face, build, clothing, and equipment. The lower storyboard guides a complete adventure sequence through the jungle and ruins. A concise prompt for this type of sheet is:

Use the upper character sheet to define Aric Vale's face, dark
tousled hair, stubble, athletic build, clothing, and gear. Maintain
his identity and outfit throughout.

Use the lower storyboard as visual guidance for the adventure
sequence, from spotting the ruins and planning the route through
climbing, exploring, claiming the artifact, and escaping. Keep the
character, environment, and cinematic style consistent.
Aric Vale character sheet above an eight-panel adventure storyboard set in jungle ruins.
The upper section defines Aric while the lower storyboard supplies the action sequence.

Example 2: Recurring Cast and Architectural Storyboard

The upper row defines four recurring professionals, while the lower storyboard moves from the empty site through planning and construction to the finished home. The short region-mapping prompt shown earlier lets the reference image carry the detailed shot plan.

Character reference sheet with four construction professionals above an eight-panel architectural storyboard.
A labeled cast sheet can define recurring roles while the storyboard guides the project sequence.

Both examples use the same beginner-friendly rule: explain what the upper and lower sections control, then let the labeled reference sheet provide the smaller details.

Voice Referencing Through a Reference Video

On RunDiffusion, voice referencing can come from clear audio contained in an uploaded reference video. H3 can use the vocal timbre and delivery heard in that video while generating the new dialogue written in your target prompt.

We use the term voice referencing because it accurately describes the creative control without promising an exact one-to-one copy of a person's voice.

The audio inside a reference video can guide:

  • Vocal timbre and delivery
  • Accent, cadence, pace, or intensity
  • Emotional performance
  • Dialogue rhythm
  • Music, ambience, or sound-effect direction

Prepare a Clear Reference Video

When voice is important, choose a reference video with:

  • One clearly audible speaker
  • A voice-forward mix with minimal competing music
  • Minimal room echo
  • Clean sound without clipping or heavy distortion
  • A delivery style close to the performance you want

A clean source gives H3 less ambiguity about which vocal qualities matter. Only use a person's voice or likeness when you have the necessary permission.

Connect the Voice to the Correct Character

State which character should use the voice heard in the reference video:

Video 1 provides the woman's measured walking pace and slow camera
movement. The voice heard in Video 1 guides the warm vocal timbre
and calm delivery of the woman defined by Image 1 for the new
dialogue written below.

For a multi-character scene, identify speakers consistently:

The woman is Speaker 1. The man is Speaker 2. The voice heard in
Video 1 guides Speaker 1.

In the advanced format, speaker IDs are written as (S1), (S2), and so on:

<Video 1> provides the walking pace, camera movement, and vocal
delivery used by <Subject 1> (S1).

Write Dialogue Exactly

Keep spoken lines short enough to fit naturally inside the clip. Write the exact sentence and identify the language when using the advanced format:

<Subject 1> (S1) looks toward the camera and says: <d>[English] Take
the morning with you.</d>

Allow time for breathing, facial reaction, and physical action. A five-second shot usually cannot support a long paragraph of dialogue.

Tell H3 What to Take From the Video

A reference video can provide visual and audible direction at the same time. State which parts should guide the result: the subject's action, camera movement, cuts, pacing, vocal delivery, rhythm, or sound.

Video 1 supplies the measured walking pace, slow half-orbit camera
movement, and warm vocal delivery. Image 1 defines the woman's
identity and clothing, while Image 2 defines the rooftop setting.

This keeps every reference focused on a clear role and helps H3 combine them without guessing.

How to Use Reference Video

A reference video can guide motion, camera behavior, cuts, rhythm, timing, voice, or sound. Add it through the video upload control in MiniMax H3 Ref. to Video, identify it as Video 1, and state what it should control.

Video 1 provides the woman's measured walking pace and slow
half-orbit camera movement. Image 1 defines the character and
clothing, while Image 2 defines the location.

Choose a clip with one readable action or one clear camera idea. A busy reference with several people, cuts, and movements gives H3 more signals to separate.

If you want to reuse a visible action from the reference video, describe that action as a reusable subject:

Transfer the controlled two-step turn from Video 1 to the woman in
Image 1 while preserving her identity and clothing.

Write the Action in Playback Order

AI video prompts work better when they describe visible actions instead of broad emotions.

This is vague:

A woman feels confident in a premium city campaign.

This is easier to generate:

The woman steps out of the elevator, straightens her cuff, looks
toward the sunrise through the glass wall, and walks past the camera
with a restrained smile.

The second prompt gives H3 a sequence it can place on a timeline.

Keep the Shot Count Realistic

For a short clip, begin with one continuous shot. Add a cut only when it reveals something new, such as a product detail, a different viewpoint, or a character reaction.

The official structured format does not timestamp the first shot. Later shots receive a cut time:

[Shot 1] A medium-wide view establishes the station platform.
[Shot 2] At 00:04.500, the camera cuts to a close-up of the ticket
in her hand.

Keep every cut time inside the selected video duration.

Use One Main Camera Idea Per Shot

Choose a clear movement:

  • Push in or pull out: The camera moves toward or away from the subject.
  • Pan or tilt: The camera turns horizontally or vertically from one position.
  • Truck or pedestal: The full camera moves sideways or vertically.
  • Arc: The camera travels around the subject.
  • Tracking: The camera follows a moving subject.
  • Static: The camera remains still.

Add speed or range only when it helps:

The camera tracks right at slow speed, keeping the woman centered
while the station columns move through the foreground.

Avoid combining conflicting directions such as static camera, handheld shake, and fast orbit in the same moment.

Prompt First and Last Frames as a Transition

When you provide both an opening and ending frame, describe how the scene moves between them.

Use this pattern:

opening state -> action begins -> visible intermediate changes ->
exact ending state

If the first frame shows a closed umbrella and the last frame shows it open, explain how the hand lifts the umbrella, the runner slides upward, the ribs spread, the canopy catches the rain, and the character settles into the final pose.

A continuous shot is often the easiest way to create a smooth bridge. Multiple cuts can work, but they add complexity when the real goal is a natural transition.

Direct the Sound With the Video

H3 generates picture and sound together. Treat sound as part of the scene instead of adding it as an afterthought.

Dialogue

Write the exact words and identify who speaks them. Keep the line short enough for the available time.

Scene Sound

Describe physical sounds near the action that creates them:

The ceramic cup touches the saucer with a light click as the
espresso machine releases a short burst of steam.

Overall Soundscape

Use this for continuing ambience and physical sounds across the clip:

overall_soundscape:
Low café room tone continues beneath occasional cup clinks, soft
footsteps, and rain against the windows.

Do not repeat spoken dialogue in this section.

Background Music

The official format calls audience-only music non_diegetic_music. This means the viewer hears it, but the characters do not.

non_diegetic_music:
Sparse muted piano at a slow tempo, joined by a sustained low cello
note that fades during the final second.

If you do not want background music, write:

non_diegetic_music:
N/A

Do not set the entire soundscape to N/A unless you want complete silence with no dialogue, ambience, or physical sound.

A Beginner-Friendly First-Frame Prompt

This example uses one image as the opening frame. Select 8 seconds and your preferred aspect ratio in RunDiffusion, then enter:

Cinematic live-action video that begins exactly from Image 1.

Preserve the woman's face, charcoal coat, red umbrella, wet street,
storefront reflections, and evening lighting from the image.

The camera holds still for the first second as she looks up from a
folded paper map. She closes the map, steps around a puddle, and
walks toward the warm storefront entrance. The camera then tracks
slowly to the right. At the doorway, she reaches for the brass
handle and gives a small relieved smile. The door opens and warm
light crosses her face. Use one continuous shot.

Sound: steady rain on the umbrella and pavement, soft wet footsteps,
the paper map folding, and a quiet door creak.

Why it works:

  • It says the image is the exact opening frame.
  • It protects the important visual details.
  • It gives the character a simple sequence of actions.
  • It uses one controlled camera move.
  • It separates scene sound from background music.

The Advanced H3 Prompt Structure

Most beginners do not need this structure for their first generation. Use it when you have several images, a reference video, multiple characters, voice referencing through video, or a storyboard that needs careful tracking.

MiniMax publishes two structured formats in its official base prompt guide and full-reference prompt guide.

Three-Part Structure for Text and Keyframes

integrated_multimodal_description:
[Shot 1] Describe the style, composition, character, action, camera,
dialogue, and synchronized sounds in playback order.

overall_soundscape:
Describe ambience, physical sounds, and non-verbal human sounds
across the clip.

non_diegetic_music:
Describe audience-only music, or write N/A.

For first-frame, last-frame, or first-and-last-frame work, MiniMax's official format also adds a frame-alignment instruction before these three sections.

Six-Part Structure for Complex References

subject_definitions:
Define each reusable person, object, environment, picture, and video
reference.

summary:
State the target video and explain how the references work together.

retention_analysis:
State what must be preserved, transferred, copied, or used as loose
guidance.

detailed_description:
Describe the final video shot by shot in playback order.

overall_soundscape:
Describe ambience and physical sound.

non_diegetic_music:
Describe audience-only music, or write N/A.

For reference-generation tasks, MiniMax says the detailed_description section is normally about 350 to 500 English words. That is guidance for complex reference work, not a reason to make every H3 prompt long. Extra detail only helps when it removes ambiguity.

Advanced Character, Motion, and Voice-Referencing Example

This example combines character images, a location image, a product image, and a reference video containing clear spoken audio. Use fictional subjects or people whose likeness and voice you have permission to use. Select 10 seconds and 9:16 in RunDiffusion before entering the prompt.

subject_definitions:
<Subject 1> is the adult woman whose facial identity and
shoulder-length copper curls come from <Picture 1>, and whose
forest-green jacket and cream shirt come from <Picture 2>. Preserve
her oval face, dark brown eyes, small mole below the left eye,
hairstyle, proportions, and clothing.
<Subject 2> is the matte-black travel cup in <Picture 3>, preserving
its tapered shape, brushed finish, silver rim, and small white
wordmark.
<Subject 3> is the rooftop garden in <Picture 4>, preserving its
pale concrete paths, tall grasses, glass rail, and sunrise lighting.
<Video 1> provides the measured walking pace, slow half-orbit camera
movement, and the warm vocal timbre and measured delivery heard in
its audio for <Subject 1> (S1). The character, clothing, and
location come from <Subject 1> and <Subject 3>.

summary:
[reference generation] Create a polished live-action product video.
<Subject 1> walks through <Subject 3> while carrying <Subject 2>.
Her movement and the camera path follow <Video 1>. Her new spoken
line uses the vocal qualities heard in <Video 1>.

retention_analysis:
<Subject 1>: fully_preserved - retain her facial identity, copper
curls, mole, proportions, green jacket, and cream shirt.
<Subject 2>: fully_preserved - retain the cup shape, finish, silver
rim, and wordmark.
<Subject 3>: fully_preserved - retain the rooftop layout, grasses,
glass rail, and sunrise direction.
<Video 1>: attribute_transfer - transfer the walking pace,
half-orbit camera path, vocal timbre, and measured delivery.

detailed_description:
The target video uses polished live-action commercial realism with
warm sunrise light and restrained color contrast.

[Shot 1] A medium-wide view opens in <Subject 3>. <Subject 1> enters
from the lower left carrying <Subject 2> at waist height. Her face,
copper curls, green jacket, cream shirt, and proportions remain
consistent with the reference images. She follows the relaxed
walking pace from <Video 1>, taking three clear steps along the pale
concrete path while the grasses move lightly in the wind. The camera
begins the slow half-orbit from <Video 1>, moving from her
front-left toward her side while keeping her face and the cup
visible. She glances down at the cup, turns it so the wordmark faces
the camera, and raises it to chest height.

[Shot 2] At 00:05.500, the camera cuts to a stable chest-up view.
<Subject 1> (S1) looks toward the camera and says with natural lip
movement: <d>[English] Take the morning with you.</d> Her delivery
follows the warm timbre and measured pace heard in <Video 1>. She
closes her lips at the end of the sentence, gives a restrained
smile, and looks toward the sunrise. The camera pushes in slowly
while <Subject 2> remains visible in the lower foreground. End with
her face, the cup, and the sunrise stable and in focus. Keep the
scene limited to <Subject 1>, <Subject 2>, and the established
rooftop environment.

overall_soundscape:
Light wind moves through the grasses while measured footsteps cross
the concrete. The cup makes one soft contact sound against her
clothing as it is raised. The dialogue remains clear above ambience
generated from this soundscape.

non_diegetic_music:
N/A
RunDiffusion Board with ChatGPT Image 2.0, MiniMax H3, and MiniMax H3 Ref. to Video tools added.

Common MiniMax H3 Prompting Problems

  • The character changes between shotsLikely cause: The references conflict or the important identity features were never stated Simple fix: Use a smaller, compatible character pack and name a few stable features
  • The opening image changesLikely cause: The image was treated as a loose reference instead of an exact first frame Simple fix: Choose first-frame mode and say the video begins exactly from that image
  • A character sheet dominates the openingLikely cause: The sheet was added as a Start Frame, or the reference remains visible during the opening transition Simple fix: Confirm MiniMax H3 Ref. to Video was used; trim a brief opening transition when the rest of the result is correct
  • A detailed storyboard is skipped or mergedLikely cause: The prompt competes with the visual plan, or the selected duration is too short Simple fix: Assign the storyboard one clear role, shorten the prompt, and select a duration that fits the sequence
  • H3 copies the wrong part of an imageLikely cause: The image contains several useful elements but no role was assigned Simple fix: State whether the image controls identity, clothing, product, setting, style, or composition
  • The motion does not match the referenceLikely cause: The source video is busy or its purpose is unclear Simple fix: Use a cleaner motion clip and name the exact action or camera move to transfer
  • The wrong character uses the voiceLikely cause: The voice heard in the reference video was not connected to a speaker Simple fix: Map the voice from Video 1 to one character and keep the same speaker ID
  • Words from the reference video appear in the resultLikely cause: The prompt did not separate vocal qualities from the source dialogue Simple fix: State that Video 1 supplies vocal timbre and delivery, then provide the new target dialogue explicitly
  • Music appears when you only wanted ambienceLikely cause: Music and scene sound were not separated Simple fix: Describe physical sound in overall_soundscape and set non_diegetic_music: N/A
  • The clip feels rushedLikely cause: Too many actions, cuts, or spoken words were packed into a short duration Simple fix: Reduce the video to one main action, one reaction, and one or two camera decisions
  • The first-to-last-frame transition warpsLikely cause: Only the two endpoints were described Simple fix: Explain the visible motion that connects the opening and ending frames
  • The product or logo changesLikely cause: Style instructions overpower the preservation request Simple fix: Define the product separately and protect its shape, material, color, and visible text

A Better H3 Testing Workflow

Do not try to solve character identity, complex motion, several cuts, dialogue, product detail, and music in the first generation. Add control in stages.

Pass 1: Prove the Character and Scene

Use one character image, one simple action, and a static or slow camera. Skip dialogue. Check the face, clothing, proportions, and environment.

Pass 2: Add Motion or Movement

Add one motion instruction or one reference video. Test a single action or camera path before combining several movements.

Pass 3: Add Voice and Sound

Use a reference video with clear spoken audio, connect its voice to the correct character, add one short line, and separate scene sound from background music. Check timing, lip movement, pronunciation, and clarity.

Pass 4: Add Cuts, a Storyboard, or a Second Character

Introduce one new source of complexity at a time. Keep your image numbers, subject labels, and speaker IDs consistent.

Pass 5: Generate the Final Version

Once the direction works, move to the final resolution or regeneration option supported by the current RunDiffusion workflow. Higher resolution can improve presentation quality, but it cannot repair a confusing prompt or conflicting reference pack.

This staged process fits a broader AI video production workflow on RunDiffusion, where concept testing, targeted retakes, enhancement, and delivery are separate production decisions.

Using MiniMax H3 on RunDiffusion

RunDiffusion handles the model environment so you can focus on the prompt, references, and result.

Add MiniMax H3 to a Board

These screenshots walk you from the RunDiffusion homepage to your first H3 generation.

1. Log in to RunDiffusion

Open RunDiffusion and select Log in in the top navigation.

RunDiffusion homepage with Log in and Free Signup buttons in the top navigation.

2. Open Boards

After logging in, select Boards in the left sidebar. Boards give you a visual workspace where you can add and run creative tools.

RunDiffusion dashboard with Boards and Open-Source Apps visible in the left sidebar.

3. Create a new Board

On the Boards page, click New Board.

RunDiffusion Boards page with the New Board tile ready to create a workspace.

4. Start from scratch

Choose From Scratch. This gives you an empty Board so you can add only the tools needed for your H3 workflow.

RunDiffusion New Board menu with From Scratch selected above the template option.

5. Set up the Board

Choose a cover, enter a clear Board title, and add an optional description. Then click Create.

RunDiffusion Add Board window with cover image, title, description, and Create controls.

6. Add a tool

Inside the empty Board, click Add another Tool. This is the exact label shown when you add the first tool to a new Board.

Empty RunDiffusion Board with the Add another Tool button at the bottom.

7. Search for MiniMax H3

In the Choose Tool window, open the Tools tab and search for H3. Then select the workflow that matches your starting point:

RunDiffusion tool picker showing MiniMax H3 and MiniMax H3 Ref. to Video search results.

8. Enter a prompt and run the tool

Enter your prompt in the Prompt field, which currently accepts up to 2,000 characters. Add a Start Frame when an image should be the exact opening composition. Use MiniMax H3 Ref. to Video when reference images or reference videos should guide the generated scenes. Add images through the image upload control and videos through the video upload control, then identify them in the prompt as Image 1, Image 2, Video 1, and so on. A reference video can provide motion and camera guidance while its contained audio can also guide voice and sound. Then choose from the duration, resolution, aspect ratio, and other options currently available in the tool. Click Run when you are ready.

MiniMax H3 tool panel with Prompt, Start Frame, Video Resolution, Aspect Ratio, and Run controls.

For your first test, use one subject, one clear action, and one main camera move. Add references, dialogue, extra actions, and cuts one at a time after the basic result works. Give every uploaded file one clear job in the prompt, and keep reference numbering consistent with the order shown in the tool.

The current RunDiffusion H3 tool cards show 5-to-15-second video generation and 768P or 2K output where offered. Available controls can evolve, so follow the choices visible in your Board.

The full official H3 structure can improve organization when a prompt becomes complex, but the principle stays the same inside RunDiffusion: identify each reference, explain what it controls, protect what must remain consistent, and describe the video in playback order.

You can also explore more RunDiffusion prompting guides for practical help with image and video workflows.

MiniMax H3 Prompting Checklist

Before generating, ask:

  • Did I choose the correct generation mode?
  • Did I select duration, aspect ratio, and resolution in the RunDiffusion controls instead of writing them as prompt instructions?
  • Did I use MiniMax H3 Ref. to Video for reference images, a character sheet, a storyboard, or a reference video?
  • Is my prompt within RunDiffusion's 2,000-character limit?
  • Does every uploaded reference have one clear job?
  • Are my character images sharp and consistent with each other?
  • Did I name the identity, clothing, product, or setting details that must not change?
  • Is the action written in the order it happens?
  • Is the amount of action realistic for the duration selected in RunDiffusion?
  • Does each shot have one main camera idea?
  • If I use a storyboard, did I map each panel to the correct shot?
  • If I use voice referencing, did I connect the voice heard in Video 1 to the correct character?
  • Did I say whether each reference video controls motion, camera behavior, timing, voice, sound, or a combination of those elements?
  • Are dialogue, scene sounds, ambience, and background music clearly separated?
  • Am I testing the idea before moving to final-resolution output?

Final Thoughts

MiniMax H3 is powerful because a prompt and reference image can work together. A detailed character sheet may carry identities, wardrobe, props, storyboards, and camera framing, while a concise prompt explains how those regions should guide the video.

Start simple. Give every reference one job. Describe the action in order. Add advanced structure only when the project needs it.

That approach gives H3 a clearer video to build and gives you a faster path from the first test to a polished result.

Try MiniMax H3 Ref. to Video on RunDiffusion