AIREITER

MiniMax H3 Prompt Guide: Official Syntax, Templates, and Fixes

Last Updated: 2026-08-17 06:55:44

A one-line "cinematic" description leaves every gap for the model to fill: omit camera instructions and provider prompting notes say you tend to get a locked-off shot; leave sound fields unspecified and unwanted speech creeps in. Real MiniMax H3 prompting is closer to writing a compact shooting script: the official guides define a fixed three-field format with shot timestamps, stable speaker IDs, and two dedicated sound fields, and the whole script has to fit a roughly 7,000-character prompt budget (provider-dependent) and a 15-second clip ceiling.

MiniMax H3 official launch page on minimax.io

Two Guides, Four Modes: Pick Your H3 Prompt Format First

MiniMax publishes two separate prompt-writing guides inside the MiniMax-H3 repository on Hugging Face, and the two cover different workflows. The base guide covers four generation tasks with no uploaded references; the reference guide takes over as soon as any uploaded image, video, or audio clip plays a role. The four base tasks map to these modes:

ModeInputAlignment instruction required in the prompt
T2VAText onlyNone; start directly with the core fields
I2VAOne first-frame image"Picture 1 is fully referenced at 0.00 seconds and belongs to [Shot 1]"
FL2VAFirst frame + last framePicture 1 at 0.00 seconds, Picture 2 at the exact end time
L2VAOne last-frame imagePicture 1 aligns with the video's end time and the final shot

Both guides end with worked examples, and the MiniMax launch post frames the model the same way: you express relationships among text, image, video, and audio inputs rather than picking tasks from a menu. Hosted endpoints flatten this into three workflows (minimax/h3/text-to-video, image-to-video, and reference-to-video on fal), but the prompt structure underneath stays the same.

The official MiniMax H3 base prompt writing guide on Hugging Face

The Base Prompt, Line by Line

A base-mode MiniMax H3 prompt has exactly three required fields after the alignment instruction, separated by a blank line. Here is the skeleton:

[alignment instruction if I2VA/FL2VA/L2VA]

integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...
  • integrated_multimodal_description carries the timeline: visual style, composition, actions, shot changes, speakers, dialogue, and diegetic sound, meaning everything visible or audible.
  • overall_soundscape summarizes ambience and physical sound in 1–4 sentences, one paragraph, no dialogue.
  • non_diegetic_music describes audience-only score in 1–3 sentences, or N/A if there is none.

A filled example, adapted from the official guide's T2VA case (a bakery scene before sunrise):

integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium-wide
shot of a small bakery interior before sunrise; warm tungsten light, flour dust in
the air. A middle-aged baker (S1), grey apron, rolled sleeves, lifts the security
shutter. The camera pushes in with small amplitude at slow speed as he places
fresh bread on a wooden counter. (S1) says in a warm, mid-pitch voice <d>[English]
First batch of the day.</d> At 00:05.000, the camera cuts to a close-up of his
hands slicing a loaf; (S1)'s words carry over from the previous shot, <scenetrans>
continuing seamlessly across the cut.

overall_soundscape: The rattling shutter, metal trays set down on stone, a doorbell
chime, footsteps on tile, and the crusty sound of a knife slicing bread.

non_diegetic_music: Acoustic guitar and upright bass at a slow tempo, gentle
dynamics fading out before the final shot.

What each line controls, per the official rules:

Prompt elementControlsRule
[Shot 1] Live-action, cinematic.Global styleStyle label opens Shot 1; examples from the guide's list (partial): Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film
A middle-aged baker (S1), grey apron, rolled sleevesSpeaker identityIdentity attributes precede the ID; the ID stays stable across shots
The camera pushes in with small amplitude at slow speedCameraMotion type + amplitude + speed, written as a natural action
<d>[English] First batch of the day.</d>DialogueLanguage tag plus exact words only, verbatim, never translated
At 00:05.000, the camera cuts to...Cut timeStrictly increasing; [Shot 1] itself never gets a timestamp
<scenetrans> continuing seamlessly across the cutAudio continuityMarks dialogue crossing a cut, at both connecting points

Shots, Cuts, and Camera Moves

Shot syntax in MiniMax H3 is positional: [Shot 1] never gets a timestamp, and every later shot starts with a strictly increasing cut time like At 00:03.500,. Approved cut language is short ("the camera cuts to", "the shot transitions to", "the shot switches to"), while cross-dissolves, fades, and wipes should appear only if you explicitly request them. If only the framing changes slightly, the guide says to use camera motion instead of a cut, because a cut should introduce new subject, space, state, or time information.

Write camera movement as a natural English action with up to three dimensions: motion type, amplitude, speed.

DimensionAccepted expressions
Motion typeZoom In/Out, Push In/Pull Out, Pan Left/Right, Truck Left/Right, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, Roll Clockwise/Counterclockwise
Amplitude"with small amplitude", "with large amplitude"
Speed"at slow speed", "at fast speed"

Medium amplitude and normal speed are omitted by convention. Note the difference in meaning: "Zoom In" changes focal length from a fixed position, while "Push In" physically moves the camera forward — the guide keeps the distinction on purpose.

Dialogue and Speaker Labels

Every vocal character in a MiniMax H3 prompt gets a stable ID — (S1), (S2), or a compound like (S1,S2) for synchronized speech — and keeps that ID across every shot. Characters who never speak get no ID at all. A speaker's first appearance should pin down enough identity for stable generation: on-screen or off-screen, age range, gender, vocal pitch, timbre, speaking rate, accent.

The exact syntax for one line:

A woman in her 30s (S2), on-screen, mid-pitch voice, says with quiet anger
<d>[English] You promised me the 15th.</d>

Four documented rules prevent the classic failures. Identity, action, and delivery stay outside <d>; inside <d> goes only the language tag and the exact words, punctuation preserved. Voiceover uses the literal phrase "says in an off-screen voiceover", immediately followed by a statement that the on-screen character's lips remain closed. Dialogue crossing a cut gets <scenetrans> at both connecting points plus a continuity phrase such as "continues seamlessly across the cut" or "carries over from the previous shot". Speech cut off by the video's end gets <cutoff>.

Quotation marks have a single reserved job: any text visible in the video (banners, signs, labels, neon, subtitles) appears in English double quotes, verbatim, untranslatable. Putting spoken dialogue in quotes is the documented reason H3 renders it as burned-in subtitles instead of speech.

The Two Sound Fields

The MiniMax launch post states that all H3 audio output is native stereo, generated with the video, so sound gets two dedicated fields instead of adjectives in the description. overall_soundscape covers ambience and physical sound (wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter) in 1–4 sentences, and N/A is reserved for the explicit request of complete silence. non_diegetic_music covers score the characters cannot hear, described by instrumentation, speed, rhythm, and dynamic changes; abstract mood words ("epic", "emotional") are explicitly discouraged, and N/A is the correct value when you want none.

Diegetic audio such as singing, an in-scene radio, or a street musician does not belong in either field; it goes into integrated_multimodal_description where the timeline lives. The split is not cosmetic: the APIMODELS how-to notes that reliable results require explicit sound constraints, and leaving non_diegetic_music unspecified can let in music you never asked for.

Reference Prompts: Labels and Retention

The reference guide takes over once uploaded assets steer the generation, and it requires six sections in a fixed order: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music. The detailed_description body is normally 350–500 English words for generation tasks, and dialogue-heavy prompts should fit the full spoken timeline even if that bends the word count.

Four label types do the work:

LabelUse
<Subject N>Visible content actually used in the output: a person, product, scene, outfit, style, or action abstracted from any asset
<Picture N>An image acting as a concrete frame anchor — first frame, keyframe, last frame, composition anchor
<Video N>A whole-video relationship: editing source, continuation point, camera or cut structure, rhythm
<Audio N>An audio signal that is copied or referenced: voice timbre, lyrics, beats, sound effects

Numbering is independent per type, so one uploaded video can be <Video 1> while its audio track is <Audio 2>. A speaking subject combines labels with speaker IDs: <Subject 3> (S1).

The summary section opens with a bracketed task prefix such as [reference generation] or [video editing + audio reuse], and a video edit must begin "The target video is an edited version of <Video 1>." Then retention_analysis assigns one marker per label: fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visuals; fully_copy, partially_copy, reference, or weak_reference for audio. Those markers are how you tell H3 that the face must survive while the outfit may change.

A minimal reference-mode skeleton in the documented order:

subject_definitions:
<Subject 1>: the woman from Picture 1, long blonde hair, light-pink shirt
<Picture 1>: first frame of [Shot 1], café window seat
<Audio 1>: voice-timbre reference for <Subject 1> (S1)

summary: [reference generation + audio reference] One short paragraph naming the
task, the target video, and each reference relationship.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved, same face, hair, outfit
<Picture 1> ([Shot 1] first frame): fully_preserved
<Audio 1> (S1 voice): reference, timbre only, no copied words

detailed_description: [1–2 style sentences.] [Shot 1] ... [350–500 words, dialogue
in <d> tags, <scenetrans> across cuts]

overall_soundscape: [Ambience and physical sounds.]

non_diegetic_music: N/A

When H3 Misbehaves: Symptom-to-Fix Table

SymptomLikely prompt causeFix
Gibberish or invented "Sims-speak"Unstructured dialogue, no speaker IDs, or undescribed non-speaking charactersAdd (S1)/(S2) IDs, wrap exact words in <d>[Language] ...</d>, and explicitly describe silent characters as not speaking
Burned-in subtitles on dialogueSpoken words written in quotation marksReserve quotes for visible on-screen text only; move dialogue into <d> tags
Dialogue assigned to the wrong characterSpeaker ID not tied to a described personGive the speaker identity attributes (age, on-screen, voice) at first appearance, keep the ID stable across shots
Score or music you never asked fornon_diegetic_music left vague or missingWrite non_diegetic_music: N/A when you want none: a fix community threads credit for killing unwanted audio
Unplanned cutsCuts implied by scene changes without timestampsUse [Shot N] At 00:03.500 with explicit cut phrases; request "no cuts" wording for single-take shots
Face, outfit, or product driftsReference role never declaredAdd a retention_analysis line with fully_preserved for that label and repeat the label at each appearance

The gibberish and subtitles rows come straight from the official syntax rules; community threads on r/StableDiffusion independently report the same fixes working in practice, including non_diegetic_music: N/A as one of the most-cited fixes for unwanted audio.

Limits and What a Prompt Costs to Test

The generation envelope matters because prompt length is finite and every iteration is billable. Documented limits across the official API and hosted providers:

ParameterValueNote
Clip durationUp to 15 s, whole secondsMinimum is 5 s on fal/EvoLink endpoints; some V2 API docs list 4 s
Resolution768p and 2K, 24 FPS2K is the default output resolution per MiniMax
Reference filesUp to 9 images, 3 videos, 3 audio12 files total; audio can never be the only reference
Prompt length~7,000 charactersfal and V2 API docs; APIMODELS lists 20,480 — check your provider
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptiveImage-to-video follows the input image's ratio

Pricing moves by provider, so treat these as dated snapshots rather than list prices. On fal, a 2K generation cost $0.26 per second as of an August 1 pricing snapshot ($1.30 for 5 seconds, $3.90 for 15), with the first five reference images free and extra images at $0.08 each. On APIMODELS, H3 runs at $0.088 per second ($1.32 for a 15-second clip), charging only successful requests. For context, per-second prices on that same marketplace:

Per-second video generation price comparison on APIMODELS, August 2026

The retry math is the real budget line: a structured 15-second dialogue prompt that lands on the second attempt costs $2.64 on APIMODELS, while an unstructured prompt that takes five attempts to get one usable take costs $6.60. To iterate prompts against a hosted H3 endpoint, AIReiter's MiniMax H3 API page runs the model with current pricing on the page.

Copy-Paste Starting Points

Two base skeletons cover most short-form work. Fill the brackets, keep the field names and blank lines exactly as written:

integrated_multimodal_description: [Shot 1] [style label]. [Composition and
subject with 1–2 identity details]. [One primary action with physically
plausible pacing]. The camera [motion type] with [amplitude] at [speed].
[Character description] (S1) says in [voice description] <d>[Language]
[exact line]</d>. At [MM:SS.mmm], the camera cuts to [new subject or space],
[continuity phrase if dialogue crosses the cut].

overall_soundscape: [Ambience + physical action sounds, 1–4 sentences.]

non_diegetic_music: [Instrumentation, tempo, dynamics] or N/A
Picture 1 is fully referenced at 0.00 seconds and belongs to [Shot 1].

integrated_multimodal_description: [Shot 1] [Style consistent with Picture 1].
The scene, subject, colors, and spatial relationships of Picture 1 remain
consistent. [Action developing forward from the first frame → result or
reaction]. The camera [motion type] with [amplitude] at [speed].

overall_soundscape: [Ambience and action sounds grounded in the image.]

non_diegetic_music: N/A

For proven prompts rather than blanks, three libraries keep source links. ecomimagelab/awesome-minimax-h3-prompts: 58 prompts (39 official, 19 community-tested), each shipped with its downloadable result video under the rule "No video, no entry." imagineVid/Awesome-minimax-h3-prompts-and-skills: 28 verified cases in six workflows, from omni-reference identity continuity to native audio dialogue. The APIMODELS gallery: 226 prompts browsable by use case with clips attached; its one-line rule is "References carry identity, the prompt carries action."

Two shortcuts exist for writing the scripts themselves. Reddit users report solid results from pasting an official guide into an LLM and naming the target mode ("rewrite this idea as T2VA using the base guide"), and a ComfyUI extension, MiniMax H3 Prompt Writer v0.3, builds the structured format inside a node workflow.

FAQ

What is the difference between the MiniMax H3 base and reference prompt guides?

The base guide covers four no-reference modes (T2VA, I2VA, FL2VA, L2VA) built on three core fields; the reference guide adds six mandatory sections and <Subject>/<Picture>/<Video>/<Audio> labels whenever uploaded images, videos, or audio steer the generation.

How do I label speakers in a MiniMax H3 prompt?

Assign stable IDs (S1), (S2) in order of first vocal event, keep the same ID across shots, use (S1,S2) for simultaneous speech, and put only the language tag plus exact words inside <d> tags.

How do I stop gibberish audio in MiniMax H3?

Structure the dialogue with speaker IDs and verbatim <d> tags, describe non-speaking characters as silent, and set non_diegetic_music: N/A when no score is wanted: the documented and community-confirmed fixes.

Should MiniMax H3 dialogue go in quotation marks?

No. Quotation marks are reserved for text visible in the video; spoken dialogue belongs in <d>[Language] ...</d>, or H3 may render it as on-screen subtitles.

How long can a MiniMax H3 prompt be?

Around 7,000 characters on fal and the official V2 API docs, though APIMODELS lists 20,480 — verify the limit for your specific endpoint before writing a 10,000-character script.

What the Syntax Still Can't Fix

Structure raises the hit rate; it does not make generation deterministic. Exact identity, logo, and product fidelity stay probabilistic: community-run API wrappers explicitly refuse to promise frame-level consistency, and the strongest prompt library examples enforce identity with wordy constraint blocks rather than measured guarantees. Inline <tags> for non-verbal sounds are community-experimental and vary by generation. Judge output by cost per approved second rather than per request, and keep the official base and reference guides open in a tab while you write.

Related reading: the MiniMax H3 release overview covers what the model is, and MiniMax H3 vs LTX 2.3 compares it against the cheapest per-second alternative.