AIREITER

AI Video Generator Prompt Adherence: 6 Models Tested

Last Updated: 2026-07-30 04:39:23

Veo 3.1's prompt adherence is scored 7.8 by one public leaderboard and 9.2 by another, and neither publishes the prompt it used. So we wrote two prompts with twenty individually checkable elements and pushed them through six models on 2026-07-30. The models turned out closer to each other than those leaderboards suggest. What separates them is the kind of instruction, not the brand.

Which AI video model follows prompts best

Seedance 2.0 Fast followed the most instructions here: 19 of 20 checkable elements, and a clean 10/10 on the harder prompt. Kling 3 Turbo and Veo 3.1 tied at 18, Wan 2.7 took 17.5, Grok Imagine 16, and Hailuo 02 Pro came last at 15.5 after ignoring an entire event sequence. Three repeat runs on the harder prompt left that order unchanged. Sora 2 could not be tested at all.

ModelLiteral promptStress promptTotal /20Cost per clipPer secondWait
Seedance 2.0 Fast9.010.019.0$0.83$0.165121–132s
Kling 3 Turbo9.58.518.0$0.45$0.09045–54s
Veo 3.19.09.018.0$1.25$0.15688–95s
Wan 2.79.08.517.5$0.40$0.080103–104s
Grok Imagine8.08.016.0$0.09$0.01543–49s
Hailuo 02 Pro9.06.515.5$0.29$0.049166–185s
Sora 2—————5 failed submits

Prices come from published credit rates — Kie's pricing table for Kling, Seedance, Wan, Hailuo and Grok, the AIReiter model pages for Veo 3.1 and Sora 2 — divided by the clip length each model actually returned, tabulated below.

Pick by what your shot needs:

  • Prompt contains a count, an ordering or a "never moves" → Seedance 2.0 Fast. It was the only model to get all ten of those elements right, which is what its price premium buys.
  • Iterating on a look, not a sequence → Kling 3 Turbo. Nearly the same obedience at roughly half the price and a third of the wait.
  • Checking whether a prompt reads at all → Grok Imagine, at $0.015 per second. Its 16/20 is mostly one missing entrance.
  • Sequenced action → not Hailuo 02 Pro, which skipped an entire event.

Sora 2 returned UPSTREAM_BUSY / Upstream service is busy or upgrading on all five submissions between 03:57 and 03:59 UTC on 2026-07-30, across both prompts and three spaced retries — no charge, no clip, no score.

How we tested: two prompts, twenty checkable elements

Every clause in both prompts maps to something you can point at in the frame. No "cinematic", no "8k", no "moody" — those cannot be scored. Each of the ten elements per prompt is marked ✓ (1 point), ~ (0.5, partially there) or ✗ (0). Nothing here scores image quality or temporal consistency: a clip can be beautiful and stable while quietly dropping half your instructions. Grok Imagine's output reads as a 3D render rather than photography, and that cost it nothing.

The literal prompt tests whether a model can assemble a described scene. Its ten elements are the ten things it names — one red vintage bicycle, leaning on the yellow brick wall, a white cat with one black ear, entering from frame right, stopping at the front wheel and looking up, the sign reading OPEN 7AM, a wide-to-medium dolly, camera at ground level, overcast light with no sun, nobody in shot.

A single red vintage bicycle leans against a yellow brick wall on an empty
street at dawn. A white cat with one black ear walks in from the right side of
the frame, stops beside the front wheel, and looks up. A wooden sign above the
bicycle reads "OPEN 7AM". The camera dollies in slowly from a wide shot to a
medium shot, staying at ground level. Overcast morning light, no sun, no people
anywhere in the shot.

The stress prompt tests counting, event order and negation, and its ten elements are: exactly three ducks, all three staying the same size, a row on a wooden table, no hands or people, the middle duck falling first, the left duck falling after it, the right duck never moving, a locked-off camera, a plain white background, hard light from overhead.

Three identical yellow rubber ducks sit in a row on a wooden table. No hands and
no people appear at any point. First the middle duck tips over onto its side;
about two seconds later the duck on the left tips over. The duck on the right
never moves. The camera is locked off: no zoom, no pan, no push in. Plain white
background, hard light from directly above.

The switch that rewrites your prompt before the model sees it

Two of these models ship with a prompt rewriter enabled by default: prompt_extend on Wan 2.7 and prompt_optimizer on Hailuo 02 Pro, both documented as defaulting to true. Submit a prompt without touching them and the text scored is not the text you wrote. Both were set to false here; with defaults on, a test measures the rewriter as well as the model.

Other parameters do not stick at all. Every clip was requested at 5 seconds, 720p, 16:9, but three models overrode that, so the per-second prices above use what came back rather than what was asked for:

ModelSentReturnedCredits billed
Seedance 2.0 Fast5s, 720p, 16:95.0s, 1280×720165 (33/s)
Kling 3 Turbo5s, 720p, 16:95.0s, 1280×72090 (18/s)
Veo 3.15s, 16:98.0s, 1280×720125 per call
Wan 2.75s, 720p, 16:95.0s, 1280×72080 (16/s)
Grok Imagine6s (floor), 720p, 16:96.0s, 1280×72018 (3/s)
Hailuo 02 Pro5s, 720p, 16:95.9s, 1920×108057

Hailuo 02 Pro's only documented inputs are the prompt and two switches, so there is nothing to set; it returned 1080p both times. Veo 3.1 returned 8 seconds regardless of what we sent, and Grok Imagine documents a 6-second floor.

What all six got right

Coarse composition was not where any of these six outputs failed. All six models put exactly one red vintage bicycle against a yellow brick wall, and all six rendered the on-screen text "OPEN 7AM" without a single wrong character. All six also honoured both negative instructions — no people appeared in any of the six street clips, and no camera moved in any of the six locked-off duck clips.

Six AI video models given the same bicycle and cat prompt, frame taken at 68 percent of each clip

Where they break: four kinds of instruction

Counting and attribute detail

"A white cat with one black ear" was followed exactly once out of six. Hailuo 02 Pro produced the only cat with a single black ear on an otherwise white body. Kling 3 Turbo and Veo 3.1 got the black ear but added a black tail; Wan 2.7 spread a black patch across both ears; Grok Imagine painted a full Siamese mask; Seedance 2.0 Fast rendered an almost entirely white cat with a faint dark smudge on one ear tip.

Cat head detail from six models, showing how each handled the instruction one black ear

Every model understood "white cat with a black marking". None except Hailuo treated "one" as a number. In these six outputs, attribute counts were markedly less reliable than coarse composition.

Event order and timing

The duck prompt asked for a specific sequence — middle falls, then left, right never moves. Four models executed it correctly: Kling 3 Turbo at 2.0s then 4.9s, Seedance 2.0 Fast at 2.0s then 3.5s, Wan 2.7 at 1.0s then 4.0s, Veo 3.1 at 1.6s then 4.8s. None hit the requested two-second gap precisely, and the right-hand duck stayed put in all six clips.

Hailuo 02 Pro did not do it at all. The middle duck never fell. Instead the left duck rotated to a profile view and swelled to nearly double size while all three stayed upright.

Duck sequence test across six models at start, mid-clip and final frame

Grok Imagine tipped the right ducks in the right order but dropped them forward onto their beaks rather than onto their sides. Veo 3.1 got the order right and then let both fallen ducks stand back up by the final frame — the event fired, the state did not hold.

Camera discipline

Locked-off cameras were universally respected; described camera moves were not. Wan 2.7 and Veo 3.1 both started the requested wide-to-medium dolly and then kept pushing into an extreme close-up, Wan ending on a bare wheel rim with the cat out of frame entirely. Grok Imagine moved so little that the push-in is barely detectable. Hailuo 02 Pro was the only model to break "staying at ground level", shooting from roughly hip height while the other five sat low.

Object consistency

Objects change size when they move. Hailuo's left duck grew to almost twice its starting size, Grok Imagine's two fallen ducks both enlarged noticeably, and Kling 3 Turbo's and Veo 3.1's tipped ducks grew slightly. Only Seedance 2.0 Fast and Wan 2.7 kept all three ducks the same size from first frame to last.

Size consistency goes unprompted, and then unnoticed.

Elements followed out of ten by model, for the literal prompt and the stress prompt

Five of six models cluster at 8.0–9.5 on the static scene and only diverge on counts and ordered events, where Hailuo falls 2.5 points. Model choice buys you little on a straightforward shot and a lot the moment a prompt contains a number.

Does the ranking hold on a rerun?

Yes, for the three models we reran. Each got two more attempts at the stress prompt, scored the same way, and the score ranges do not overlap.

ModelRun 1Run 2Run 3MedianRange
Seedance 2.0 Fast10.09.510.010.09.5–10.0
Kling 3 Turbo8.58.08.08.08.0–8.5
Hailuo 02 Pro6.56.5—6.5n=2

Kling 3 Turbo painted a blue-grey wall instead of the requested plain white background in all three runs, which is a stable preference rather than sampling noise. Hailuo 02 Pro reproduced its exact failure both times: the middle duck never fell and the left duck swelled to nearly double size. Hailuo's third run was blocked by a daily credit cap, so its median rests on two samples.

Repeat runs of the stress prompt for Seedance, Kling and Hailuo

One new behaviour showed up in reruns that single passes had hidden. Kling 3 Turbo's middle duck fell and then stood back up before the final frame in run 2, the same state-reversion Veo 3.1 showed on its only run. Events fire; states do not always hold.

What a retry costs when a model ignores you

Price and speed do not line up. Hailuo 02 Pro was the second-cheapest model per clip and by far the slowest at 166–185 seconds, while Kling 3 Turbo delivered its 18/20 in 45–54 seconds. Per-clip prices also flatter the short-clip models: Veo 3.1 costs $1.25 against Seedance 2.0 Fast's $0.83, but Veo returns 8 seconds and Seedance 5, so per second the two are nearly identical at $0.156 and $0.165.

Kie pricing table showing Seedance 2.0 Fast at 33 credits per second

The 17 clips behind this article cost 1,387 Kie credits plus 250 AIReiter credits, which at 200 and 100 credits per dollar is $9.44; the five failed Sora 2 submissions cost nothing. Generating is the cheap part. Generating three times, because element four went missing and no checklist caught it, is not.

As one r/SoraAi poster put it: "No matter how clear, detailed, or restrictive I make the prompt, SORA consistently ignores basic visual instructions." Element-level scoring turns that into a list of which instructions, which is something you can act on.

Run this test on your own shot list

  1. Rewrite one real prompt so every clause is checkable — replace "cinematic" with the camera height, the light direction and the frame the move ends on.
  2. Split it into a numbered element list before you generate. Ten elements is enough; write them down or you will grade from memory and score generously.
  3. Turn off the rewriters. Set prompt_extend to false on Wan, prompt_optimizer to false on Hailuo, and check whether your platform has its own before you compare anything.
  4. Send the identical string to two or three models at matched duration and resolution, and note what each one silently overrides.
  5. Review the full clip, using the first, middle and final frames as checkpoints — brief state failures happen between samples. Then rerun only the model that missed the elements you actually care about. A missing black ear does not matter; a duck that stands back up does.

If you want to reproduce our runs directly, the model pages for Veo 3.1, Seedance 2 Fast and Sora 2 list the exact model IDs and per-second credit prices used above.

FAQ

Which AI video generator is most accurate?

In this test, Seedance 2.0 Fast, at 19 of 20 checkable elements and the only clean pass on counting, ordering and negation. On straightforward scenes accuracy was near-identical across five of six models, so the ranking only separates them on sequenced prompts.

Is Seedance better than Veo 3.1 or Sora 2?

Seedance 2.0 Fast outscored Veo 3.1 by one point, 19 to 18, at two-thirds the price per clip, and both executed the duck sequence in the correct order. We cannot rank Sora 2 — five submissions on 2026-07-30 all failed upstream, so it has no score here rather than a low one.

Is Kling 3 Turbo worth it for prompt-heavy work?

Yes, when speed matters more than sequencing. Kling 3 Turbo scored highest of all six on the literal prompt at 9.5/10 and returned in 45–54 seconds, but it lost a point on the stress prompt for a blue-grey wall where the prompt asked for plain white.

Does a longer prompt improve adherence?

Not by itself. The literal prompt in this test is longer than the stress prompt and every model scored higher on it, because its clauses describe a static arrangement rather than counts and event order.

The part this test cannot answer

Three runs are enough to show that these scores are not sampling noise, and not enough to make them a benchmark. Only Seedance, Kling and Hailuo were rerun; Veo 3.1, Wan 2.7 and Grok Imagine still rest on one attempt each. Both prompts are also single-shot, dialogue-free and under ten seconds, which leaves the instruction types most likely to break a real edit — multi-shot continuity, spoken lines, a named character reappearing — completely untested. The ranking above tells you who follows a described shot. Whether it survives a shot list is the next test, and it is a different one.

Related reading: Seedance 2.0 vs Kling 3.0 vs Sora 2 vs Veo 3.1 · AI camera movement prompts, tested