The second change is what we hand the model.
Passing a character sheet, a room plate and a wardrobe plate as three separate references asks the model to composite — to take a person from one image, a room from another, and clothes from a third, and integrate them. It often works, but it is three chances to drift and it tells the model nothing about what happens next.
A storyboard tells it the whole arc at once. One image, the character already in the room already wearing the clothes, with the poses of that take laid out as panels in order. Identity, setting and wardrobe arrive pre-integrated, and the movement is shown rather than described.
The constraint is which model can receive it. Only `seedance_2_0` accepts `image_references` alongside a start frame — up to nine images, twelve reference files total. `minimax_h3` accepts references but refuses to combine them with a start frame. `kling3_0`, `wan2_7` and `veo3_1` have no reference channel at all: a start frame is the only image they take.
So there are two routes, and they are worth testing against each other rather than picking blind.
Route 1Storyboard-derived start frame
Recommended
kling3_0 · 2.0 cr/s (std) · 2.5 (pro)
Generate the storyboard, then crop panel one and use it as the start frame. The storyboard never reaches the video model — it is a planning artifact that guarantees the poses are coherent in that specific room, and it gives you a start frame with character, set and wardrobe already integrated.
For Cheapest by a wide margin. 10s at std is 20 credits. Kling has been the reliable engine on this account.
Against The model sees one frame and a paragraph. It is told what happens next rather than shown it.
Route 2Storyboard as a live reference
seedance_2_0 · 9.0 cr/s at 1080p
Pass the storyboard as an image_reference alongside the start frame, so the model can see every pose in the take while generating it.
For The only route where the model actually sees the intended movement. This is the approach that has worked before on complex character motion.
Against Four and a half times the price. A 10-second take is 90 credits against Kling's 20 — the full film would run past 400 credits on video alone.
Do not pick between them on argument. Setup D is the hardest take in the film — four distinct poses, a locked camera and a body that rotates — so generate it both ways and look. Kling at 10s std is 20 credits; Seedance at 10s is 90. One hundred and ten credits settles whether a live storyboard reference is worth 4.5x, and the answer transfers to every other setup. If Kling holds, the whole film runs on route 1 for 66 credits of video.