Keep characters and speech clear
Help the video model identify the intended speaker, keep spoken words separate from directions, and review voice consistency.
Studio Motion
Describe the person the model can see
A name such as “Robert” does not tell the video model which person in a picture is Robert. Use a short visible description with the name: “Robert, the adult dad with wavy brown hair, a short beard, and a blue shirt.” For Leo, use “Leo, the young son with tousled brown hair and a green T-shirt.”
Add position only when it matches this scene’s picture. A character can move from left to right between pages. Choose artwork where the intended speaker’s face and mouth are visible.
Keep words and directions separate
| Control | Example |
|---|---|
| On-screen speaker | Robert, the adult dad in the blue shirt with a short beard |
| Exact words | It’s nine o’clock. |
| Voice direction | Calm adult male voice, warm tone, lightly firm delivery |
| Scene action | Robert looks toward Leo; Leo listens with his mouth closed. |
Only dialogue belongs in Exact words. Do not put “Robert says”, clothing, acting instructions, or camera directions there unless you want those words spoken. Use consistent identity and voice wording whenever the same character speaks. The app formats speech separately from scene instructions, but generated speech can still get the speaker or delivery wrong.
Plan short turns
One speaker and one short line per clip is easier to control than several rapid speaker changes. Use the pacing hint, leave time for a reaction, and choose ten seconds when a line needs more room. A five-second clip should not be asked to contain a full back-and-forth conversation.
For automatic movies, set the voice approach during planning and check each page’s speaker and exact words before starting the run. For manual generation or replacement, review the same settings in Sound in this scene and the voice editor before confirming.
Choose narration when that fits the story
Video narrator requests an off-screen voice. Add narration separately lets you record, import, or generate narration after the video is ready; this is useful when you want to control the audio independently from lip movement. Check the page’s video sound so you do not accidentally play two voices at once.
Listen before accepting
Replay each completed clip and compare it with the neighbouring pages. Verify who speaks, what is said, and how it sounds. Check Generation details when investigating a mismatch, and replace only the affected scene.
Consistent descriptions and references help; they do not lock a generated voice or guarantee identical characters across clips. Shared cast profiles and automatic picture checks are future improvements, not current beta controls.