You just talk to the assistant in normal language — no buttons, no code, no settings. This guide shows you exactly what to type so your character (same face, same outfit) and your objects (same product, same prop) stay consistent from scene to scene.
AI video tools tend to "forget" what your character looked like the moment the next shot starts — you ask for the same person and get a stranger. The fix is almost entirely about how you phrase your request. This guide gives you the exact words to type, from the quickest one-message trick up to building a reusable model of your character.
How to read this: pick your level from the cheat-sheet, then follow When → What you type → Why it stays consistent → What you'll get. The green boxes are things you copy and paste straight into the chat.
Consistency isn't one trick — how much effort you put in depends on how long your video is and how often the character comes back.
| What you're making | Effort | Best approach |
|---|---|---|
| A short clip (a few seconds, 2–5 quick shots) | Simple | One message with a strong "anchor" description — ask for it as a single continuous clip |
| A real multi-scene short (6–15 shots) | Medium | Make one "hero" image first, then tell the assistant to reuse that exact character/object in every scene |
| A character who returns across many videos | Advanced | Ask the assistant to build a reusable model of them from reference photos |
| The same product/prop in lots of settings | Medium | Give the product photo once, ask to keep it identical in each scene |
The tool has no memory of a look between requests unless you give it one. Simply re-describing "Mika in a red jacket" five separate times gives you five different people. So you always do one of three things: (1) ask for a single continuous clip, (2) reuse the same reference image in every scene, or (3) have a reusable model trained. Everything below is one of those three.
For a short piece, the easiest win is to describe your character very specifically once and ask for the whole thing as a single short clip. Made in one pass, the look holds by itself — nothing to set up.
A short idea you can tell in a few quick shots (roughly up to 5 shots / ~15 seconds).
Make me one short continuous video, about 12 seconds, of this character. Keep her looking exactly the same in every shot: Mika, a 28-year-old bike courier — short black hair, light-brown skin — wearing the same red hooded jacket, black cargo pants, and white sneakers the whole time. Shot 1: at dawn she clips a matte-black coffee tumbler into her bike cage. Shot 2: she weaves through morning traffic. Shot 3: she coasts downhill, jacket flapping, the same tumbler still in the cage. Only the camera and setting change between shots — her face and outfit stay identical.
Asking for one continuous clip means it's created in a single pass, so the character can't drift between shots. The repeated outfit words and the "only camera and setting change" line leave almost no room to reinvent her.
A short multi-shot clip where Mika and her tumbler stay on-model throughout — the strongest consistency for a short piece with the least effort. (Behind the scenes there's a length limit around 15 seconds / 5 shots; for anything longer, use Level 2.)
You want a handful of separate scenes but don't want to manage images yourself yet.
Create a 5-scene story about Mika the bike courier delivering a matte-black coffee tumbler across the city. Lock her look and keep it identical in every scene — same red hooded jacket, black cargo pants, and white sneakers. Use the first scene as the reference for the rest, and flag any scene where she stops looking like herself.
Saying "use the first scene as the reference for the rest" tells the assistant to anchor every following scene to that first image instead of starting fresh each time. Asking it to "flag any scene where she stops looking like herself" turns on a quiet consistency check.
A set of scenes anchored to one look, with a heads-up if any shot drifts.
For a proper 6–15 shot short, the reliable pattern is: make one great "hero" image first (your character in the exact outfit), then in each later message tell the assistant to reuse that exact image. This is the single biggest consistency upgrade most people are missing.
Make one clean full-body reference image of my character, plain background: Mika, a 28-year-old bike courier — short black hair, light-brown skin, wearing a red hooded jacket, black cargo pants, and white sneakers. Neutral pose, sharp focus, even lighting. This is my reference image for the whole project — we'll reuse it in every following scene.
One approved "source of truth" image is worth more than any description. Everything after this points back to it.
A canonical picture of exactly what Mika looks like.
Use the character from the reference image you just made — same face, same red hooded jacket, black cargo pants, and white sneakers. Now put her crossing a steel bridge at golden hour, seen from a low angle. Don't change her outfit or face; only the setting and camera change.
Then for the next scene: "Same character from that reference, same outfit — now she's riding through light rain with her hood up." And so on.
Phrases like "use the character from the reference image" or "the same Mika from image 1" tell the assistant to build each scene from that picture rather than imagining a new person. That's the whole game.
Every scene built off the same face and outfit.
Animate that bridge scene into a short clip — Mika pedals forward, camera tracking alongside her. Keep her exactly as she looks in the image — same face, same red jacket and black cargo pants.
Animating from your approved still carries the locked look into the motion, instead of generating a fresh character for the video.
Two shots should feel continuous (a push-in, a pan, a move that carries across the cut).
Make the next shot start from the last frame of that clip and continue the motion — keep pushing in on Mika's face as she brakes. Same look, no jump between shots.
"Start from the last frame" hands the end of one clip to the beginning of the next, so the cut feels invisible and the look can't shift.
Same idea, applied to a product/prop or a specific garment. Pick the phrasing that matches how exact you need it.
Here's a photo of my product. Place this exact coffee tumbler into each of these scenes, keeping its shape, color, matte-black finish, brushed-steel lid, and logo placement identical — only the background changes: 1) on a wooden cafe counter in morning light 2) in a bike cage on a rainy street 3) beside a laptop on a desk at night 4) held in a gloved hand on a snowy trail
Spelling out "keep its shape, color, finish, and logo identical" tells the assistant to transplant the real object into new backgrounds rather than draw a new tumbler each time. Object identity is the thing viewers spot instantly, so be explicit.
The literal same tumbler in every scene.
If you have a real photo of the jacket and it has to match precisely:
Here's a photo of Mika and a separate photo of the red hooded jacket. Put this exact jacket on her, matching its real color, texture, and cut.
If you don't have a garment photo and just re-describe the outfit each time, expect small color/texture wobble — that's normal. For an exact match, give the real photo like above.
Pick one fixed phrase for the wardrobe — "the same red hooded jacket, black cargo pants, and white sneakers" — and paste it into every scene request word for word. The more identical the wording, the less the outfit drifts. Only change pose, camera, and setting between shots.
When Mika comes back across many scenes or many separate videos, reusing one image starts to strain — she'll drift under big changes in pose, angle, or lighting. The lasting fix is to have the assistant build a reusable model of her from several reference photos. After that, you just say her name and she comes out looking like herself.
I'm attaching several reference photos of Mika. Train a reusable model of her from these so she looks identical in future videos — then use it for everything in this project. Call it "Mika".
A trained model learns Mika's actual identity, so it holds up even when the pose, camera angle, and lighting change a lot — far beyond what a single reference image can do.
A reusable "Mika" you can call on in any later request in that project.
Using the Mika model, make a scene of her sprinting up a fire escape at night — low angle, red hooded jacket, dramatic lighting.
Give it a variety, not 20 copies of the same headshot:
Avoid: watermarks, heavy text, and busy backgrounds. Give at least 8 photos; 12–16 is the sweet spot. The same trick works for a signature outfit — train on photos of the outfit so it comes back exactly.
Save a brand style for this project: main colors red, black, and white, and always include "red hooded jacket, black cargo pants, matte-black tumbler". Apply this style to everything I make in this project from now on.
Once a brand style is saved and active, its colors and wardrobe keywords get folded into every request automatically — you don't have to retype them.
Keep these four and fill in the amber blanks. Reuse the same wording every time — that repetition is what holds your look together.
CHARACTER (keep identical in every shot): [Name], a [age]-year-old [role] — [hair], [skin/eyes/build], wearing [signature outfit — be specific: colors + items]. Only these change per shot: camera angle, pose, and setting. Face and outfit stay exactly the same.
Here's a photo of [the object]. Place this exact item into each scene, keeping its [shape, color, finish, materials, logo placement] identical. Only the background changes.
Use the [character / object] from [the reference image / image 1 / the clip you just made] — same look, same outfit. Now [new setting + camera + action]. Don't change the face or outfit.
Make the next shot start from the last frame of that clip and continue the motion — [what the camera/subject keeps doing]. Same look, no jump between shots.