Webinsane

AI video

The face is the brief.

AI video models no longer fail by being ugly. They fail by being helpful.

, Founder & Creative Director6 min read

There is one sentence I typed into a prompt box more often this month than any other: keep the face intact, don’t change anything, just animate. I typed it over vertical fashion reels, over a Renaissance portrait asked to open its mouth and let out a stream of flowers, over two craftsmen working on a wooden boat, over a slow dolly-zoom out from a face that had to stay exactly the face it started as. Several hundred generations, most of them in Kling 3.0 and motion-control models, and most of them in service of a single question.

How do you get a model to animate a person without turning them into somebody else?

That is the question that decides whether AI video character consistency is a curiosity or a production tool. This is the method we arrived at, and the reasoning behind it.

The models have stopped failing the way we expected.

A year ago, a bad AI video shot was obviously bad. Hands melted, physics gave up, faces slid. That is mostly behind us. What fails now is subtler, and for brand work more dangerous: the model improves things. The jawline gets a little stronger. The eyes get a little larger. The jacket becomes a slightly better jacket. No single change is wrong on its own, and together they produce a person who is no longer in the brief.

For a campaign built around a real ambassador, a founder or a recurring character, that is the whole job failed. It also means the instinct most people bring to prompting works against you. Describing more, and describing better, is exactly the wrong move. Every adjective spent on how someone looks is an invitation to change how they look.

So the craft is not description. It is constraint.

Start from a picture someone has already approved.

We almost never use text-to-video for work a client will sign off. We make the still first, generated or photographed, and get it approved as a picture, because a still is cheap to iterate and a video is not. Only then do we animate it, using that approved image as the start frame. The model is no longer inventing a person. It is continuing one.

When the camera move matters, we pin the end as well. A dolly-zoom with only a start frame gives the model a long, open road to drift down; with a start and an end frame, it only has to find the journey between two fixed points. Faces survive journeys far better than they survive open endings. The vertical push-outs we produced this month held identity only once both frames were supplied, and I would now treat start-and-end framing as the default rather than the advanced option.

Write the motion, and leave the person alone.

Once the frame is fixed, the prompt has one job: say what happens next. One pumps while the other measures. They pull on each side. She walks along the rocks at dusk, looking out to sea. The image already says what everyone looks like. The words should spend nothing on it.

Two smaller lessons came out of the log. The first is to write in the language you think in. Some of our best direction this month was written in Montenegrin, and the models handled it well; a director describing motion in their own language is simply more precise than one translating in their head. The second is to treat automatic prompt enhancement with suspicion on identity-critical shots. It is useful for motion, but it tends to add style vocabulary, and style vocabulary is exactly what moves a face.

One photograph is not a person.

A single reference leaves the model to guess the far side of a face, the profile, how the hair falls when the head turns. It guesses beautifully, and wrongly. For anything with a real person in it, we now supply a character sheet: four or five images of the same face from the front, at three-quarters and in profile, plus a full-length shot for proportion and clothing.

For scenes with two people, each gets a sheet, and we add a clean plate of the location with nobody in it. That last image turned out to matter more than I expected. It tells the model what belongs to the room and what belongs to the people, and in our boat-workshop scenes it was the difference between two recognisable craftsmen and two plausible strangers.

Borrow choreography rather than describing it.

Words break down on complex movement: a dance, a walk with a turn, two people passing something between them. Motion control approaches the problem from the other side. You hand the model a reference clip for the movement and your character sheet for the identity, and it transfers one onto the other.

We used it to put a specific person into the choreography of black-and-white fashion reels, keeping their own face and hair while taking the camera language of the reference. The lesson was about precision of intent. “In the same style as the reference” will quietly copy the wardrobe too. “Keep my face, hair and clothes; take the movement and the grade” produces something very different. It also helps to match the aspect ratio of the source clip, since reframing during transfer is where bodies start to distort, and to replace one figure at a time in group shots rather than asking for everyone at once.

Length is where drift compounds.

A four-second shot that holds identity is worth more than a twenty-second shot that holds it for fifteen. We now generate in four- to six-second pieces and cut them together, the way live action is shot in setups rather than one unbroken take. The edit is also where the work comes back to people: pacing, sound, type, the decision about which take actually carries the idea.

What this means for brands.

AI video is ready for brand work in a particular shape: short, art-directed shots, built from approved stills, with identity constrained rather than described. That shape covers a great deal of what a brand needs every month: social cutdowns, product loops, campaign variants for each channel. It is how our AI production work delivers video in days rather than weeks, with a person signing off every frame that leaves the studio.

It is not yet the right tool for long dialogue scenes that need continuity across minutes, and we say so in the first meeting. The quickest way to lose trust in this work is to promise it the one thing it still does badly. The volume side of the story is in Generation is cheap. Judgement isn’t.

The models no longer fail by being ugly. They fail by being helpful. Most of the job now is politely declining the help.

Asked often.

How do you keep a character consistent in AI video?

Animate from an approved still rather than from text, supply both start and end frames for camera moves, give the model a multi-angle character sheet, prompt for motion instead of appearance, and keep each shot to four to six seconds.

What is motion control in AI video?

Motion control transfers the movement from a reference video onto your own character. The reference clip supplies the choreography and the character images supply the identity.

Is AI video good enough for brand campaigns?

For short, art-directed shots such as social content, product loops and campaign variants, yes, with human sign-off on every frame. Long scenes with continuous dialogue are still better shot traditionally.

Written by

Tomo Vukasović

Founder & Creative Director

Tomo Vukasović founded Webinsane in Belgrade in 2004 and has led its design work since: product and UX for companies like Pilatus Aircraft and ScholarshipOwl, brand systems, and now AI production, meaning image, video and code made with generative tools at brand quality.

If this is the work you need done, we should talk.

Become a client