
From Brief to Keyframes: My AI Visual Development Process
My AI visual development process begins with story and production needs—not prompts. Here is how I move from a creative brief to coherent cinematic keyframes.
A woman wakes up in a familiar bedroom. At first, nothing seems wrong. Then she wakes up again—and one small detail has changed. When I developed *A Dream of a Dream*, I needed those moments to feel connected while gradually making Mia, and the audience, question what was real. A beautiful image on its own couldn’t do that. I had to decide what each shot revealed, what it concealed, and how it would lead into the next one. That’s how I approach visual development. Before I generate a keyframe, I work out what the scene needs the viewer to see and feel. Here’s how I turn that decision into a visual sequence.
From Brief to Keyframes: My AI Visual Development Process
A successful AI-generated keyframe does more than look cinematic. It solves a specific storytelling problem.
When I receive a brief—or begin developing one of my own films—I do not start by writing prompts. I start by identifying what the audience needs to understand and feel. The visual system comes after that: character, environment, light, camera position, lens behavior, palette, texture, and movement all have to support the same intention.
This is how I turn an initial idea into keyframes that can guide a film, commercial, trailer, or generative video sequence. To make the decisions tangible, I will use the visual development of A Dream of a Dream, my film about false awakenings, as an example.

1. I translate the brief into visual problems
Creative briefs often arrive as a mixture of narrative information, brand requirements, references, technical constraints, and abstract words: cinematic, premium, energetic, intimate, surreal, realistic.
Those words are useful, but they are not yet visual decisions. “Cinematic” could describe a quiet 105mm close-up, a handheld wide shot, a controlled studio composition, or an anamorphic landscape. “Premium” can become polished and lifeless if it is not connected to a specific world.
I break the brief into questions:
Who or what is the emotional center of the scene?
What changes during the moment?
What information must be immediately legible?
What should remain ambiguous?
Is the image selling a product, establishing a world, revealing character, or creating tension?
What parts will be filmed, generated, composited, or animated later?
What continuity constraints already exist?
This step keeps the work from becoming a collection of attractive but disconnected images.

2. I define the visual system
Before generating individual shots, I establish a shared language for the project. Depending on the brief, that system may include:
color palette and contrast range;
lighting motivation and time of day;
camera format and lens character;
composition rules;
texture and level of realism;
environment design;
character identity and wardrobe;
motion language;
elements to avoid.
For a personal film, the system may begin with an emotional memory. For a commercial project, it also needs to respect brand identity, product visibility, and the production pipeline. In both cases, the purpose is the same: every frame should feel as if it belongs to one world.
I prefer specific visual references over large, unfocused moodboards. A reference should answer a question—about light falloff, camera distance, production design, movement, or texture—not simply communicate taste.

3. I map the sequence before polishing images
When the work involves multiple shots, I create a simple sequence map. This can be a shot list, beat sheet, or rough storyboard depending on the project.
At this stage I define the function of each frame:
establish the space;
introduce the subject;
direct attention;
reveal a change;
intensify emotion;
create a transition or resolution.
This prevents a common problem in AI filmmaking: producing many variations of the same visually impressive shot while leaving essential narrative beats uncovered.
I deliberately vary shot scale and point of view. A wide frame establishes geography. A medium shot connects body language to environment. A close-up gives emotional access. An insert can make a tactile detail—fabric, skin, water, an object—carry narrative weight.
The order matters. A keyframe is not only an image; it is a proposed edit. In A Dream of a Dream, an eye close-up puts us inside Mia’s perception. A bedroom frame establishes a familiar space. A repeated awakening asks us to notice what has changed; a distorted frame makes that uncertainty physical. Each image has a different narrative job. Four beautiful portraits of Mia would not tell the same story.
4. I design the shot, not just the prompt
Prompting is only one part of the process. The creative decision is the shot itself.
For every keyframe, I consider:
subject position and eyeline;
camera height and angle;
shot size;
lens and depth of field;
foreground, middle ground, and background;
direction and quality of light;
pose and weight distribution;
screen direction;
negative space;
likely camera or subject movement;
how the frame will connect to adjacent shots.
The prompt then becomes a production instruction for that design. I describe relationships and physical behavior, not only aesthetic adjectives. For a false-awakening scene, useful information includes exactly where Mia is in relation to the bed, what she can see, where the light comes from, and which visual detail changes when the scene repeats.
This is where knowledge of photography and filmmaking becomes essential. Generative tools respond more reliably when the creative intention has been translated into observable visual conditions.

A real example: directing repetition
In A Dream of a Dream, the dramatic problem is that Mia believes she has woken up, but the room and her senses can no longer be trusted. The first awakening needs to feel ordinary. If the bedroom already looks frightening, the later breaks have nowhere to go.
I begin with stable geography: where the bed sits, what Mia faces when she opens her eyes, the direction of the light, and the path toward the door. When an awakening repeats, I can keep part of that composition familiar and alter one thing at a time: her eyeline, a reflection, a light source, or the distance to the exit. The sequence can become more unstable without losing the viewer completely.
A macro image of an eye, a waking shot, an image that repeats with a subtle change, and a moment of physical panic each answer a different visual question. The frame has to carry its place in that progression. A dramatic image that reveals the nightmare too early may be striking on its own but wrong for the scene.
5. I generate for exploration, then narrow the field
Early generation is exploratory. I may test different ways to express the same beat: a wider lens versus a compressed frame, frontal light versus backlight, static composition versus handheld immediacy.
But exploration needs limits. Once a direction begins to work, I stop introducing unrelated aesthetics and concentrate on the variables that still need resolution.
My selection criteria usually include:
narrative clarity;
emotional accuracy;
character and wardrobe continuity;
spatial logic;
believable anatomy and physics;
lighting consistency;
camera coherence;
suitability for animation or compositing;
alignment with the brief.
An image can be beautiful and still be wrong. If it contradicts the emotional beat, breaks continuity, or creates a production problem downstream, it is not the right keyframe.

6. I evaluate the production path
Keyframes are often treated as final artwork, but in production they can serve several different purposes. A frame might guide live action, become a direct input for image-to-video generation, establish an environment for compositing, or communicate an idea to a director and client.
I adapt the frame to its next use.
For animation, I look for a readable pose, clean silhouette, coherent depth, and enough visual information for movement. For hybrid production, I consider where live-action elements will sit, how the light will integrate, and whether perspective supports the composite. For a pitch, I prioritize immediate communication of tone and story.
This avoids treating every image as an isolated endpoint. Visual development is most valuable when it reduces uncertainty for the next creative decision.
7. I build coherence at sequence level
Once the strongest frames are selected, I review them together rather than one by one.
I check whether:
the sequence has visual rhythm;
shot scales vary with purpose;
screen direction remains understandable;
the lighting progression makes sense;
characters remain recognizable;
the environment maintains geography;
color and texture belong to the same world;
each frame contributes new information.
Sometimes the best individual image has to be removed because it disrupts the sequence. Sometimes a quieter frame becomes essential because it creates contrast before a more intense moment.
This stage is closer to editing than image generation. It is where a group of keyframes begins to function as a cinematic structure.
A keyframe is a decision
Generative AI can produce an enormous number of images, but volume does not equal visual development. The value lies in deciding what the project should look like, why it should look that way, and how that decision can survive across a sequence and production pipeline.
My role is to build that bridge between an idea and an executable visual world. I use generative tools to explore and produce, but the process is guided by story, photography, production design, continuity, and editing.
The result I aim for is not a folder of images. It is a clear visual direction that a director, studio, or production team can evaluate, develop, and use.

Work with me
Explore my selected film and commercial work or learn more about my practice. I collaborate internationally on cinematic visual development, keyframes, AI-generated environments, character work, and generative film sequences. Contact me to discuss a project.
