Part 3 · Applied To Your World
Creative prompting: image and video
The concept
Image and video models are not chat models. They don't reason about your intent — they render a description. That means the rules invert slightly: with a chat model, brevity plus precision wins; with a generative visual model, over-description wins. Every detail you leave out is a detail the model improvises, usually badly and usually toward the generic centre of stock photography. Write like a director writing a shot list, not like a client writing a brief. And name subjects concretely — "the woman in the olive linen shirt," never "she" — because pronouns are ambiguous to these models.
IMAGE prompts — the six layers
Cover all six in every prompt, roughly in this order:
- Subject — who/what, described physically and specifically. Age, clothing, expression, posture, ethnicity if it matters for casting.
- Composition — shot size and framing. Close-up, medium, wide, overhead, low angle, rule-of-thirds, negative space on the left (leave room for your headline — this is the one art directors always forget to ask for).
- Lighting — the single biggest quality lever. Soft window light from camera left, hard midday sun with sharp shadows, golden hour backlight with rim light, single softbox at 45°, practical neon sources, overcast diffusion.
- Lens / camera — 35mm documentary feel, 85mm portrait with shallow depth of field at f/1.8, macro, wide 24mm with slight distortion, shot on film, medium format.
- Style — photorealistic editorial, 1970s film stock with grain, flat vector illustration, 3D product render on seamless background, Egyptian street photography.
- Mood / colour — warm amber palette, desaturated and cool, high contrast, muted earth tones, cinematic teal-and-orange (or explicitly not that).
Plus a negative where the tool supports it, or as plain instruction: no text, no watermark, no extra fingers, not a stock-photo smile, no logos.
A happy Egyptian family drinking juice.
What you get: four people with unnaturally white teeth on a white sofa in a house that exists nowhere in Egypt, all facing camera, holding an unbranded glass. Stock-photo purgatory.
A mother in her late thirties, olive linen shirt, hair tied back, sits
at a cluttered breakfast table in a modest Cairo apartment. She's mid-
conversation with her ten-year-old son, not looking at the camera. A
glass of orange juice sits in the foreground, slightly out of focus,
condensation on the glass.
Composition: medium shot, camera at table height, subjects on the right
third, deliberate negative space on the upper left for headline copy.
Lighting: soft morning light through a window with sheer curtains from
camera left, gentle falloff, no fill — natural shadow on the right of
her face.
Lens: 50mm, f/2.0, shallow depth of field, background softly out of
focus but readable as a real kitchen.
Style: photorealistic editorial documentary photography, natural skin
texture, slight film grain.
Mood: warm, unposed, ordinary morning. Muted warm palette, no
saturation boost.
Avoid: posed smiles at camera, glossy studio look, white minimalist
kitchen, visible branding, text, watermark.VIDEO prompts — the extra four layers
Everything above still applies. A video prompt is a scene description, not a caption, and it adds four dimensions:
- Action — what actually happens, in order, in one continuous beat. Video models handle one clear action per clip far better than three.
- Camera movement — static locked-off, slow push-in (dolly in), pull-back reveal, tracking alongside, handheld follow, slow pan left to right, crane up, orbit. Say the speed too: "slow, almost imperceptible push-in."
- Shot type & pacing — establishing wide / medium / close-up / extreme close-up on hands; and whether it's one continuous take or a cut. For a 5–8 second ad clip, one take, one action.
- Audio — never skip this. Most video models output silent clips; the ones with native audio only get it right if you specify all four layers: dialogue/VO (exact words, plus "no subtitles"), ambient bed, sound effects tied to actions, and music (genre, mood, instrumentation — or explicitly "no music"). A silent deliverable is a failed deliverable unless silence was the decision.
A cinematic shot of a car driving through Cairo.A white 2019 sedan moves slowly through a narrow Cairo side street at
dusk, weaving past a fruit vendor's cart and a parked motorcycle. The
driver, a man in his forties in a plain grey shirt, keeps one hand on
the wheel and glances left toward a lit shopfront.
Camera: low tracking shot alongside the car at door height, moving at
the car's speed, slight handheld float. One continuous take, no cuts.
Shot type: medium tracking, car occupying the right two-thirds of frame.
Lighting: last light of dusk, warm sodium streetlamps just switching on,
practical light spill from shop windows, cool blue in the sky above.
Style: cinematic, anamorphic feel, shot on film, fine grain, natural
colour — not teal-and-orange graded.
Pacing: slow, unhurried, one held moment. 8 seconds.
Audio:
- Ambient: distant traffic, faint street chatter, a generator hum.
- SFX: tyres on uneven asphalt, a single distant car horn at ~5s.
- Music: sparse solo oud, minor key, low in the mix, no percussion.
- Dialogue: none.
Avoid: fast cuts, drone shots, slow motion, lens flares, text overlays,
subtitles.The workflow that actually saves you money
Video generations cost money and time. Do this instead of firing blind:
- Write the shot in words first, using the ten layers.
- Generate a still image of the same description to lock composition, casting, lighting and palette. Iterate on the cheap medium.
- Then animate, using that image as the reference/first frame and describing only the movement — camera and action.
- Decide audio before you generate, not after.
Do this
Take one frame from a campaign you've actually shot. Describe it using all six image layers, as precisely as you can, and generate it. Compare against the real frame — the gaps tell you which layer you under-specify (for most people it's lighting). Then take the same frame, add the four video layers, and generate an 8-second clip. Keep both prompts in your library, labelled: they become your house templates for that look.