Founder on-screen walks through his office admitting he never filmed the video video, made for SaaS, optimized for TikTok in 9:16 路 Voiced in English
A software platform that turns one photo and a written prompt into finished talking videos. Users upload a single image of a face, describe the scene in plain text, and get back lip-synced clips with the same person across every shot, ready to be cut into ads, pitches, and social content.
The buyer is a founder, ecommerce operator, or agency producer who needs a steady stream of video creative and has no time, budget, or appetite for shoots. They have watched competitors ship native-looking ads every week, and they know a real face on camera converts better than stock footage, but the camera setup is the thing that never gets scheduled.
A founder ad works here because the message and the medium are the same proof. The person on screen claims the video was never filmed, and the footage itself is the demonstration. For a product whose entire pitch is skipping production, showing the output while describing the process closes the argument in under thirty seconds.
The prompt
Hyper-photorealistic vertical 9:16 video in three scenes (15s, 10s, 10s), cinematic handheld documentary look, no music, no on-screen text. SUBJECT, identical in every scene: a Caucasian man in his mid-20s, medium build, fair slightly flushed skin with realistic texture and small imperfections, light blue eyes, light brown hair pulled back into a small low bun with a few loose strands, short well-groomed light brown beard and moustache, plain black crew-neck cotton t-shirt, no jewelry, no glasses. Natural confident relaxed energy, direct eye contact with the lens. SETTING for all scenes: a real modern startup office in late afternoon, large window with soft natural daylight from camera left, white desks, monitors slightly out of focus in the background, a plant, cables, coffee cup, a few papers, lived-in and authentic, not staged. Scene 1 (15s): he stands mid-frame and talks straight to the camera while walking slowly forward two or three steps, then stops. Small natural hand gestures, a light shrug on the last line, a brief half-smile, natural blinking and micro-expressions. Dialogue, native fluent English, casual confident tone, lip-synced: "I'm the founder of this company. I never filmed this video. No camera, no crew, no budget. Just one photo and some text." Camera: handheld with slight operator breathing, medium shot from chest up, subtle slow push-in, shallow depth of field, 35mm look, realistic motion blur, one small natural focus adjustment. Scene 2 (10s): he sits on the edge of a desk, one foot on the floor, counts on his fingers while listing, then opens both hands in a wide disbelief gesture on the last line. Dialogue: "A normal video ad? A studio, a videographer, an actor, two days of editing. Three thousand dollars, easy." Camera: handheld medium shot from chest up, slight sway, shallow depth of field, ends with his hands lowering back down. Scene 3 (10s): he is mid-walk across the office when the shot starts, weaving between two desks, talking to the camera over his shoulder, then turning to face it and dropping backwards to sit on a desk edge. He grabs a phone off the desk, lifts it briefly, tosses it back down. Energetic, casual, never static. Dialogue: "Here? One photo of your face. One paragraph describing the scene. That's the whole process." Camera: handheld operator walking alongside him, strong natural bounce, one quick reframe as he turns, a brief autofocus hunt, slight roll of the horizon, imperfect framing that briefly crops the top of his head. Audio in all scenes: clear close voice, quiet office room tone, faint keyboard and distant street noise, no music. Avoid: plastic skin, beauty filter, CGI look, warped hands, extra fingers, readable phone screens, changing face, hairstyle or clothing, on-screen text, logos, other people in focus, robotic text-to-speech delivery.Why this works
The opening line is a claim the viewer can audit in real time: "I never filmed this video." Every other ad asks to be believed; this one asks to be inspected, and the inspection is the watch time. The harder you look for the seam, the longer you stay, and by the time you give up looking, the pitch is over.
The authenticity load is carried by deliberate imperfection. Handheld bounce, operator breathing, an office with cables and coffee cups instead of a set, and a walking shot where the framing briefly clips the top of his head. The same face holds across three separately generated scenes, which is the technical claim the product actually needs to prove. The gestures are scripted against the words: he counts the cost items on his fingers as a receipt overlay stacks up on screen, and the three thousand dollar total slams in with the spoken number. Word-synced captions keep the argument legible with sound off.
The structure is a cost contrast, and the edit makes each side concrete. The old way gets a line-item receipt; the new way gets the actual input, a photo thumbnail and then the full generation prompt typing itself into the product's prompt box as he says "one paragraph describing the scene." Ending on the literal prompt turns the ad into a tutorial: the viewer leaves knowing exactly what they would type. At 28 seconds vertical, it fits the platforms where the target buyer is already studying everyone else's ad creative.
Ship one like this
Start with this prompt, swap your product in. Two minutes from idea to download.
Last updated August 2026