20s0 views

Talking head coach concedes the AI tell has moved, in a green panelled studio mid-episode video, made for SaaS, optimized for YouTube in 16:9 路 Voiced in English

An AI video platform that produces finished clips from a written brief. The brief describes the room, the lighting, the wardrobe, the camera coverage, the cutting rhythm and the exact spoken lines, and the output is a graded video with mixed audio rather than a set of assets to assemble.

The buyer here is further along than the outright skeptic. They accept that the technology moved, they have seen a demo that impressed them, and their remaining objection is narrower and more technical: the failure modes they can still name. Motion, faces, hands, continuity between shots. They want to hear those limits acknowledged honestly rather than promised away, because a vendor who claims everything is solved reads as a vendor who has not shipped anything.

The podcast excerpt fits that buyer because it lets the product concede. A guest who says a year ago you were right, and this specific part is still hard, earns the sentence that follows. The format also removes the ad frame entirely: two people arguing at minute forty-one of an episode are not performing for anyone, which is exactly why the argument gets heard.

The prompt

prompt.txt458 words
Photoreal 20-second podcast clip that looks like an excerpt pulled from the 41st minute of a two-hour episode, shot on a full-frame cinema camera with 35mm and 50mm primes at T2.0, set in a calm upscale studio room with deep forest-green wainscoted walls and classic moulding panels, a long solid oak table with visible grain running through the foreground, dark navy velvet armchairs, and warm practical lighting only: a fabric-shaded floor lamp and a small table lamp glowing softly in the background, plus a gentle soft key on each speaker and clean natural fill. No LED strips, no neon, no colored gels, no haze or smoke, no shelves or clutter or decorative objects anywhere, just bare panelled wall, lamp glow, and a single out-of-focus plant, background falling into soft shadow while faces stay evenly and naturally lit with accurate skin tones. Two boom arms with black broadcast microphones sit large in the foreground of every frame, a tablet on a stand and a phone face-down on the table, half-empty water glass. Two men in their thirties sit face to face mid-conversation, already warmed up and casual, natural skin texture with pores and stubble, slight forehead shine, no exaggerated expressions, occasional glances away while thinking, small hand gestures over the table, a shrug, a sniff, one of them nudging his mic without noticing. Four angles only: over-the-mic medium on the host, over-the-mic medium on the guest, wide two-shot from the side of the table, and a loose close-up on hands and face. Cut every 1.5 to 3 seconds with hard cuts, no transitions, zero dead air, natural overlap where they talk slightly over each other, and consistent J-cuts and L-cuts so the incoming voice lands a half-second before the picture changes and the outgoing voice tails over the next shot. Dialogue is flat, real, conversational, no punchline delivery: HOST - "I just think people overstate where it's at. Like, generated video, I watch it and something's off. The eyes, the way the hands move, whatever, I can't always name it but I clock it"; GUEST - "Yeah, a year ago, sure. The tell used to be motion, faces mostly. That's mostly solved now, it's lighting continuity that's still hard"; HOST - "You'd say that though, that's your whole business"; GUEST (small laugh, not big) - "I mean, we did the spot that's running on this episode, so, yeah"; HOST - "Right, I forgot that was you guys"; GUEST - "Nobody mentioned it, which is kind of the point". Audio is raw podcast sound, close-mic proximity, faint room tone, chair creak, no music bed, no sound design. Camera locked on sticks with only micro-drift, very light grain, neutral clean grade, clip ends mid-sentence on the host starting his next thought.

Why this works

Nothing in this clip performs. The host does not set up a punchline, the guest does not deliver one, and the strongest line in the piece, nobody mentioned it, which is kind of the point, is thrown away rather than landed. That flatness is the persuasion. A viewer who has learned to recognise the shape of an ad gets no cue to brace against, so the argument arrives before the defences do.

The room does an unusual amount of work. Green panelled walls, an oak table, two lamps and nothing else: no LED strips, no haze, no shelf of props. Those absences are deliberate, because clutter and coloured light are the ingredients that make generated footage look generated. Holding four fixed angles across the whole clip reinforces it, since real episodes are cut from a small number of locked cameras and every extra angle would imply a shoot that never happened. The result reads as coverage rather than as a series of renders.

The argument itself is structured as a concession. The guest agrees the host was right a year ago, names the failure that is genuinely still hard, and only then mentions that the spot on this episode was theirs. Ending mid-sentence, on the host starting his next thought, is what sells the excerpt: real clips are cut out of something longer, and a clean closing line would have exposed the whole thing as an ad with a beginning and an end.

Ship one like this

Start with this prompt, swap your product in. Two minutes from idea to download.

Last updated August 2026

More AI Podcast Videos examples