15s0 views

Brand spokesperson turns a host's AI skepticism back on him under amber studio neon video, made for SaaS, optimized for YouTube in 16:9 路 Voiced in English

An AI video platform that turns a written description into a finished clip. The description covers everything a shoot would: the room and its lighting, the people and what they are wearing, the camera angles, the cutting rhythm, and the exact words spoken out loud. What comes back is a graded, mixed video rather than a storyboard or a stock assembly.

The buyer is a marketer or founder who has already seen generated video and dismissed it. They watched something with dead eyes and melting hands eighteen months ago, filed the whole category under not yet, and stopped checking. They are not arguing about price. They are convinced they can spot it, and that conviction is the only thing standing between them and a trial.

A podcast two-hander fits because the objection has to be voiced by someone who is not selling. A host says the exact sentence the viewer is thinking, out loud and slightly rudely, and the guest answers it. When the clip then reveals that the studio, the microphones and both faces were generated, the proof is the artifact itself, not a claim about it.

The prompt

prompt.txt299 words
Ultra-photorealistic 4K podcast clip shot on a cinema camera with 35mm and 50mm primes, shallow depth of field, natural skin texture with visible pores and micro-expressions, filmed in a warm modern studio: dark acoustic foam panels, amber neon glow on the back wall, walnut table, two broadcast microphones on boom arms in the foreground of every frame, subtle lens flare from a soft key light, faint haze in the air. Two people sit face to face across the table: the HOST (late 30s, male, beard, headphones around his neck, black tee, leaning back skeptical, coffee mug in hand) and the GUEST (early 30s, female, blazer over t-shirt, calm, confident, laptop closed beside her). Shoot coverage from four angles: wide two-shot over the mic, medium single on host past the mic windscreen, medium single on guest, and tight close-up on eyes and hands. Edit in fast reel-style cuts every 1.2 to 2.5 seconds with zero silence, constant overlapping speech, natural interruptions and laughs, and J-cuts and L-cuts where the next speaker's audio starts a beat before we see them and the previous speaker's voice lingers over the incoming shot. Dialogue: HOST - "Come on, AI video still looks fake, plastic eyes, weird hands, I can spot it in two seconds"; GUEST (smiling) - "Can you? Because you've been looking at it for the last thirty seconds"; HOST (leans in, laughs, glances around the room) - "Wait... what?"; GUEST - "This whole clip. The studio, the mics, me, you. Rendered." Then a slow subtle push-in on the host's stunned face as the room lighting flickers imperceptibly. Audio design: crisp broadcast-quality voices with mic proximity effect, subtle room tone, low lo-fi bass bed, no dead air. Handheld-stabilized micro-movement on every shot, film grain, teal-and-amber color grade, bold captions synced word by word.

Why this works

The clip spends its first seven seconds arguing against itself. The host says plastic eyes, weird hands, I can spot it in two seconds, which is the objection almost verbatim as the viewer would phrase it. Nothing is being sold yet, so there is nothing to resist. By the time the guest answers, the audience has already agreed with the skeptic, and agreement is what makes the reversal land instead of feeling like a trick.

The reveal is structural, not verbal. This whole clip. The studio, the mics, me, you. Rendered. works only because everything before it was ordinary: a bearded host with headphones round his neck, a coffee mug lifted mid-sentence, a guest half smiling while she waits her turn. The close-up on the mug and hands is the most important shot in the piece, because it is the one nobody would bother to fake. The wide two-shot arriving on the last line then shows the whole room at once and invites the rewatch, which is where the actual persuasion happens.

Cutting every one to two seconds with overlapping speech hides the seams a longer take would expose, and it matches the rhythm of clipped podcast content people already scroll past. Word-by-word captions carry the exchange for a silent autoplay, and the amber neon and teal grade give the room a recognisable identity across four camera angles. At fifteen seconds it is short enough to finish before the doubt returns.

Ship one like this

Start with this prompt, swap your product in. Two minutes from idea to download.

Last updated August 2026

More AI Podcast Videos examples