The Diary Of A CEO —
Vitamin D Expert: The Supplement World Is Giving The WRONG Advice! | Dr Stasha Gominak
- Platform
- youtube
- Length
- 120:19
- Views
- 5.4M
- Median cut
- 1.96s
- Face on screen
- 80%
- Dark frames
- 72%
channel-intro
Cold-Open Kinetic Caption Interview
medium effortSteal this
Light and frame so that one third of every interview shot is intentionally black, then treat that void as a permanent caption plate — big condensed caps, one word per line coloured red — and cut between two camera sizes every two seconds so the type keeps re-entering.
A long-form two-person interview is front-loaded with a 60-90 second cold open cut entirely from the best soundbites, each one carrying a large kinetic caption on the empty side of frame. Roughly every two seconds the edit flips between a medium and a tighter angle of the guest, punctuated by handheld illustrated-prop inserts and a torn-newspaper name-card title card that establishes credentials. The look is dark, low-key and warm-neutral so white and red type sits hot against the background.
Frame one is already mid-sentence: guest in a medium two-shot framing at a low pale table, and a stacked three-line title ("THE / VITAMIN / WORLD") slams in bottom-left over her body while she talks. No logo, no host intro, no establishing wide — the caption does the hook work and the eye is pulled to text before the viewer has decided whether to stay.
37 shot changes in 90 seconds, median 1.96s between cuts, so it is a near-metronomic two-second flip. Cuts are made on sentence beats rather than on movement, and many are angle changes within one continuous take of speech (jump-cut between an A and B camera), which lets the words run uninterrupted while the picture keeps refreshing. A face is present 80% of the time; the illustrated inserts and the title card are the only breathing spaces.
Heavy condensed sans, all-caps, mixed weights and sizes within one phrase so the operative word is 2-3x the surrounding text, plus occasional italic for emphasis. Predominantly white with a thick dark drop shadow for legibility, one accent colour per line (blood red on the guest, pink on the host) reserved for the single stress word. Numbers appear as huge standalone figures. Type is stacked and left- or right-aligned into whatever side of the frame is dark, never over the face. Title card breaks style deliberately: ransom-note torn newsprint, hand-cut portrait, faux-documentary collage. Grade is low-key and desaturated with warm practical highlights; 72% of frames read dark, blacks crushed at the frame edges.
- Medium presenter, guest seated frame-left, table edge in lower third, right third of frame falling into deep shadowThe workhorse. The dark negative space on camera right is deliberately reserved as a caption plate.
- Close presenter, shoulders-up 3/4 profile, shallow depth, background lamp bokehCuts in on emphasis lines; the tighter size makes a caption change feel like a new beat even on the same sentence.
- Reverse single of the host, medium close, looking off-frame right, cooler and flatter lightReaction / listening cutaway that resets the rhythm and proves it is a conversation, not a monologue.
- Handheld top-down macro insert of hand-drawn illustration boards and a hand on a bedsheet, motion blur, slight tiltTactile B-roll break every 15-25s; kills caption fatigue and gives a visual for an abstract idea.
- Title card: cut-out portrait of the guest on a blurred newspaper-collage background with torn-paper name and credential stripsDelayed credentialing — identity lands only after the hook has already worked.
- Set the room dark: light only the subject's face and one warm background practical, leaving one third of frame in near-black so type has a home.
- Shoot the interview on two locked cameras at different focal lengths, same eyeline side, running continuously; add a reverse single on the host for reactions.
- Before editing, pull the 8-12 strongest sentences from the full conversation and lay them end to end as the cold open — no intro, no pleasantries.
- Cut on the words: alternate medium and close roughly every two seconds so each new clause gets a new frame size.
- Caption every line by hand, breaking phrases across 2-4 lines, and blow up the single most important word in each phrase; colour only that word.
- Always park the type in the dark side of frame, and switch sides when you switch cameras so it feels designed rather than automatic.
- Insert a handheld top-down shot of a physical hand-drawn prop every 15-25 seconds as a texture break; keep it slightly blurred and moving.
- At around 25-30 seconds, drop a torn-newsprint title card with a cut-out portrait plus a second card listing credentials, then return straight to the interview.
- Grade dark: crush blacks, desaturate skin slightly, keep the practical lamp warm so the white type reads as the brightest thing on screen.
- Two cameras: one medium/wide on the guest, one on a longer lens for the close single (plus a third or repositioned single for the host reverse)
- Fast primes or a 24-70/70-200 pair for shallow background separation
- Key light with soft modifier from the guest's front-side, plus a warm practical lamp behind for bokeh; flag the opposite side to keep it black for captions
- Lavs on both people plus a boom as backup
- Table and low-contrast chair set in a dark room with textured wall/window panel behind
- Phone or small gimbal-free handheld camera for the top-down illustrated-prop inserts
- Hand-drawn prop cards/illustrations made in advance
- Edit: Premiere/Resolve plus After Effects (or a caption template pack) for kinetic type and the torn-paper title card
- All audio: whether there is a music bed, sub-drops or whooshes under the caption hits and the title card
- Whether cuts land on musical beats or purely on speech
- Whether the inserts are original footage or licensed/animated assets
- Exact number of cameras (the host reverse may be a repositioned single from a separate pass)
- Whether captions are auto-generated then restyled or fully hand-animated
Written by anthropic/claude-opus-5 from the frames and measurements above. Technique only, and it cannot hear the video.