Restyle the world, keep the person — the third video route
The talking-head ceiling had two known sides: the machine must not fabricate a presenter, and captions-around-raw-footage is the weak honest fallback. A xiaohongshu tutorial supplied the missing third route, and the owner extracted its general form: keep the real person (cutout, voice, gestures) from raw footage; regenerate everything around them in a reference-driven style. The presenter — the one un-fakeable, algorithm-favored ingredient — stays real; the environment becomes designed motion. Talking-head freedom: record once, restyle infinitely.
The tutorial's pipeline: Pinterest style references → hand them to Claude with a much-iterated prompt (analyze the references and write a detailed Video Generation Prompt to transform my raw video into a motion graphic with the exact same style, with VO SCRIPT + four STYLE slots: Movement/Aesthetic/Feel/Color, and a strict output format: Summary / Sync Rules / Scene-by-scene / Output Spec) → Google Flow with a video-conditioned model (Omni Flash) → collage animation, real face and voice kept, auto shot-changes.
The owner's three-part generalization (sharper than the tutorial itself): the model is swappable (any video-conditioned generator, or a mechanical fallback — matting + a Remotion-built world), the style is swappable (collage is one reference set; the brand-skin picks another), the prompt pattern is the asset — a two-stage structure where a strong LLM first reads real reference images into precise generation language, instead of a human hand-writing style descriptions. Note the tutorial's own "no realistic humans" output spec: it states our no-fabricated-presenter law from the generation side — the only human in the cut is the real one.
Compiled into the machine as skills/style-transfer-video (Route C). Kin: apply-your-own-theory (the format truth: run a person, not a logo), um-as-skills (source hierarchy — a practitioner's iterated prompt is tier-2 teaching worth harvesting whole).