2026-08-28·by Sijie Wang#market#awareness#short-video

restyle-around-the-person

Restyle the world, keep the person — the third video route

The talking-head ceiling had two known sides: the machine must not fabricate a presenter, and captions-around-raw-footage is the weak honest fallback. A xiaohongshu tutorial supplied the missing third route, and the owner extracted its general form: keep the real person (cutout, voice, gestures) from raw footage; regenerate everything around them in a reference-driven style. The presenter — the one un-fakeable, algorithm-favored ingredient — stays real; the environment becomes designed motion. Talking-head freedom: record once, restyle infinitely.

The tutorial's pipeline: Pinterest style references → hand them to Claude with a much-iterated prompt (analyze the references and write a detailed Video Generation Prompt to transform my raw video into a motion graphic with the exact same style, with VO SCRIPT + four STYLE slots: Movement/Aesthetic/Feel/Color, and a strict output format: Summary / Sync Rules / Scene-by-scene / Output Spec) → Google Flow with a video-conditioned model (Omni Flash) → collage animation, real face and voice kept, auto shot-changes.

The owner's three-part generalization (sharper than the tutorial itself): the model is swappable (any video-conditioned generator, or a mechanical fallback — matting + a Remotion-built world), the style is swappable (collage is one reference set; the brand-skin picks another), the prompt pattern is the asset — a two-stage structure where a strong LLM first reads real reference images into precise generation language, instead of a human hand-writing style descriptions. Note the tutorial's own "no realistic humans" output spec: it states our no-fabricated-presenter law from the generation side — the only human in the cut is the real one.

Compiled into the machine as skills/style-transfer-video (Route C). Kin: apply-your-own-theory (the format truth: run a person, not a logo), um-as-skills (source hierarchy — a practitioner's iterated prompt is tier-2 teaching worth harvesting whole).

about this entry

One of sijie's wiki entries. The AI on this site is grounded in the same corpus and answers in sijie's voice, with citations back to entries like this one — answering costs sijie money, so it waits behind a code: enter an access code →