2026-08-28·by Sijie Wang#market#awareness#communication

what-transfers-to-agent-content

What transfers to agent-produced content

utter-manual produces AI-voiceover video scripts and written posts — no live body, no real-time room-read, no nerves. So the {live}-only craft (breathing, gesture, eye contact, posture, props, conversation games, vocal drills) is out of scope for the machine — it belongs to the human residual (cold-start, rapport-and-presence). This note is the filtered core the generator and the video-production-pipeline VO-direction actually consume. It is L3 context (context-derives-from-the-sop): the delivery register a cell's generator is conditioned on, and the thing that pushes a draft off the prior's center (slop-is-a-context-deficit).

Prosody direction (write these as explicit script/TTS direction)

  • End statements on a low pitch — phrase so lines land, not phrasing that invites uptalk; no trailing "…right?"
  • Vary the dials as a highlighter — slow for the important line, short passion bursts (never default), volume up for energy / down for lean-in, pitch low for authority / high for warmth. Volume variety + pitch variety + pace + tonality = engagement.
  • Energize the whole sentence — sustain energy/air to the final word; no vocal fry, no decaying tails ("clarity lives on the edges").
  • Bake emotion in explicitly ("face = remote control") — script the emotional intent of each line; flat = un-felt.
  • Strategic pauses = whitespace — script silence after a punchy line and after the first word ("Hey. [pause]"); pause replaces filler and gives processing time.
  • Melody (off monotone) · keyword emphasis (mark the one load-bearing word) · calibrate energy to "10–20 people in the room," run hotter than feels natural (the camera drains 50–80%) · a chosen "vocal image" (voice casting = a personality decision).

Scripting structure & clarity (VO scripts AND written posts)

  • Declarative short sentences — "Twitter, not a 17-text thread"; kill hedges and qualifier-piles; cut the imposter disclaimer ("here's one idea," not "this might be dumb, but").
  • Drop fillers / non-words; say it once, clearly — but keep one intentional core-message repetition (the anchor).
  • Simplify numbers into pictures ("more than the population of Belgium"); simple language over jargon.
  • Structure with a framework — 3-2-1 · PREP · CCC · CLEAR. The scaffold that lets a script skeleton write itself and reads coherent, not rambling (clarity-and-frameworks).
  • Analogy / metaphor / simile engine — connect unknown→known.
  • Story before content — open with a story (first 30 seconds decide it), relive don't report (present tense + one vivid sensory detail), cut to the peak, then "the reason I'm telling you this is because…" → the lesson-for-them. Honor the 15-minute rule (short/decisional → straight answer, not a parable) and VAKS sensory layering (storytelling-craft).
  • Content mix 33-33-33-1 · Head + Heart aligned · cliché + twist · headline-first · on-screen enumerate / contrast / diagram as the visual analog of gesture (for video-production-pipeline).

Intent & process (upstream of the script)

  • Pre-flight question: "what do I want the viewer to think, feel, and do?" — collapses rambling; the intent field of the context pack.
  • Write to serve the viewer, not to look impressive (audience-conscious, connection-mode) — the transferable core of the nerves fix, and the register that reads as genuine rather than marketing (converges with the reddit-sop answer-shaped rule and the outsider-founder-earns-voice humility).
  • Record-and-review your own generated output in isolated channels (the-practice-engine): a transcript pass for filler & structure, an audio pass for prosody & tonality, a say/show-congruence pass. This is the agent-content analog of the corpus's flagship improvement loop — and a natural soft-judge (J) component in ledger-components.

How it wires into the machine

  • In the context pack: these are the voice/register/structure directions the generator is conditioned on — the delivery half of the anti-slop L3 layer.
  • In the video cell (video-production-pipeline): the prosody block becomes explicit VO/TTS direction; the on-screen enumerate/contrast becomes the visual layer; the story-before-content shapes the Hook→Problem→Demo→Proof→CTA fill.
  • As a soft-judge: the record-and-review passes (filler / prosody / say-show congruence) are advisory J-checks on a draft — recorded, never blocking (ledger-components).
  • The standard: "sounds right AND feels right" (the-practice-engine) — for agent content that means prosody direction + structure + emotional framing, not just accurate words. Delivery is the multiplier on the message (the perception gap, communication).

Parent: communication.