Script-anchored captions and dual text

Captions that follow the words instead of the clock: how takes, alignment and dual text keep subtitles correct through re-records.

The problem with timecoded captions

A timecoded subtitle file is correct until the audio changes. Re-record one line and every caption after it is offset — which is why weekly formats quietly rot.

Words as the coordinate system

Hypit aligns each take to word positions, then resolves captions from that map. Change a take and the captions move with it; no retiming, no re-export.

Dual text

A line can carry two forms: what is spoken and what is displayed. Use it for pronunciation, brand names, translated captions or tighter on-screen wording.

Step by step

  1. 1. Write segments, not timings

    Divide the script into segments and keep one or more takes per segment. You never type a timestamp.

  2. 2. Mark what should differ on screen

    Write dual text where the spoken and displayed versions should differ, e.g. <D | Dee> so captions read the way you want them read.

  3. 3. Bind the caption track

    The caption track renders from the script, so styling comes from the stylesheet rather than from per-line overrides.

  4. 4. Re-record freely

    Swap a take and re-run. Alignment rebuilds, captions re-position, and everything anchored to selections follows.

After this

Once captions are script-anchored, the caption track becomes a design decision rather than a post-production chore.

Keep reading

FAQ

Do I need exact transcripts?

No. Alignment is per take; you provide the script and the delivery, not a manual transcript.

Can captions be styled per word?

Yes — because selections are named ranges of words, emphasis styling can be anchored to the same ranges that drive b-roll and effects.

How does this help multi-language work?

Localised captions come from a localised script; the pacing and styling stay in the stylesheet, so variants do not drift.