How it works — Captions — VidVertex Docs
Docs/Captions/How it works

How it works

Rendered TikTok variant for Founder Hours: the cover moment, hook 'Fix the offer before ads', clean captionsClick to enlargeRendered TikTok variant for Founder Hours: the cover moment, hook 'Fix the offer before ads', clean captions
Founder Hours ·TikTok
  1. Transcribe (Content tab): click Transcribe now. VidVertex extracts the audio and runs a local Whisper model on your CPU. No API, no upload, no cost. The result is a word-level transcript — every word with its exact timing.
  2. Style (Captions tab): pick a look, size, words per line, position.
  3. Render: captions are burned into the video as the topmost layer. Twelve brand × platform variants transcribe once — the transcript is cached per stack.
The caption band of five brands' TikTok variants at the same moment: impact punch, word pop, golden glow, karaoke neon and poster shadow looks on one clipClick to enlargeThe caption band of five brands' TikTok variants at the same moment: impact punch, word pop, golden glow, karaoke neon and poster shadow looks on one clip
Ecom Playbook · impact_punchPricing Lab · word_popOffer Desk · golden_glowFounder Hours DE · karaoke_neonFounder Notes · poster_shadow
FIG 1Five brands, five caption looks — burned into that brand's variant of one clip.

The speech model

The standard model ships with the app, so the first transcription works offline and without waiting. Larger models (for tricky audio) are downloaded once on demand — the status line names the size and shows the download progress; afterwards they're local too. The Model dropdown and the line naming where the active model comes from sit on the Content tab, next to Transcribe now.

NOTE

Captions are a bonus, never a blocker. A render whose model isn't available still finishes — without captions, with a warning in the log. So VidVertex checks before a render starts: the render confirmation (and the CLI) warns when the model is missing and the machine is offline, and notes when the first render will download it once.

Expect roughly a few seconds of processing per ten seconds of speech on a normal CPU.