Captions — FAQ & troubleshooting — VidVertex Docs

Captions

Transcription is slow. It runs on your CPU. Long videos take proportionally longer — the transcript is cached per stack afterwards, so it's a one-time cost per clip.

"Speech model not available — no internet connection." The standard model ships with the app; a larger model you selected must be downloaded once. Either go online for the first run, or switch back to the bundled model in the Model dropdown on the Content tab. A render you confirm anyway runs without captions and says so in the log.

A name is spelled wrong in the captions. Fix it in the transcript editor (Content tab). Your correction keeps the timing and survives re-renders — the transcript is never silently re-transcribed over your edits.

TIP

The transcript is the basis for more than the captions — the silence cut, the music ducking and the caption keyword pick all read the same word timings. Fixing it once fixes all of them.

Captions come out in the wrong language. Check three places: the transcription language (if the auto-detection guessed wrong, set it explicitly and re-transcribe), the brand's content language (that's what translation targets), and the hook/endcard language pools — a brand whose language has no pool falls back to the flat list, which may be in another language. The Hook tab and the render confirmation both warn about that gap.

NOTE

A caption style chosen in the stack editor wins over the brand's pinned style — the pin is only the default for stacks that made no choice. Clear the stack's style to get the brand look back.