Skip to content

Typing it yourself or getting suggestions

When to skip the Text Overlay analysis, how the AI builds its suggestions, the caption language, variants, regeneration and the instructions field.

In Text Overlay, the text placed on each video comes either from you or from the analysis model. Both paths combine: you can ask for suggestions, keep one, rewrite it, or type everything by hand without paying anything. Here is how to choose, and how to get the most out of the suggestions.

Typing the captions yourself

In the Analyse step, the Type the captions myself button skips the analysis. You go straight to the Text step, where each video has its field Type the caption, or pick a suggestion below. No credit is spent on the analysis; only the final render will be billed.

It is the right choice when:

  • you already know what you want to write, for example a hook reused across a whole series;
  • the video has no usable speech (music only, noise);
  • you are importing just a few videos and want to move fast.

You can also analyse only part of the batch: a video that was not analysed shows This video has not been analysed. Run the analysis in the previous step, or type the caption yourself.

How the suggestions are built

For each analysed video, the Studio sends:

  1. the audio track to Deepgram, which transcribes it (Transcription);
  2. up to eight frames downscaled to 512 px, picked to represent the video (Representative frames), to the analysis model, along with the transcript;
  3. a reading of lip movement (Reading lips) that tells the model who is speaking on screen.

The video itself never leaves. The model then writes several Suggestions for each requested type, each filed under a style: Curiosity, Humour, Question, Reaction, Situation or First person. A single analysis per video covers every ticked type.

The What the model will read block recalls, before launch, the number of calls to Deepgram and to the model, and the cost: 1 credit per requested text type, for each video. Nothing is billed twice: an already-paid analysis is found again when you come back to the session (N already-paid analyses restored: nothing is charged again).

The caption language

The Caption language setting of the Analyse step applies to the whole batch: Automatic uses the language most videos speak, French and English force the language of the suggestions, whatever the video's own. For a bilingual series, analyse the two groups separately.

Choosing, editing, asking for a variant

In the Text step:

  • Use places the suggestion in the field; it is marked In use. The text stays freely editable (Suggested by the model, freely editable.).
  • Variant asks for another wording in the same style, different from everything already suggested. If the model repeats an existing suggestion, the tool says so (No new variant).
  • Back to the chosen suggestion undoes your edits to the text.
  • Copy grabs the text to use it elsewhere.

A variant is a new request to the model: it costs 1 credit for the type involved.

The instructions field and regeneration

The Extra instructions field steers the writing: "make it more mysterious", "funnier tone", "don't mention Tinder", "first person", "add an emoji", "under 50 characters". Fill it in before the analysis, or afterwards: the Regenerate with these instructions button reruns a video's suggestions with your instructions. Like a variant, a regeneration is a request to the model, billed 1 credit per requested type.

When the model suggests nothing

If the model declines the video (The model declined to process this video.) or suggests nothing (The model suggested nothing for this video.), the credits of that request are refunded automatically. Type the caption by hand, or run that video's analysis again. A message such as The model did not answer in time: run this video's analysis again. signals a passing error, refunded as well.

See also

Was this article helpful?