Skip to content

Auto captions, step by step

Transcribe and burn animated captions into several videos - pick a style, confirm the batch placement, transcribe, review, then start the render.

Auto captions transcribes your videos and burns animated captions into them, several videos at a time. Style and position are set once for the whole batch, each voice can get its own color, and you correct the transcript before anything is burnt in. Five steps: Upload, Configure, Transcribe, Review, Export.

The Configure step: batch style and placement in the frame
The Configure step: batch style and placement in the frame

Step 1: Upload

Drop your videos, a folder or a ZIP into Videos to caption. Lone audio files are skipped: only videos are captioned. Click Continue to the settings.

Step 2: Configure

The style

Under Caption style, pick a preset:

  • Neon conversation: three words at a time, a neon color per voice, glow and outline.
  • Plain: white, crisp outline, no glow or animation.
  • Impact: one word at a time, very large, thick outline.

The Appearance section refines every detail: Caption font (or Upload a font, TTF, OTF, WOFF, WOFF2), Text size, Uppercase or Original case, Outline and shadow, Neon glow, Words shown at once, Light animation (a 120 ms fade). As soon as a setting departs from the preset, the style becomes Custom. Captions always fit on a single line: a long sentence is split into several successive captions. Confirm with Confirm the batch style.

The placement

Under Placement in the frame, a sample caption sits on a frame from one of your videos. Drag it with the mouse or a finger, or use the arrow keys. Change the preview frame lets you check another moment or another video. One placement applies to every video in the batch; if the batch mixes vertical and horizontal formats, the position is applied proportionally. Click Confirm the batch placement: it is required before moving on.

Language, speakers and side files

Under Language, speakers and detailed settings:

  • Transcription language: Automatic detection, French or English.
  • Detect the speakers automatically: each voice gets a color role from its timbre (Female voice, Male voice, Uncertain voice). It is a visual cue, not an identification of people.
  • Use white when unsure: an unresolved voice stays white instead of being colored.
  • Also export the SRT and VTT files: in addition to the burnt-in captions, never instead of them. Upload SRT/VTT files reuses existing captions, matched by video name.

Step 3: Transcribe

Click Transcribe. Only the audio track of each video goes to Deepgram; the video stays in your private space. The Match the speakers section then lets you listen to a sample of each voice and set its color for the whole video. Audio already transcribed the same way is not billed again; changing the language or the speaker detection needs a new transcription, and the tool tells you so (Redo N transcriptions?).

Step 4: Review

This step is optional but worth it: you fix a word, adjust a start or an end, hide a caption, rename a speaker or replace a term across the batch. It is detailed in Correcting the transcript.

Step 5: Export

  1. Click Start rendering and confirm the cost.
  2. The video service burns the captions in without re-encoding the rest: frame rate, resolution, colors and sound stay those of the original.
  3. Download each video, a selection, or Download everything (ZIP). Retry the failures reruns the failed videos.

A video with no detected speech is exported without captions, and the tool says so. If you correct captions after a render, the warning Captions changed since this render appears: Re-render the changed video brings the archive up to date.

Cost

1 credit per minute of video, every started minute with a 2-second allowance, transcription included within your plan's minutes. A failed video is refunded.

See also

Was this article helpful?