Skip to content
Text Overlay

The right caption on every video, written by AI, reviewed by you

Upload your vertical videos. Clipfer transcribes, looks at a few frames, reads lips to know who is speaking, and suggests several captions per video — Snapchat band, POV, POV TikTok or a plain sentence. You choose, edit or type your own, adjust the position, and export the whole batch.

Sign-ups opening soonSee how it works

Sign-ups not open yet · opening soon

clipfer.com/montage-snapchat
Write or review each video's caption, and see it on the picture before export.
01How it works

Six steps, from upload to export

Upload, analyse, text, adjust, check, export. The analysis is paid once per video and per text kind, then the render per minute.

  1. 01

    1. Upload the videos

    One or several vertical videos. Each video shows its length, its format and, once chosen, the caption kept.

    The card of an uploaded video: name, length, format, caption kept.
  2. 02

    2. Let the AI suggest, or write

    Pick the text kinds, add instructions (“more mysterious”, “no Tinder”, “under 50 characters”), run the analysis. Five to eight suggestions per kind, sorted by angle: curiosity, situation, reaction, question, humour, first person. Or type your own text.

    The Text panel: the caption kept, the suggestions and the extra instructions.
  3. 03

    3. Adjust the position and style

    The band or the text goes where you want — top, center, bottom, or to the percent —, with a warning if the text leaves the safe area. Size, line height, colours, outline, display time: set once for all videos, or video by video.

    The vertical preview with the Snapchat band placed on the picture.
  4. 04

    4. Check and export

    Every video is checked before rendering. The text is drawn by the same engine as the preview, then burned in without touching the picture. Files come out with a neutral name.

    The step bar of Text Overlay, the caption chosen and the check pending.
02What you get

What you get

01

Four text kinds

Snapchat (translucent black band, Helvetica Neue), POV (“POV: …”, outlined TikTok Sans), POV TikTok (“pov: …” in lowercase, at the top) and plain Text. One file per chosen kind, a single analysis.

02

The AI knows who is speaking

Audio transcription, up to ten frames, and lip reading: the person on screen and the one off screen are told apart. Captions only tell what really happens.

03

The right language, the right format

Spoken language detected (or French / English forced): a suggestion in the wrong language is rejected. Each kind's format is enforced, no hashtag, ninety characters at most.

04

A whole series in one batch

Same settings for all videos, or “Apply to N videos”. Variant, regeneration with new instructions, and hand-written text whenever you prefer.

03Example result

A Snapchat band placed on the video

The text is burned onto the original video, intact underneath, and the file comes out anonymised.

Files
MP4 with a neutral name (vid_…), cleaned metadata, slightly different size at every export; ZIP with a mapping report.
Styles
Snapchat: band at 60% opacity, six lines at most. POV and Text: TikTok Sans SemiBold with an outline. POV TikTok: smaller, at the top.
Checks
Every render passes a quality check before download; an analysis already paid is found in the cache, never counted twice.
A vertical video with a black band and the caption “wait, did he really just say that?”.
The Snapchat style, as it will be rendered.
04FAQ

Questions about Text Overlay

Which text kinds?

Snapchat, POV, POV TikTok and Text. You can pick several before the analysis: each video then gives one file per kind, for a single analysis.

Does the AI make things up?

No: it gets the transcript, a few frames and who is speaking, with the instruction to invent nothing. You review, and you always have the last word.

Can I write my own text?

Yes, at any time: type it, or ask for a variant or a new series with your instructions. Two hundred and eighty characters at most.

Does the Snapchat style work everywhere?

The Snapchat render uses Helvetica Neue, which only exists on Mac and iPhone: that is where it renders. The POV and Text styles use TikTok Sans, available everywhere.

What leaves my computer?

The audio goes to Deepgram for transcription, and up to ten reduced frames go to Anthropic (Claude) for analysis. The video stays in your private space.

The right caption on your next video.

Sign-ups are opening soon.

Sign-ups opening soonI already have an account