Generate On-Screen Captions from Video
Use the browser's built-in AI to draft short descriptions of visible video content, then edit and style them.
Generate on-screen captions looks at video frames and drafts short descriptions of what is visible. It complements speech-to-text: it describes the picture rather than transcribing dialogue.
1. Select a video clip
Add a video to the timeline, select the clip, and expand Generate on-screen captions in the selected clip's tools panel.
2. Check browser support and generate drafts
The tool uses Chrome's built-in Prompt API and an on-device model. If it reports that the model is downloadable, its first use prepares that model on your device. If the tool says the API is unavailable, use a supported Chrome setup or choose another way to write captions.
Select Generate captions and wait for the model to review the clip. It creates short caption drafts with time ranges. Read each draft carefully: the model can misread a frame or describe details that are not useful to viewers.
3. Edit, style, and add to the timeline
Edit the caption text in the draft fields. Choose Editable text for text clips you can restyle in the editor, or SVG graphic for image-based caption cards. When using generated styles, the selected style is shared across the captions for this clip. Adjust font, size, colors, background opacity, placement, and alignment, then choose Add text clips to timeline or Add SVG clips to timeline.
This tool does not create a spoken-word transcript. For dialogue and narration, use Speech to Text (Whisper) instead.
Ready to start?
Apply what you just learned directly in Video Weaver.