SmoothyApp can transcribe a Premiere sequence or an audio/video file on your computer, then let you edit the resulting captions before export. Its local caption tool is free and does not require Studio credits.
Premiere also has its own Speech to Text workflow. This guide explains the separate SmoothyApp workflow, which is useful when you want a local Whisper model and the same caption editor for sequence audio and standalone files.
1. Install the app and prepare the source
Download SmoothyApp for Apple Silicon Mac or Windows x64. If you plan to use sequence audio, follow the Premiere connection guide, open the bundled panel and keep the desktop app running.
In the desktop app’s Captions tab, choose either the Premiere sequence source or Audio / Video File. The file source works without Premiere. For sequence audio, select the tracks that contain the speech you want transcribed. Background music and duplicate microphone mixes can make the transcription less useful.
2. Choose a language, model and engine
Set the language to match the recording, or use automatic detection. For a language other than English, choose a multilingual model rather than an English-only model.
The engine controls include Auto, GPU and CPU. Available acceleration depends on the platform and installed hardware. Auto can retry on CPU when GPU processing fails. A dedicated GPU is not required to use the CPU engine.
The first run may download the caption engine and selected model. After setup, local transcription can run without uploading the media. A larger model may need more memory and processing time; test a representative recording before committing to a long batch.
3. Generate and review
Click Generate Captions. The app extracts audio from the selected source and runs Whisper locally. Completion time depends on the source duration, selected model and machine. Cancel is available if you need to stop the run.
Listen while reviewing the output. Correct names, technical terms, punctuation and mistakes caused by overlapping speakers. Automatic word timestamps are a starting point, not a guarantee of exact timing.
You can change line length, line count and maximum cue duration, then reformat the existing transcript without running Whisper again. Use transcript-wide Find and Replace for recurring spelling errors. Check each replacement in context.
4. Correct cue timing and export
Individual captions have editable text and timing. Valid timing changes update the export. Check that every cue has a start before its end and that the words remain readable at normal playback speed.
- Save SRT writes a subtitle file you can keep with the project.
- Send to Premiere imports captions as a caption track in the connected sequence.
An SRT contains text and time ranges. It does not carry animated word styling or a finished motion-graphics design. Style the caption track in your editor after import.
Fix a mismatch at the handoff
If captions start late, first check whether your file begins at the same point as the sequence audio. Captions generated from a trimmed export use that file’s time zero. Moving them onto a different timeline may require an offset.
If words are inaccurate, confirm the language and selected audio tracks before changing models. If GPU processing repeatedly fails, try the CPU engine. For manual subtitle use outside Premiere, see the Resolve SRT workflow.