Script-first audio and captions

Text to speech with SRT or WebVTT subtitles

Create an MP3 voice track and timed subtitle file from the same script in one generation. Download both files for your video or web player.

  • MP3 plus subtitle file
  • SRT or WebVTT output
  • One script, one generation
1

Enter text

2

Voice and delivery

1.0×
.5×
100%
0%200%
0Hz
-50+50
Matched sample files

Hear the MP3 and inspect its subtitle timing

The audio and both caption examples use the same script. Download either subtitle format to inspect it in your editor.

1
00:00:00,000 --> 00:00:04,320
Hello. This is a short voice preview from TTSMint.
WEBVTT

00:00.000 --> 00:04.320
Hello. This is a short voice preview from TTSMint.

Generate voice and captions in three steps

The subtitle option is already enabled on this page.

01

Add and review your script

Use punctuation and short sentences to make pauses and caption breaks easier to review.

02

Choose SRT or WebVTT

Use SRT for many editors and players, or WebVTT for a web-video workflow. Choose your voice and delivery settings too.

03

Generate and download both

Sign in, generate once, then download the MP3 and the separate subtitle file from the result panel.

Script-first TTS

Your submitted script is the source for both the new voice track and its subtitle cues. This avoids running a second transcription pass over the generated audio.

  • Useful for narrated videos and lessons
  • Separate MP3 and subtitle downloads
  • Same text source for audio and captions

Not speech-to-text

TTSMint does not accept an existing recording on this page. Use a transcription product when you need captions for audio or video that has already been recorded.

  • Review timing and line breaks before publishing
  • Check names and pronunciation
  • Generation requires Google sign-in and sufficient Credits

Speech and subtitle questions

What this workflow creates, which format to choose, and what to review.

How do I generate speech and subtitles together?

Enter your script, keep Generate subtitles enabled, choose SRT or WebVTT, and generate. After a successful request, the result panel provides separate MP3 and subtitle downloads.

What is the difference between SRT and WebVTT?

Both formats store timed caption cues. SRT is widely supported by video editors and players. WebVTT is designed for web video and uses a slightly different timestamp syntax.

Are the subtitles created from the same script as the audio?

Yes. TTSMint creates the speech and subtitle timing from the submitted script in the same generation request. It does not transcribe an existing recording.

Do I need to review the subtitle timing?

Yes. Always preview the files together and check line breaks, names, pronunciation, and timing before publishing. Results are generated automatically and may need editing.

Can I add subtitles to an audio file I already have?

No. This page is a script-first text-to-speech workflow, not speech-to-text. It creates new audio and captions from your text rather than transcribing an uploaded recording.