Add the subtitles
Drop a .srt or .vtt file, choose one from your device, or paste the text.
Subtitles to text
Subtitle files are full of numbers, timing lines and tags. Paste an SRT or VTT file and get readable text back, with timestamps kept or removed and repeated auto-caption lines cleaned up. Nothing is uploaded.
Runs in your browser. Files are not uploaded.
Review the transcript before you export
Drop a .srt or .vtt file, choose one from your device, or paste the text.
One line per subtitle, the same with timestamps, or paragraphs that break where the speaker pauses.
Copy the text into your notes, or download it as a .txt file.
WEBVTT
Kind: captions
Language: en
00:00:00.000 --> 00:00:02.310 align:start position:0%
welcome<00:00:00.480><c> back</c><00:00:00.960><c> to</c><00:00:01.200><c> the</c><00:00:01.440><c> show</c>
00:00:02.310 --> 00:00:02.320 align:start position:0%
welcome back to the show
00:00:02.320 --> 00:00:05.200 align:start position:0%
welcome back to the show
today<00:00:02.800><c> we're</c><00:00:03.100><c> talking</c><00:00:03.400><c> about</c><00:00:03.700><c> attention</c>
00:00:08.400 --> 00:00:11.000 align:start position:0%
so<00:00:08.700><c> what</c><00:00:08.900><c> actually</c><00:00:09.300><c> helps</c>welcome back to the show
today we're talking about attention
so what actually helpssrt to txt
The same file can become a clean script, a timestamped log or readable paragraphs. Pick the layout for what you'll do with the text next.
Cue numbers, timing lines, the WebVTT header, NOTE and STYLE blocks, formatting tags such as <i> and <font>, and positioning codes like {\an8} are all removed. Speaker tags such as <v Anna> become "Anna:" so you can still tell who is talking.
Line breaks inside a subtitle exist only to fit the screen, so each line becomes its own line of text, and lines in paragraph mode flow together.
Without timestamps you get a clean script for reading, translating or searching. With timestamps, each line starts with its time, such as [00:01:05], which makes quotes easy to verify against the video.
Paragraph mode joins lines into running text and starts a new paragraph after a pause of two seconds or more between subtitles. Chinese and Japanese lines are joined without spaces.
Automatically generated captions, such as YouTube's, are written as rolling lines: every cue repeats the previous line, some last only a few milliseconds, and word-level timing tags sit inside the text.
With "Remove repeated lines" on, EarScribe strips the word timing and skips any line identical to the one before it, so each sentence appears once.
[00:00:00] welcome back to the show
[00:00:02] today we're talking about attention
[00:00:08] so what actually helpswelcome back to the show today we're talking about attention
so what actually helpsReview the transcript before you export
Subtitles rarely name speakers; add names where they matter.
Auto-captions often lack punctuation; add it before publishing.
Keep timestamps for any line you plan to quote.
Keep the subtitle file; plain text loses the timing.
Keep timestamps when you'll quote from the text later; they tell you where to check the video.
Paragraph mode starts a new paragraph after a pause of two seconds or more, which often marks a new thought or speaker.
Auto-generated captions repeat each line as it scrolls. Leave "Remove repeated lines" on for those.
Choose "One line per subtitle" or "Paragraphs". Cue numbers, timing lines and formatting tags are removed, and only the spoken text is kept.
Yes. "Lines with timestamps" puts each line's start time in front of it, for example [00:01:05].
Auto-generated captions scroll: each cue repeats the previous line and adds a new one. EarScribe skips a line when it matches the line just before it, which removes the repetition and leaves everything else alone.
Yes. In paragraph mode, Chinese and Japanese lines are joined without adding spaces, and if a file shows garbled characters you can pick its encoding, such as GB18030 or Shift_JIS.