Translating one subtitle file is a small job. The reason people go looking for a way to do it is that it is never one file.
A subtitle file carries a single language, so the second language is a second file and the third is a third. Put a course or a series behind that and the arithmetic gets loud: twelve episodes in three languages is thirty-six files, each one with the same terminology decisions in it, all of them named carefully enough that you can still tell them apart in six months.
This guide covers the whole job, for any subtitle file and any place you are putting it back: what to export, how to describe the work once, the two kinds of instruction that people mix up, and what to check on the first file before you commit to the rest.
What you are working with, and what to export
Subtitles almost always live in a small text file that sits next to the video rather than inside it. Open one and you find a numbered list of cues, each with a start and end timecode and the text that appears on screen for that stretch.
You get that file from wherever the video was finished. A video editor exports it from the captions panel, in Premiere, Final Cut or DaVinci Resolve. A platform exports it from the subtitle screen of the video, in YouTube Studio, Vimeo or your LMS. A transcription tool produces one if the video never had captions to begin with.
Whichever route you take, ask for .srt. That is the format AI Glot reads, up to 4 MB for a single file. Subtitles are only text, so an hour of dialogue is a few tens of kilobytes and that ceiling is rarely the thing that stops anyone. If your download offers .vtt, .sbv or .ass instead, go back one menu, because the same export screen almost always offers SRT too.
One case is a genuine dead end. If the subtitles are burned into the picture, they are pixels rather than text and there is no file to translate. You need the caption file the editor still has, or a transcription of the audio, before anything here applies.
How many jobs is this, really
Here is the fact that decides the shape of the whole task. A subtitle file holds exactly one text track. There is no second field for Spanish and nowhere to add one, so translating a subtitle file replaces the text of each cue in place and gives you back a single-language file.
We know that one from the inside. An early version of our own planner once offered to add a second language as an extra track in the same file, which is a thing the format cannot do, and the suggestion had to be corrected at the source rather than patched later. Languages never collapse into one file. Three languages is three jobs, whatever tool you use.
Episodes behave in the opposite way, and this is the part that saves the most time. A set of subtitle files goes in as one ZIP archive, up to 200 files and 20 MB, with 4 MB for each file inside it. Subtitle files have no shared structure that has to match, so an eight-minute lesson and a ninety-minute webinar sit in the same archive happily and get the same instruction and the same glossary.
| What you have | How many jobs |
|---|---|
| One video, one language | One |
| One video, three languages | Three |
| Twelve episodes, one language | One, as a ZIP |
| Twelve episodes, three languages | Three, one ZIP each |
So the number of jobs is always the number of target languages. Everything else about volume is solved by putting files in an archive.
- 1ExportOne SRT per videoIn your editor
- 2Say what you needLanguage, and what to skipFree
- 3Read the planCues, words, exact costFree
- 4ApproveThe only step that spendsSpends credits
- 5Check one filePlay sixty seconds of itBefore the rest
- 6Upload the trackOne language, one fileWhere the video lives
Three ways to run it, and they end in the same place
The route to pick depends on how many files you have, not on how technical you are.
For one or two languages, doing it yourself in the browser is genuinely the fastest thing available and there is no reason to open a terminal. Above that, the repetition is what costs you: eight uploads, eight instructions, eight approvals, eight downloads, and the errors are clerical rather than linguistic. The German file saved over the French one is the classic.
That is where the third route earns its place. Our command line tool takes a file directly, so an AI agent on your machine can run the whole set and leave you a folder with one subtitle file per language. The loop, the prompt to hand your agent and the trap that makes a waiting script hang forever are all written up in the guide to translating YouTube subtitles into multiple languages, so there is no point repeating them here.
Which instruction goes where
This is the one part of the workflow that fails quietly, so it is worth thirty seconds of attention. There are two places to say what you want, they are read at different moments, and a sentence in the wrong one does nothing at all.
The plan instruction decides what gets translated. It is read once, against the shape of your file, and it is where scope belongs: the target language, which cues to include, anything to leave out. “Translate the text of every cue into German, and skip the credits at the end.”
The batch instructions are applied while each individual line is written. They can only describe something visible inside one line, because that is all that is in view at that moment. Line length, tone, a name that must stay exactly as it is. Batch instructions change how the text is written, never which text is chosen. You get up to 1,500 characters of them on a paid plan.
Put it the wrong way round and nothing errors. A line-length rule in the plan instruction is simply not consulted when the line is written, and “only translate the second half of the file” in the batch instructions has no file in view to apply to. Both times you get output that ignored you, with no warning that it did.
The subtitle rules worth writing down
Subtitles have constraints that ordinary prose does not, and almost all of them are line-level, so they belong in the batch instructions. Four are worth writing down every time.
A maximum line length. Under 42 characters per line is the usual working figure, and it exists because a viewer is reading while also watching. Ask for it and the translation is written to fit, rather than written long and trimmed afterwards.
A maximum of two lines in a cue. Three lines cover the picture, and on a phone they cover most of it.
Brevity over literalness where the cue is short. A cue that is on screen for a second and a half cannot carry a longer sentence, and nothing in the pipeline can extend the timing to make room. Saying “prefer a shorter phrasing when the line would not fit” gives the engine permission to make the right trade rather than the faithful one.
Any name that must not move. A product name, a character name, a feature you always leave in English.
1200:04:02,120 --> 00:04:04,600Sélectionnez ensuite le dossier de destination completChoisissez ensuite le dossier de destinationThe reason the length rule works at all is that it is a request, not a validator. Nothing counts your characters and rejects a line, which is exactly why the first translated file is worth opening rather than trusting.
Keeping a name identical across a whole series
Across an episode, grammar is rarely what goes wrong. What goes wrong is that your product is called one thing at minute four and something else at minute fifty-one, and that the person who translated episode two is not the person translating episode nine.
Put the terms that must not drift into a workspace glossary once: product names, character names, the house translation of the phrase you say in every intro. It then applies to every file in every language, including the episode you have not recorded yet.
A glossary matters more for subtitles than for most content, for a mechanical reason. Each line is written with the brief and the surrounding cues in view, which is plenty for ordinary dialogue, and not enough for a name that appears exactly once in an hour. Nothing in that single cue says the word is a product rather than a common noun. A glossary is the thing that says it, and it says it the same way in episode nine as in episode two.
This is also where a general machine translator struggles, and the mechanism is worth knowing. Its glossary substitutes a term wherever it matches the text, so the sentence around the substitution is never reconsidered. A glossary carried in as meaning while the line is written lets the term inflect into the sentence instead of being dropped into the middle of it.
What to check on the first file, before you run the rest
Translate one file. Then spend five minutes on it, because everything you find here you would otherwise find thirty-five times.
- Compare the cue count. The original and the translation should have the same number of cues, and the same first and last timecode. Cue numbers and timecodes are structure rather than text, so they are copied through untouched and one cue in gives one cue out.
- Load the track and watch a minute. Not the beginning, which everyone over-polishes. Pick a minute in the middle where somebody talks quickly.
- Look at the shortest cues. Sort by duration if your tool can, or just scan for the one-second cues. A longer wording breaks there first.
- Search the file for one name that had to stay put. One search tells you whether the glossary or the instruction actually landed.
- Read three lines out loud. Subtitles are read at speed, and a sentence that is correct but stiff is obvious the moment you say it.
If something is off, the fix is usually one clause in the batch instructions rather than a different approach. Then run the rest.
Putting the tracks back
Every player and platform works the same way: you add a language to the video, then upload the subtitle file for that language. There is no screen anywhere that takes several languages at once, which is the same constraint as the file itself, showing up at the other end.
Two habits save real time here. Name the files with a language code, in the pattern episode-04.es.srt, so the set stays sortable and a mis-upload is visible rather than mysterious. And keep the English file as the only source. Every new language starts from it, never from a translation of it, because a second-hand translation carries the first one’s compromises forward without telling you.
What a second language actually buys you
Two separate arguments, and the evidence for each is different, so it is worth keeping them apart.
Discoverability is the reason to translate. YouTube states on its own translation tools page that “on average, over two-thirds of a creator’s audience watch time comes from outside of their home area”, and that translated titles and descriptions can appear in search results for viewers who speak those languages. A video carrying one track is findable in one language. The same page is the reason to translate the title and the description as well as the subtitles: they are three lines of text next to an hour of speech, and they are what a search result is made of.
Attention is the reason to have a text track at all. In a YouGov poll of 1,000 US adults run at the end of June 2023, 63% of under-30s said they prefer subtitles on when watching in a language they already know. And captions are not a preference for everyone: the World Health Organization projects that 2.5 billion people, or one person in four, will be living with some degree of hearing loss by 2050.
Be precise about what that second group proves, because it is easy to stretch. Those figures are about captions in a language the viewer already speaks, so they argue for having a text track, not for evidence that a Japanese track earns Japanese viewers. The YouTube figure is the one that speaks to translation. Together they say something simpler than either: a text track earns attention, and a text track in someone’s own language earns attention in a market where you currently do not appear.
When this is the wrong tool
For one short video in one language, when somebody on your team speaks it, ask them. They will phrase it better than any pipeline, they know what the video is for, and it will take them less time than reading this article.
- Someone on the team speaks it
- They know what the video is for
- Twenty minutes of their afternoon
- A hundred and twenty files
- The same terms in every one of them
- And another twelve videos next month
The case for running subtitles through AI Glot is volume, and volume only. A course library, a season, a back catalogue, or a language set you have to repeat every month: that is where the file handling becomes the whole cost, and where it is worth noting that the file handling is not the part you are paying an expert for. Plenty of teams use it as a first pass with human review on top, which puts the expert time on language instead of on renaming downloads.
If you want to see what it does to one of your own files first, the subtitle translator takes a file without an account. Then create a free account when you are ready to run a set, and check what a set costs before you plan one.
The limits quoted here are the ones in force when this was written: 4 MB for a single SRT file, and 200 files and 20 MB for an archive of them. Word counts and costs always come from the plan for your own file, which is free to produce and free to correct.