Skip to content

Subtitles

Subtitles are the reason this app exists. Getting a media server to pick the right one, without being asked twice, is the whole job.

Verdict What it is
Full The complete dialogue. What most people mean by “subtitles”
SDH Subtitles for the deaf and hard of hearing — full dialogue plus sound effects, speaker names and music cues
Foreign Language Covers only the parts spoken in another language. Usually what a “forced” track really is
On-Screen Text Translates signs, letters and captions in the picture, not speech
On-Screen SDH Text On-screen text with SDH conventions applied
Commentary A subtitled director’s or cast commentary

Audio tracks get their own set: Main, Alternate, Commentary and Descriptive (audio description for blind and partially sighted viewers).

The dominant signal is caption density — how much text a track carries per minute of runtime. A full track has a cue every few seconds across the whole video. A foreign-language track has a few dozen cues clustered around two scenes. The gap between those is enormous and does not depend on anyone having labelled the track honestly.

SDH is separated from Full by its conventions rather than its size: bracketed sound effects ([door creaks]), music notes on lyric lines, and speaker names in capitals followed by a colon. Commentary tends to give itself away in the first minute, because commentary participants introduce themselves and ordinary dialogue essentially never does. A greeting alone, like “hi, everyone”, isn’t enough, because characters walk into scenes saying that too. Where there is no introduction it shows in the vocabulary instead — crew roles, talk about “the character” or “the shot”, a take being called on set. A commentary track written with SDH conventions, speaker names and (GUNFIRE) included, is called Commentary and gets the hearing-impaired flag as well as the commentary flag — it is both, so a media server can offer it as both.

Once a file is matched to a film or series, its cast and crew help as well. People recording a commentary talk about each other by their real names, while the film’s own dialogue calls the same people by their characters’ names. A track that keeps naming several of the people who made the film reads as commentary. Names that are also a character’s are ignored, and so is anyone appearing as themselves. The names only ever count towards commentary, and they never relabel the only dialogue track in a language. This needs a match in the rename pane, so it doesn’t run when working offline. With OMDb, which lists no characters, only full names count.

Every verdict carries a confidence and a list of the evidence behind it, visible in Organise’s panel. A weak call is shown as weak rather than quietly presented as fact.

Some calls are only possible in context. Two English tracks where one has four times the cues of the other tells you more than either does alone, so tracks are also compared against their siblings before the verdicts settle.

That comparison also settles a forced track in an episode with a lot of foreign dialogue. On its own, a forced track that busy can read like full subtitles. When it’s labelled forced and another track in the same language has at least twice as many lines, it’s taken at its label: the other track is the full dialogue, so this one can only be part of it.

Some files carry the same subtitle track twice, usually because their source listed one stream in two places. A track with the same language, the same lines at the same times and the same text as an earlier one is badged Duplicate and set to be removed. The earlier track is kept, so nothing is lost. For image tracks the text is compared on the OCR sample; if OCR cannot run, identical timing alone decides it.

Two versions of one track can still count as copies when they break a line into cues differently, for example one line shown as two cues on one track and as one on the other. The Duplicate evidence says where this happens. A track with a line the other doesn’t have is never a duplicate, because removing it would lose that line.

Some files carry a US and a UK English version of the same subtitles, with the same lines at the same times but “meters” on one and “metres” on the other. A copy like that is still suggested for removal, but it isn’t removed without asking. It’s shown as Suggested, and the detail panel lists the lines that are worded differently, with the changed words highlighted. Choose Keep or Remove for whichever version you want.

For image subtitles, only a sample of each track’s lines is read, so a difference can hide in the lines that weren’t. A copy that breaks a line into cues differently was made separately rather than listed twice. If not every line could be compared, it’s shown as Suggested too, even when every line that was read matches.

This distinction affects almost everything else.

Text subtitles — SRT, ASS/SSA, VTT, and Matroska’s internal text formats — are literally text. Reading them is cheap and the words are exact.

Image subtitles — PGS and VobSub — are pictures of text. There are no words in the file, only a bitmap per cue. To read them at all, the images have to be run through OCR.

That has three consequences:

  1. They are read with Tesseract, the OCR engine, which is included in the installer. It ships with English language data, which reads Latin-alphabet subtitles; scripts such as Japanese, Chinese and Korean are left to the track’s language tag. If OCR cannot run at all, these tracks are classified on their timing alone. Timing still separates a forced track from a full one fairly reliably, but it cannot tell SDH from Full, because that difference is entirely in the words.
  2. They are slow. OCR is the single most expensive thing the app does. File Verdict OCRs an evenly spread sample of cues rather than all of them, which is enough to classify a track without reading a two-hour video front to back. The evidence panel says how many were sampled.
  3. They cannot be edited line by line. The Edit subtitles cue editor only works on text tracks. There is nothing to edit in a bitmap, until you convert it to text.

Inspect Subtitles, in the track panel, shows a sample of the track’s actual cues with their timings, and lets you reclassify from there. For an image track these are the OCR results — the quickest way to confirm OCR is working at all, and to see how well.

Edit subtitles opens every cue of a text track and lets you delete lines. The usual targets are a “Subtitles by SomeGroup” credit or an advert baked into a downloaded file.

Picture subtitles make Plex transcode on many devices: the player can’t draw them itself, so the server burns them into the video as it plays. Text subtitles play as they are.

Convert to text, in a picture track’s panel, reads the text out of every picture, not just the sample a scan reads, and makes a text (SRT) track from it. That text track takes the picture track’s place:

  • it’s written into the file instead of the picture track when you Apply or encode, so nothing changes until then
  • it keeps the picture track’s verdict, its name, any Keep or Remove or flags you set, and any timing you set with Sync subtitles, and if the picture track was the default subtitle, the text track is instead
  • its panel says which track it was read from, and the picture track’s panel says it has been replaced

A film’s track takes about half a minute: most of that is taking the track out of the file, the rest is reading it, which the panel counts as it goes. You can carry on with other files meanwhile.

Text recognition isn’t 100% accurate. Hover over Convert to text and its tooltip says so, and so does the text track’s panel afterwards. Checked line by line against the pictures on a film’s track, about 97 words in 100 came out right, and about one line in seven had a word wrong. FileVerdict fixes the mistakes it makes every time: a capital I read as “|”, a two-line subtitle run into one line, and a line split in two where the picture was redrawn. What’s left is the odd misread word, “walt” for “wait” or “Its” for “It’s”, which can’t be told apart from a real one that way. The text track can be opened in Edit subtitles like any other, so it’s worth a look there before you apply.

Only English picture subtitles can be read for now. Tesseract, which does the reading, is included with English only; other languages need their own recognition data, which isn’t included yet. For a track in another language the button is there but can’t be used.

Each file in Organise has + Add subtitle beneath its tracks. Choose Upload file… and pick an .srt, .ass, .ssa, .vtt or .sub file. The language is guessed from the filename.

The subtitle is classified as soon as it is added, exactly like an embedded track, and appears as a normal row with its own verdict, evidence and Keep/Remove control. It is muxed into the file when you apply.

A downloaded subtitle file is often timed for a different release, so every line appears too early or too late. Sync subtitles lines it up while you watch it play: see Syncing subtitles.

In Settings → Naming & Tracks:

  • Scan and read subtitle track contents — off makes scans near-instant, but they only look at each track’s label, not its contents, so verdicts are less accurate until you read the tracks. See Scanning speed
  • Prefer SDH subtitles over full subtitles by default — which one becomes the default track when both exist
  • Remove tracks not in your preferred language
  • Remove duplicate subtitle tracks — on by default; see Duplicate tracks
  • Keep only one subtitle track per language — stricter: one track for each language and verdict, whether or not the others are copies
  • Remove commentary tracks by default