Scanning speed
If you have scanned a folder off a network share and wondered why it took minutes, this page is the answer — and the setting that fixes it.
Where the time goes
Section titled “Where the time goes”Classifying a subtitle track means reading it. In a Matroska file, subtitle data is not stored in one block at the front; it is interleaved through the whole file, next to the video it belongs with. Reading the subtitles therefore means reading the file, end to end.
That is not an implementation choice that could be improved on. There is no index that lets you skip to just the subtitles, and no cheaper route to the data — asking ffprobe for nothing more than the subtitle timestamps still takes half a minute on a large file, because it has to demux the container to find them.
So the cost of a full scan is the cost of reading every byte of every file.
A real measurement
Section titled “A real measurement”A folder of thirteen files, 36 GB total, on a gigabit SMB share:
| Time to scan | 373 seconds |
| Throughput | 96 MB/s |
96 MB/s over gigabit is the link running flat out. None of that time was wasted; it was spent moving the files across the network as fast as the network allows. A faster app cannot beat it. A faster network can.
On a local SSD the same scan is far quicker, because the constraint moves. If your library is local and scans still feel slow, the likely cost is OCR on image-based subtitles rather than I/O — see Subtitles.
The order files finish in
Section titled “The order files finish in”Every file’s rename suggestion is looked up first, all together, so the rename rows fill in within seconds. Then the files are read one at a time, from the top of the list down, and each is judged while the next is being read. So the first file is ready as soon as it has been read, and the rest follow in the order they’re listed, while the network still runs flat out.
Measured on TV episodes of about 1.7 GB each, on a gigabit share: the first was ready after 20 seconds, and the rest followed every 18 seconds or so. Six read at once, which is how earlier versions scanned, finished about 6% sooner in total, but none of them was ready until 90 seconds in.
On a local SSD, reading a file takes a second or so, and judging it is most of the work. Several files are judged at once there, so a file that’s quick to judge can finish a moment before the one above it.
The setting
Section titled “The setting”Settings → Naming & Tracks → Scan and read subtitle track contents.
Leave it on and nothing changes: every subtitle track gets a full verdict, as described above. A track whose label disagrees with what its contents read as, such as one titled “SDH” that reads as plain dialogue, is marked Suggested so you can check which is right.
Switch it off and a scan reads headers and does a title lookup, nothing more. The same thirteen-file folder above then scans in 2.4 seconds.
What you still get with it off:
- every track listed, with its language, codec and name
- audio verdicts, which are read from the headers and cost nothing
- subtitle verdicts from each track’s label, described below
- rename suggestions, which come from the filename and the metadata lookup
- the ability to remove tracks, re-flag them and rename the file
What you give up is accuracy on subtitles. With contents unread, a track is judged on what its label says:
- a title such as “SDH”, “CC”, “Forced” or “Commentary”, or the matching flag in the file, is believed and applied
- a track labelled with nothing but its language, such as “English” or no title at all, is taken as the main subtitles when it is the only one
- where several tracks are labelled with only the language, the cue counts most files record in their header are used to tell them apart: the busier of two is the SDH one, as when the contents are read
- anything the label and header cannot settle is marked Suggested, and a track whose label says nothing recognisable shows as not scanned
These verdicts show from label in place of a confidence figure, because nothing about the contents was measured. A wrong label gives a wrong verdict, which is the trade for not reading the file.
Reading tracks on demand
Section titled “Reading tracks on demand”An unscanned track is not a dead end. Each file’s row carries a read N unscanned link that reads that file’s subtitle tracks there and then, and the verdicts appear as normal.
This is offered per file rather than per track deliberately. The expensive part is opening the file at all — once that is paid, the remaining tracks in it are nearly free. Reading them one at a time would pay the same cost several times over.
Which way round to leave it
Section titled “Which way round to leave it”Leave it on if you are doing what the app is for: deciding which subtitle tracks to keep.
Turn it off if you are renaming a library, tidying audio tracks, or feeding files into Compress — and read individual files’ subtitles on demand when a decision actually needs one.
Tracks already set to be removed
Section titled “Tracks already set to be removed”Independently of that setting, a subtitle track that is already destined for removal — because it fails Remove tracks not in your preferred language, for instance — is not read at all. Its contents cannot change the outcome, so nothing is spent on them. Those tracks also show as not scanned, and are not counted as decisions waiting on you.