Subtitle & Caption Workflow: Timing, Formats, Specs

By BlockReel Editorial Team Guides, Post-Production
Subtitle & Caption Workflow: Timing, Formats, Specs

Subtitles and captions are the last text assets a film has to get right, and they are the ones most often built in a panic during delivery week. They decide whether a deaf viewer can follow the film, whether a distributor accepts the package, and whether the story travels past its original language. This guide covers the practical side: how timing actually works, which formats belong to which delivery, and how to keep dozens of language versions straight without shipping the wrong file. For the complete overview of the finishing pipeline, see the Deliverables & Archiving Masterclass: Mastering, Localization, and LTO.

Bong Joon-ho put the cultural case plainly at the 2020 Golden Globes: "Once you overcome the one-inch-tall barrier of subtitles, you will be introduced to so many more amazing films." The technical case is less quotable but just as real. A subtitle that lands two frames late, runs across a hard cut, or asks a viewer to read 28 characters per second breaks immersion exactly like a bad edit does. Captions carry a second obligation on top of that: in the United States, caption quality for televised and internet-delivered programming is regulated, not merely encouraged.

Subtitles, SDH, Closed Captions, and Open Captions

These four terms get used interchangeably in casual conversation and mean four different deliverables on a spec sheet.

Subtitles translate spoken dialogue (and, where needed, plot-relevant on-screen text) for a viewer who can hear the audio but does not speak the language. They assume the viewer hears music, effects, and tone of voice, so they carry no sound descriptions.

SDH (Subtitles for the Deaf and Hard of Hearing) are subtitle-formatted files that add what a hearing viewer gets from the mix: speaker identification, relevant sound effects, and music cues. SDH exists because streaming platforms deliver text as subtitle-style files rather than broadcast caption data, so SDH is the streaming-era equivalent of closed captions. It is rendered by the player like any other subtitle track.

Closed Captions (CC) are the broadcast lineage: caption data carried in or alongside the video signal under the EIA/CEA-608 (legacy analog and SD digital) and CEA-708 (digital) standards, switchable on and off by the viewer. They include dialogue, speaker IDs, sound effects, and music cues.

Open Captions (OC) are burned into the picture and cannot be turned off. They are common on festival screeners, social cutdowns, and territories where the exhibitor cannot handle a separate text track. Because they live in the image, every change means a re-render of the master, and every language means another render. That dependency is why textless masters matter, as covered in The Localization Pipeline: Textless Elements, M&E Mixes, and Deliverables Planning.

For most 2026 deliveries you will produce an English SDH file for accessibility, translated subtitle files per territory, and a broadcast caption file only if a linear broadcaster is in the chain.

The Compliance Floor You Cannot Negotiate

Accessibility text is not a courtesy line item in the US market.

* The CVAA and 47 CFR 79.4. Programming shown on television with captions must be captioned when it is later distributed over the internet by the video programming distributor. That is the rule that pulls captions into streaming deliveries for anything with a broadcast life.

* FCC caption quality standards. In its February 2014 Report and Order, the Commission adopted four quality benchmarks for televised captions: accuracy, synchronicity, completeness, and placement. Those four words are the vocabulary a QC vendor will use when it rejects your file, so build your internal QC around them.

* ADA exposure. Streaming platforms and distributors carry their own accessibility obligations, which is why their delivery specs are stricter than anything the FCC requires. Contractual specs, not statutes, are what will actually bounce your file.

Practical takeaway: treat the distributor spec as the governing document, and treat the regulatory floor as the reason that spec exists.

Transcription: The Foundation

Transcribe from the final audio mix after picture lock. Transcribing from an earlier cut guarantees rework, because every timing decision downstream is anchored to frame numbers that will move.

* Human transcription remains the reliable baseline for overlapping dialogue, heavy accents, technical jargon, and dense production sound. A trained transcriber makes judgment calls about punctuation and speaker attribution that materially change readability.

* AI-assisted transcription (Whisper-based tools, Rev, Trint, Resolve's built-in transcription) is a legitimate first pass and can cut hours off a documentary with fifty interview subjects. It is a draft, not a deliverable. Accuracy collapses on crosstalk, non-standard accents, proper nouns, and any scene where music sits close to dialogue.

* Do not paraphrase. Netflix's English style guide is explicit on this point for SDH: keep transcription as close to the audio as possible, and when reading speed forces a cut, reduce and condense rather than rewrite.

A working rule: budget AI transcription at roughly 60 to 80 percent usable on clean dialogue, and near zero on anything with layered sound. Human cleanup time is the real cost, not the transcription fee.

Timing (Spotting): The Numbers That Actually Matter

Timing is where amateur subtitle work becomes visible. Vague advice about "comfortable reading speed" is useless; here are the parameters professional specs are written in. Netflix's published Timed Text Style Guide is the most widely mirrored reference in the industry, so its numbers are a defensible default when a distributor has not given you their own.

ParameterCommon professional default
Reading speed, adult programmingup to 20 characters per second
Reading speed, children's programmingup to 17 characters per second
Character limit per line42 characters
Lines per subtitle event2 maximum
Minimum duration5/6 of a second per event (20 frames at 24 fps)
Maximum duration7 seconds per event
Minimum gap between events2 frames, at every frame rate
Note that the industry moved off words per minute years ago. Netflix removed words-per-minute language from its English guide in favour of characters per second, because CPS accounts for word length and works across languages. If a vendor quotes you a WPM target, they are working from an old template.

Beyond the raw numbers:

* In-times sit on the first frame of audio, or within a frame or two of it, read off the waveform rather than by ear.

* Out-times want breathing room. Where no subtitle follows immediately, holding roughly half a second past the end of dialogue reads far more comfortably than cutting on the last syllable.

* Shot changes win ties. When dialogue does not cross a cut, the subtitle should not cross it either. A subtitle straddling a hard cut reads as a glitch. Where dialogue starts on or just after a shot change, pull the in-time back to the first frame of the new shot.

* Chaining. When the gap between two consecutive events falls under half a second, extend the earlier out-time to two frames before the next in-time. A run of even gaps looks calm; a run of ragged 4-frame and 9-frame gaps flickers.

* Line breaks follow grammar. Break after punctuation, before conjunctions, and never between an article and its noun. Keep each line of a dual-speaker event as a self-contained clause.

* Frame rate is part of the file. A 23.976 file conformed against a 25 fps master drifts progressively and will fail QC. Confirm frame rate and whether timecode is drop-frame before the first event is placed.

๐Ÿ’ก Pro Tip: Give translators the picture, not just the script. Sarcasm, off-screen reactions, and who a line is aimed at are frequently invisible in text. Script-only translation is the single most common source of subtitles that are technically correct and dramatically wrong.

Format Reference

There is no universal format. The distributor's spec sheet decides, and one film routinely ships four or five different encodings of the same timed text.

* SRT (SubRip). Plain text, sequential numbering, comma-decimal timecodes, no positioning or styling. Universally supported, which is why it is the default for YouTube, Vimeo, and internal review. Never assume it satisfies a professional delivery.

``` 1 00:00:05,000 --> 00:00:08,250 This is the first subtitle event.

2 00:00:09,500 --> 00:00:12,000 And this is the second. ```

* WebVTT (.vtt). The HTML5 web standard. Similar shape to SRT with period-decimal timecodes, plus cue positioning, styling, and voice tags. The default for HTML5 players and HLS/DASH sidecar text tracks.

* TTML / DFXP (.ttml, .dfxp, .xml). The W3C XML standard for timed text. Rich styling, positioning, and metadata. It is the base layer nearly every professional profile builds on.

* IMSC (.ttml). The current W3C interoperable profile of TTML (Text and Image profiles) and the format most large streaming platforms now ask for by name.

* iTT (iTunes Timed Text, .itt). Apple's TTML-derived profile for the Apple ecosystem. Widely accepted by other platforms too, which makes it a common intermediate master. It is `iTT`, not `ITP`, on a spec sheet.

* SMPTE-TT (SMPTE ST 2052-1). The SMPTE profile of TTML for professional and broadcast delivery, including the mechanism for carrying legacy 608 data inside timed text.

* EBU-TT and EBU-TT-D. The European Broadcasting Union's TTML profiles, EBU-TT for exchange and archive, EBU-TT-D for distribution to players.

* SCC (Scenarist Closed Caption). A hex representation of CEA-608 caption data, still the workhorse for North American broadcast delivery. Frame-rate specific and unforgiving.

* MCC (MacCaption Closed Caption). Telestream's format for CEA-708 data, used where 608 alone cannot carry the required services.

* STL (EBU N19 subtitle exchange, .stl). The long-standing European broadcast binary format. Do not confuse it with Spruce STL, a different text-based format sharing the extension.

* ASS/SSA (Advanced SubStation Alpha). Enormous styling and animation control, originating in fansubbing and karaoke. Excellent for open-caption design work, not a professional delivery format.

* Embedded 608/708. Caption data multiplexed into the essence itself, for example as ancillary data in an MXF or in an MPEG transport stream. Required by most linear broadcasters.

๐Ÿ’ก Pro Tip: Build one authoritative timed master, ideally in a TTML-family format such as iTT or IMSC, and generate SRT, VTT, and SCC from it. Editing four sidecar files in parallel is how a corrected typo ships in three languages and not the fourth.

Version Control That Survives Twelve Languages

Once a film has an English SDH, six translated subtitle tracks, a broadcast caption file, and two rounds of distributor notes, informal file management fails. This is the same discipline described in the On-Set Data Policy: Who Owns Drives, Backups, and Deliverables, applied to the smallest and most easily lost assets in the package.

Naming Conventions

A filename should tell you everything without opening it:

`{Title}_{ISO639-1 language}_{type}_{purpose}_{version}_{date}.{ext}`

* Language: ISO 639-1 codes (`en`, `es`, `fr`), with region where it matters (`es-419` for Latin American Spanish, `pt-BR` for Brazilian Portuguese). Regional variants are where most mix-ups happen. * Type: `SUB`, `SDH`, `CC`, `FN` (forced narrative). * Purpose: `NETFLIX`, `BCAST-US`, `FEST`, `YT`. * Version and date: `v03_20260410`.

Example: `MyFilm_en_SDH_NETFLIX_v03_20260410.itt` and `MyFilm_es-419_SUB_v02_20260408.srt`.

Never ship a file named `final`, `final2`, or `FINAL_USE_THIS`.

Single Source of Truth

Nominate one storage location as authoritative and make everything else a copy. Caption files are tiny, so there is no excuse for them living only in an editor's Downloads folder or a translator's inbox. Include them in the 3-2-1 backup discipline you use for masters: three copies, two media types, one offsite.

Git works genuinely well for SRT, VTT, TTML, and other text-based formats. Line-level diffs show exactly which cue changed, blame shows who changed it, and reverting a bad translation pass takes seconds. Binary formats such as SCC and MCC do not diff usefully, so version those by regenerating from the text master rather than hand-editing.

Change Log and Approvals

Maintain one tracking sheet per title with a row per language: current version, source of the last change, who approved it, delivery date, and open notes. When a distributor sends corrections at 6pm on a Friday, that sheet is what stops someone patching v02 while v03 is already in the pipeline.

Archive the Sources, Not Just the Deliverables

Archive the final approved file for every language and format, plus the upstream material that makes regeneration possible: the original transcript, the project files from the captioning tool (for example `.cap` from CaptionMaker or the Resolve project with its subtitle track), the glossary or style sheet used by translators, and a short specification document recording what each deliverable was built to. Five years later, a new platform will want a format that does not exist yet, and the source files are what make that a one-day job instead of a full rebuild.

๐Ÿ’ก Pro Tip: Store a language-neutral dialogue list alongside the timed files. A timecoded, untranslated transcript of every line and sound cue is the cheapest asset in the archive and the fastest route to a new language version.

Where This Sits in the Post Schedule

Captioning is not a delivery-week task. It is a scheduled department with hard upstream dependencies.

Handoff Points

* Picture lock. The picture-locked master is the prerequisite. Any picture change after spotting begins invalidates the timings, and re-conforming a 1,400-event file is not a trivial fix. The realities of lock slipping are covered in Picture Lock Reality: Contracting and Scheduling Sound When Lock Moves.

* Final mix. SDH needs the final mix to describe sound accurately, and the levels in that mix determine which effects are worth captioning at all. If the mix is still moving, see Dialogue Editing Workflow: Cleaning, Matching, and Continuity and Loudness and Compliance: How to Avoid Rejections and Too Quiet, Too Loud Notes for what has to settle first.

* Textless elements and M&E. Localization needs textless picture for on-screen text replacement and an M&E mix for dubbing. See M&E Deliverables: How They're Built and Why Distribution Requires Them.

Rough Scheduling Guide

For a 100-minute feature, plan on transcription and English SDH origination taking several working days including QC, translation of each additional language taking a comparable span per language when done by a professional subtitler working from the English template, and a full real-time watchdown per file before sign-off. Running languages in parallel across vendors compresses the calendar; running them through one person does not.

Budget accordingly. Per-minute rates vary widely by language pair and turnaround, so get quotes early rather than assuming captioning is a rounding error against the finishing budget.

Tools

* Dedicated subtitling software: EZTitles, Ooona, Telestream CaptionMaker and MacCaption, and Subtitle Edit (free and open source, and genuinely capable for SRT/VTT/TTML work). * Free spotting tools: Aegisub remains useful for ASS/SSA styling and rough spotting, though development has slowed. * NLE-native: DaVinci Resolve and Adobe Premiere Pro both handle creation, timing, and export of common formats directly on the timeline, which is efficient for a single-language deliverable and awkward for a twelve-language matrix. * Transcription services: Rev, Trint, and Whisper-based local tools for first-pass text.

QC: What a Watchdown Actually Checks

A file is not finished when it exports. It is finished when someone has watched the whole film with it running.

* Linguistic pass. Spelling, grammar, punctuation, consistency of character names and invented terminology, and for translations, a native-speaker review for register and cultural fit. * Timing pass, in real time. Not scrubbed. Watch for events that flash, linger, cross cuts, or land ahead of a joke and kill it. Spoiling a punchline by a second is a real and common failure. * Technical pass. Character limits, line counts, frame rate, timecode start, character encoding (UTF-8 unless the spec says otherwise), and format validation against the target profile. * Accessibility pass. For SDH and CC: are speakers identified where the visual does not make it obvious, are plot-relevant sounds captioned, are music cues present, and is anything captioned that a hearing viewer would not actually be able to hear. * Placement. Move events off burned-in text, hardcoded titles, and faces where the format allows it. Placement is one of the FCC's four quality criteria for a reason.

The same QC mindset applied to sound is laid out in QC for Audio: Phase, Sync, Distortion, and Dialogue Intelligibility, and the wider finishing checklist lives in The Ultimate Film Finishing & QC Guide 2026 (Conform to DCP).

Localization Notes

Transcreation Over Translation

Literal translation fails on humour, idiom, wordplay, and register. Professional subtitlers adapt meaning to fit both the culture and the character limit, which is a different skill from document translation. Supply a glossary of proper nouns, invented terms, and preferred renderings before the first language starts, or every subsequent language will re-litigate the same decisions.

Forced Narrative Files

Forced narrative (FN) subtitles cover only the text a viewer must read regardless of their language setting: foreign dialogue inside an otherwise English film, location cards, signs that carry plot. They are a separate deliverable from the full subtitle track and are frequently forgotten until a distributor asks for them.

Templates

For multi-language work, originate a timed English template first and have every translator work into it. The timings stay consistent across languages, the QC burden drops, and a timing correction propagates instead of being fixed eight times.

Common Mistakes to Avoid

* Starting before picture lock. Every frame that moves after spotting is rework. * Treating AI output as a deliverable. It is a draft with confident-sounding errors in exactly the places that matter. * Working from words per minute. Modern specs are characters per second. * Ignoring the distributor spec. Ask for it in writing before origination, not after rejection. * Editing sidecar files independently. Corrections diverge across formats. * Skipping the real-time watchdown. Reading a file is not the same as watching it. * Forgetting forced narratives and regional variants. `es` and `es-419` are not interchangeable. * Leaving no source files in the archive. Format requirements will change; the ability to regenerate is the asset.

Interface & Handoff Notes

What You Receive (Upstream Inputs)

* Picture-locked video master at the delivery frame rate, with visible timecode burn-in on the review copy. * Final audio mix, plus the M&E mix if localization is in scope. * Dialogue script or continuity script, ideally as-broadcast rather than a shooting draft. * Textless elements where open captions or localized graphics are required. * Distributor specification documents for every territory and platform in the deal.

What You Deliver (Downstream Outputs)

* English SDH file in the specified profile. * Translated subtitle files per language and territory, including regional variants. * Forced narrative files where applicable. * Broadcast caption files (SCC, MCC, or embedded 608/708) where a linear broadcaster is in the chain. * Open-caption picture master if requested. * QC report covering the linguistic, timing, technical, and accessibility passes. * Archive package with source transcripts, project files, glossary, and the specification each deliverable was built to.

Top 3 Failure Modes

1. Picture changes after spotting. A single inserted shot shifts every event downstream of it and forces a full re-conform and re-QC.

  • Wrong spec, wrong format. Delivering 25 fps SCC against a 23.976 master, or SRT where IMSC was specified, means automatic rejection and a delivery date missed.
  • Errors that reach the audience. Typos, mistranslations, and unreadable timings surface in reviews and support tickets, and pulling and replacing a text track across a live catalogue is far more expensive than one more watchdown.

    Browse This Cluster

    - Deliverables & Archiving Masterclass: Mastering, Localization, and LTO

  • The Localization Pipeline: Textless Elements, M&E Mixes, and Deliverables Planning
  • M&E Deliverables: How They're Built and Why Distribution Requires Them
  • The Ultimate Film Finishing & QC Guide 2026 (Conform to DCP)
  • QC for Audio: Phase, Sync, Distortion, and Dialogue Intelligibility
  • Dialogue Editing Workflow: Cleaning, Matching, and Continuity
  • Loudness and Compliance: How to Avoid Rejections and Too Quiet, Too Loud Notes
  • On-Set Data Policy: Who Owns Drives, Backups, and Deliverables

    Next Steps

    Ready to see how this fits into the bigger picture? Start with the complete guide.

    ๐Ÿ“š Complete Guide: Deliverables & Archiving Masterclass: Mastering, Localization, and LTO

    ---

    Sources: Netflix Partner Help Center, Timed Text Style Guide (General Requirements, Subtitle Timing Guidelines, English (USA) guide); FCC Report and Order FCC 14-12 on closed captioning quality; 47 CFR ยง 79.4.

    ---

    ยฉ 2026 BlockReel DAO. All rights reserved. Licensed under CC BY-NC-ND 4.0 โ€ข No AI Training.

  • Originally published on BlockReel DAO.