Most people will see your video muted
Video in a feed plays silently. What that means for the edit, subtitles and the first three seconds — and why it is solved on the shoot, not in post.
Video on social media autoplays, and it autoplays silently. A minority of viewers turn the sound on, and they do it only once something has caught them. In practice that means the opening seconds have to work with no audio at all — and that a talking head without subtitles communicates precisely nothing at that moment.
None of this is news. Yet video is routinely delivered as though someone were watching it in headphones with full attention.
Subtitles are not an extra
A subtitle is not an accessory for deaf and hard-of-hearing viewers, though it serves them too. It is the primary carrier of the message for most of the audience. Without one, a video communicates only what is visible — for an interview shot in an office, that is the office.
Several practical consequences follow. The subtitle has to be legible on a phone, so large enough and contrasty enough. It must stay out of the bottom fifth of the frame, which the player interface covers. And it has to be timed to speech rather than dumped in paragraphs — a two-line block that disappears before it can be read is worse than none.
The first three seconds decide, without sound
Openings are often built on a line somebody says. Muted, that leaves a person opening their mouth. The reverse works: start with movement, a detail or a situation that raises a question on its own, and bring speech in afterwards.
For commercials it is harsher still. Ten seconds is a commonly bought length, and within it there is no room to build atmosphere first. The message has to be legible from the first second — including whose message it is.
The brand must be seen, not heard
When the company name appears only in a closing audio signature, most viewers never learn it. So the logo belongs in the picture — not necessarily across the whole frame; discreet and continuous is enough.
The same goes for the product. If the viewer is meant to grasp what the video is about, they have to grasp it with their eyes. The audio track is a bonus, not the load-bearing structure.
What this means for the shoot
None of the above can be reliably retrofitted in the edit suite. Compositions with room for a subtitle, vertical variants of the key setups, product close-ups usable as an opener — all of it is shot, not generated.
In practice it adds one line to the shooting script: a note on each setup saying whether it will run muted and whether it will be cropped to portrait. That costs half an hour of preparation and saves a reshoot.
And where sound does decide
The opposite extreme exists too. Video playing on a reception screen, a trade-fair stand or in a cinema is listened to — and there, unintelligible or uneven sound is the first thing a viewer notices. So the mix is made per target medium; European television asks for a different loudness from social platforms, and a platform will normalise a wrong level in its own way.
One video for every medium sounds like a compromise. In reality it is one master and a few derived versions — an hour’s work, provided somebody planned for it.
