Why Do Multilingual Virtual Conferences Lag? (And How SFU Architecture Fixes It)
By the SmartMeet team · · 4 min read
Multilingual virtual meetings lag for two main reasons: media servers that decode, mix and re-encode every stream before sending it on, and weak connections at the speaker's or interpreter's end. A Selective Forwarding Unit (SFU) removes the first problem by forwarding each stream without re-encoding it. Good equipment and wired connections remove most of the second.
Few things undermine an international meeting faster than delay. When an English speaker is interpreted into Arabic, every extra fraction of a second between the floor audio and the interpreter's ear is a fraction of a second the interpreter must hold in memory. Delay that builds up along the chain leaves listeners further and further behind the speaker's slides and gestures.
Where the delay comes from
In an interpreted meeting, the audio makes two trips: from the speaker to the interpreter, and from the interpreter to the listener. Delay added anywhere on the first trip is added again on the second. There are three main sources:
- Server processing: servers that decode, mix and re-encode media add delay at every hop.
- Network conditions: packet loss and jitter on Wi-Fi or congested links force buffering, which adds delay and cuts out syllables.
- Device and audio settings: Bluetooth headsets, heavy audio processing and overloaded laptops each add their own delay.
The ITU-T G.114 recommendation on one-way transmission time advises keeping one-way delay below about 150 milliseconds for conversation to feel natural. Interpretation is more demanding than conversation, because the interpreter's output depends on hearing the floor promptly and clearly.
MCU vs. SFU: two ways to route a meeting
- Speaker and interpreters
- Server decodes every stream
- Server mixes and re-encodes
- One composite stream to everyone
- Speaker and interpreters
- Server forwards each stream
- Each listener receives the streams they chose
An MCU receives every participant's audio and video, decodes them, composes a single layout, re-encodes it and sends the result to everyone. That design suited older hardware endpoints, but the transcoding adds delay and server load, and everyone receives the same mix, which does not suit language channels that each listener chooses for themselves.
An SFU does not decode or re-encode. It receives each participant's streams and forwards them to the participants who need them. Each language channel stays a separate audio stream, so a listener who chooses Arabic receives the Arabic interpreter and, if they want it, the floor audio quietly underneath, while a French listener receives a different combination.
| Aspect | MCU (mixing) | SFU (forwarding) |
|---|---|---|
| Server work per stream | Decode, mix and re-encode | Forward without re-encoding |
| Added server delay | Higher, because of transcoding | Lower, because there is no transcoding |
| Language channels | Hard: everyone receives the same mix | Natural: each listener receives the channels they chose |
| Weak connections | One quality for everyone | Each receiver can be sent a lower quality without affecting others |
How SFUs handle participants on weak connections
Forwarding raw streams to everyone would overwhelm attendees with slow connections. SFU platforms use techniques such as these to keep audio flowing:
- Simulcast: a camera publishes several resolutions at once, and the SFU sends each viewer the one their connection and screen can handle.
- Pausing unseen video: video that a participant is not displaying does not need to be delivered to them, which frees bandwidth for audio.
- Audio first: audio streams are small, and keeping them clean matters more than video resolution for interpretation.
Most browser-based platforms build on WebRTC, the open standard for real-time audio and video in the browser. WebRTC carries speech with the Opus codec, which is designed for low delay.
What organizers can fix on their side
Server architecture handles delivery, but much of the delay in real meetings comes from the speaker's and interpreter's own setup:
- Use wired connections. Speakers and interpreters should connect over Ethernet, not Wi-Fi.
- Use wired or USB headsets. Bluetooth adds delay and can switch to low-quality audio modes when the microphone is on.
- Close heavy applications. An overloaded laptop delays audio processing.
- Rehearse with real interpreters. Ask them whether the floor audio arrives clearly and promptly, and fix problems before the day.
SmartMeet is built on the LiveKit SFU and WebRTC. Each interpretation channel is its own audio stream alongside the floor audio, so switching language changes only what a listener receives, and listeners set how loud the original audio plays under the interpretation. For the wider planning view, see how to evaluate video conferencing platforms for large events.
Conclusion
Lag in multilingual meetings is not inevitable. Forwarding streams through an SFU instead of mixing them on the server removes a major source of delay and makes per-listener language channels natural. Wired connections, proper headsets and a rehearsal with your interpreters take care of most of the rest.
Frequently asked questions
How much delay is acceptable in remote simultaneous interpretation?
As little as possible. ITU-T G.114 advises keeping one-way delay below about 150 milliseconds for natural conversation, and interpreters need the floor audio to arrive promptly so they can keep pace with the speaker.
Why does audio quality matter more than video resolution for RSI?
Interpreters rely on clear consonants, intonation and rhythm to work accurately in real time. Choppy or muffled audio increases their cognitive load and leads to errors, while lower video resolution rarely does.
What is the difference between an SFU and an MCU?
An MCU decodes, mixes and re-encodes all streams into one composite stream for everyone. An SFU forwards each stream without re-encoding, so each participant receives only the streams they need.
How do SFU platforms handle attendees with slow internet?
They send lower-resolution video to those attendees, for example through simulcast, and avoid delivering video they are not displaying, which leaves bandwidth for audio.
- RSI Latency
- SFU Architecture
- WebRTC
- Multilingual Webinars