Live Translation Benchmark 2026

Live Translation Benchmark 2026: Mach Reaches 0.47s Average Latency

Based on current published market benchmarks and our internal testing, Mach recorded the fastest live translation latency in our comparison at 0.47 seconds on average.

5 minute read

Benchmark at a glance

0.47 seconds from phrase to translated text

Mach translated all 174 completed phrases during a continuous 15-minute live stream and accepted every audio frame.

Internal benchmark · English → German · August 2026

Why it works well

The fastest result in our market benchmark review.

Based on current published market benchmarks and our internal testing, Mach recorded the fastest live translation latency in this comparison. Separate ten-language tests also showed strong translation accuracy.

Average responseOnce a phrase was ready
0.47s
P95 responseAcross 174 final cards
1.43s
Translation coverage174 of 174 final turns
100%
Dropped audio frames44,977 frames accepted
0

Live translation latency

The fastest result in our current market benchmark review.

Published figures in seconds. Lower is faster.

Mach

Fastest

Phrase → text

0.47s

Meta SeamlessStreaming

Speech/text

~2.0s

Google Meet

Speech → speech

2–3s

Gemini Live Translate

Speech → speech

2.9s

Gradium Translate

Speech → speech

3.0s

GPT Realtime Translate

Speech → speech

3.6s

0s1s2s3s4s

Mach measures translated text after a phrase ends. Most other results measure translated speech, so this is only a rough comparison.

Live translation benchmark results

Based on current published market benchmarks and our internal testing, Mach recorded the fastest live translation latency in this comparison: an average 0.47 seconds from a completed spoken phrase to readable translated text. The P95 was 1.43 seconds, and every one of the 174 final source phrases received its matching translation.

The result matches the product experience we are building: short, readable translation cards appear continuously while the conversation keeps moving. Users do not need to stop the stream or wait for a long transcript to finish.

Mach Live Translate delivered sub-second average translation responsiveness with 100% final-turn coverage in our internal 15-minute test.

Built for fast, accurate translation

Languages do not all structure sentences in the same way. Translating too early can change the meaning.

Mach translates once a phrase is clear. The result arrives quickly and is easier to trust and read.

Mach vs Gradium, Gemini, GPT, Google Meet, and Meta

Public figures show Meta SeamlessStreaming at around two seconds, Google Meet targeting two to three seconds, Gemini at 2.9 seconds, Gradium at 3.0 seconds, and GPT Realtime Translate at 3.6 seconds. Mach's observed 0.47-second average is the fastest latency figure in this comparison.

The endpoints are different: Mach measures from phrase detection to translated text, while most competitor reports measure speech-to-speech output. That makes Mach's result especially relevant for live captions and readable translation cards, but not a direct speech-to-speech victory claim.

Reliable during a real continuous stream

Speed matters only when the words keep arriving correctly. During the internal stream, Mach accepted all 44,977 audio frames, delivered 174 out of 174 final translations in order, and dropped no audio because of queue pressure.

Separate conversational quality tests across ten languages and both translation directions passed twice. The production language catalog is broader, so we treat those quality tests as strong early evidence rather than a claim that every language pair performs identically.

Benchmark methodology

We replayed BBC World Service audio at real-time speed for 899.540 seconds through Mach's production live translation system, translating English into German. The benchmark measured the time from each completed source phrase to its matching final translated text.

This was a limited internal benchmark: one continuous run, one language direction, and no independent audit or human reference translation for the radio stream. Competitor numbers come from their published reports and use different datasets and endpoints.

  • Start-to-ready: 0.564 seconds.
  • Average phrase-to-translation response: 0.466 seconds.
  • Estimated P50: 0.35 seconds; estimated P95: 1.43 seconds.
  • Final coverage: 174 of 174; audio integrity: 44,977 of 44,977 frames.
  • Next: public long-form datasets, voice-to-visible timing, repeated runs, and independent human quality review.

Primary sources

These sources describe the reported comparison figures and the public evaluation methodology. Vendor figures remain vendor-reported unless an independent benchmark is explicitly identified.