Image & Video Generation

Sub-second live AI dubbing: the MLS / FanCode / NASCAR playbook

Rajat Gautam18 min readUpdated
Share

Key Takeaways

  • Four named deployments anchor the live AI dubbing space: Camb.ai dubbed MLS NEXT Pro in 4 languages (April 2024, developmental league, not Apple TV first-team); FanCode + Camb.ai delivered fully AI-driven Hindi commentary for Caribbean Premier League 2025; NASCAR + Camb.ai ran Spanish-language radio of the Mexico City Cup Race as a single-event pilot in June 2025; and the Trophée des Champions on January 8, 2026 became the first European football match with live AI-dubbed commentary, French to Italian on Ligue 1+.
  • Two tiers buyers must not confuse: sub-500ms conversational (interpreter, two-way conference, voice agent on live broadcast) and 2-10 second one-way commentary (most production broadcast deployments). Vendor numbers measure different points in the pipeline, and Deepdub has published both a 10-15 second end-to-end figure and a ~125ms first-time-to-audio figure for the same product.
  • The architecture has six legs: capture, streaming ASR with partial hypotheses, MT with wait-k policy, low-latency TTS, cross-lingual voice cloning, mix and transport out. Best-case glass-to-glass at 400-1050 ms when fully parallelised; viewer-side reality is typically 2-10 seconds when LL-HLS transport is in the chain.
  • Camb.ai is the only vendor with 12+ verified on-air deployments across pro and developmental sports from 2024 into 2026, and the only one with a tier-1 European football fixture. Deepdub Live is the credible enterprise competitor, with TPN Gold Shield and SOC 2 Type II disclosed.
  • The tier-1 line was crossed in January 2026 by the Trophée des Champions, but no full-season tier-1 commitment and no flagship US major-league deployment (NFL Sunday Night, an NBA national window, an F1 Grand Prix live feed) has happened in public as of August 2026. One marquee fixture is not a season.
  • Procurement reality: 5-10x price step from on-demand to live, broadcast SLAs (99.95+ uptime, latency percentiles, drift ceiling, RTO), TPN+ for studio-adjacent buyers, named-account engineering with 24/7 NOC, pre-event glossary management for proper nouns.
  • Ten failure modes to test in a paid pilot: voice consistency under context pressure, emotional tone preservation, accent and dialect handling, proper-noun accuracy, profanity filtering, multi-speaker diarisation, music + voice mixing, ASR in noisy environments, language ID latency, latency drift over event duration.
Sub-second live AI dubbing: the MLS / FanCode / NASCAR playbook

The live AI dubbing tier is real, it is deployed on air, and almost nobody outside the broadcast-tech press is reading the technical details accurately. Four deployments to anchor on.

MLS NEXT Pro x Camb.ai, April 2, 2024. Orlando City B versus Inter Miami II at IMG Academy. Streamed free on MLSNEXTPro.com in four languages: original English plus AI-generated French, Spanish, and Portuguese (MLS official, Camb.ai post). Note that this is MLS NEXT Pro (the developmental league), not the MLS Season Pass first-team product on Apple TV. The two are different deployments.

FanCode (Dream Sports group) x Camb.ai, September 2025. Caribbean Premier League 2025 Hindi commentary, fully AI-driven via Camb.ai. FanCode broadcasts 17,000+ live matches per year including 6,000+ cricket matches; this is the largest confirmed cricket-tier deployment in 2026 (Dream Sports newsroom, Variety coverage).

NASCAR Mexico City Cup Race, June 2025. Camb.ai supplied live Spanish-language translation of Motor Racing Network radio broadcasts. NASCAR called it a first. Translated feeds were available alongside English audio on the NASCAR app. The deployment is a single-event pilot with stated plans to expand. Sourcing is Camb.ai + trade press; no dedicated nascar.com press release was published.

Trophée des Champions, January 8, 2026. Paris Saint-Germain versus Olympique de Marseille at Jaber Al-Ahmad International Stadium in Kuwait City, carried on Ligue 1+ with live Italian commentary translated from the French feed by Camb.ai (Ligue 1 announcement, Sports Video Group). Both the league and the trade press called it the first European football match with AI-dubbed live commentary. This is the most consequential reference in the set, because it is a senior first-team fixture between two top-flight clubs rather than a developmental league, a press conference or a radio feed.

From these four, plus the Ligue 1 / LFP Media multi-year deal (March 2025), Australian Open 2024-2025, Eurovision Sport, European Athletics, India Today Group, the Milano Cortina 2026 Paralympic Winter Games, and YES Network in development, Camb.ai has cleared the informal 12-events-or-more reference threshold that separates broadcast-grade from demo-grade. No other vendor has matched that public deployment count for sub-second live AI dubbing as of August 2026.

This piece is the engineering playbook behind that tier: the architecture, the latency budget, the vendor map, what procurement actually requires, what's hard at the live tier, and the buyer-side checklist for any team scoping a live AI dubbing build in 2026.

For the on-demand (asynchronous) AI dubbing tier, see our broader AI dubbing and voice cloning piece.

Two tiers buyers must not confuse

The single most common mistake in scoping a live AI dubbing build in 2026 is conflating two different tiers.

Tier 1: sub-second conversational. End-to-end glass-to-glass latency under 500 milliseconds. Required for two-way interpreter use cases, voice agents, simultaneous interpretation in conferences. Architecture is aggressive: streaming ASR with partial-hypothesis triggering, wait-k policy in machine translation, sentence-level streaming into TTS, often SRT or WebRTC transport. Vendors at this tier in 2026: Camb.ai DubStream (sub-500ms vendor claim), OpenAI Realtime / gpt-realtime (~300-500ms), Cartesia Sonic-3.6 (sub-90ms vendor claim, roughly 166-190ms once the network is in the path), ElevenLabs Flash v2.5 with Speechmatics streaming ASR (~200ms vendor claim, 288ms P50 time-to-first-audio in a 2026 third-party benchmark).

Tier 2: 2-10 second one-way commentary. Acceptable for one-way commentary dub of live sports, news, conferences. Architecture is more conservative with bigger buffers, deeper translation context, higher voice-quality output. Most production deployments operate here. Deepdub Live is the sharpest illustration of the measurement problem, because the same vendor now leads with a different number than it launched with. The launch material quoted 10-15 seconds end-to-end (Deepdub announcement); the current live-dubbing page leads with ~125ms FTA, first time to audio (Deepdub Live). Those are two measurement points roughly two orders of magnitude apart and the page does not reconcile them. Neither figure is a correction of the other, and neither belongs in a contract until the vendor states in writing which one is glass-to-glass at the viewer. AI-Media LEXI Voice operates under 10 seconds. SyncWords' Live Dubbing operates in a similar range.

The vendor's quoted number is almost never apples-to-apples with the next vendor's. Camb.ai's sub-500 milliseconds is the conversational tier or the TTS-first-byte measurement. Deepdub's 10-15 seconds is glass-to-glass viewer-side including HLS packaging. Same product category, different measurement points, very different numbers. Buyers should ask each vendor specifically: median and 95th-percentile end-to-end latency measured at the viewer endpoint, with date and event name.

The architecture: three stages plus three transports

Every live AI dubbing pipeline has the same six legs. End-to-end latency equals the sum of these legs minus the parallelism wins the vendor's engineering team can extract.

Leg 1: Capture and ingest. Source audio enters as SDI / AES (broadcast tier) or SRT / RTMP (streaming tier). 50-150 milliseconds depending on transport.

Leg 2: Streaming ASR. Whisper-class but with streaming partial-hypothesis output. Production options in 2026 include NVIDIA Riva (Parakeet TDT streaming, ~30-80ms per chunk; multilingual ASR with Whisper-Large + Canary) and NVIDIA's newer Nemotron Speech Streaming, a roughly 600M-parameter cache-aware streaming conformer transducer built for exactly this job, Faster-Whisper (CTranslate2 backend, ~22ms per chunk on RTX 4090 at INT8 with Large-v3 Turbo), Speechmatics Real-Time on Ursa 2 (partials under 500ms, 55+ languages), AssemblyAI Universal-Streaming, Camb.ai's proprietary BOLI front-end. The win at this leg is partial hypotheses: emit a partial transcript every 100-300ms so downstream legs can start working before the speaker finishes the sentence.

Leg 3: Machine translation. Two regimes. Cascaded with wait-k policy is what most production systems use: MT triggers when ASR emits a stable partial hypothesis of 3-5 words. Latency 100-200ms per chunk. End-to-end speech-translation (S2ST) is the future direction: direct audio to audio without an intermediate text representation. OpenAI's gpt-realtime is the leading example. Currently hard to scale to 100+ languages with the named-entity accuracy that broadcast requires.

Leg 4: Low-latency TTS. This is the largest single chunk of the compute budget, roughly 62 percent of pipeline compute time per one published IEEE measurement. Production options in 2026 with publicly stated latency:

  • Cartesia Sonic-3.6 (state-space architecture, not transformer; shipped August 2026, after Sonic-3 in October 2025 and Sonic-3.5 in May 2026): sub-90ms vendor latency claim, roughly 166-190ms measured once the network is in the path (Cartesia Sonic page)
  • ElevenLabs Flash v2.5: 75ms inference vendor claim, and still their real-time model. Eleven v3 reached general availability in March 2026 but is a larger, higher-fidelity model that they do not recommend for real-time, so keep it out of a live chain
  • Deepdub Lightning 2.5: 200ms inference latency, runs on NVIDIA TensorRT-LLM
  • Camb.ai MARS-Flash (the low-latency member of the MARS8 family, January 2026): sub-150ms TTFB design target, with ~100ms TTFB reported in production at 600M parameters (MARS8 technical report)
  • NVIDIA Riva TTS / Magpie-class: integrated with NVIDIA Holoscan for Media for broadcast workflows

Leg 5: Cross-lingual voice cloning. The source speaker's voice fingerprint must persist across the translated output. Camb.ai's MARS family takes a 5-second reference (MARS5 / MARS6) or as short as 2 seconds (MARS8) and preserves voice identity across language outputs. This is what makes the commentator sound like the same human in Hindi or Spanish that they sounded like in English. Critical for sports because the voice is the identity.

Leg 6: Mix and transport out. The translated voice is mixed against crowd noise, music bumpers, sponsor stings. Then transport: SRT (1-2 second latency), LL-HLS (2-4 seconds), MPEG-DASH, or WebRTC (under 500ms but rare for one-way broadcast). The transport choice often dominates the perceived viewer-side latency.

End-to-end latency budget, best case.

StageBest-case msNotes
Capture + jitter buffer50-100SDI faster than SRT
Streaming ASR (partial)100-300Riva Parakeet, Speechmatics partials
MT (wait-k=3)100-200Triggered on stable partial
TTS first-byte75-150Vendor claims. Measured figures run higher
Mix + transport out100-300SRT lower than LL-HLS
Total glass-to-glass~425-1050 msPipeline fully parallelised, on vendor-claimed legs

Camb.ai's sub-500 millisecond claim implies aggressive parallelisation between ASR partials, MT wait-k, and TTS streaming, plus SRT or WebRTC out rather than LL-HLS. For broadcast TV with LL-HLS to viewer, sub-second total is at the edge of feasible; sub-2-second is more realistic at the viewer endpoint.

Vendor map, August 2026

The sub-second live AI dubbing field has cohered into about ten serious players. Honest evaluation matrix below. ("Sub-second claim" is vendor-published unless flagged. Component-level TTS benchmarks now exist, including the Artificial Analysis speech arenas, but there is still no independent end-to-end benchmark for live dubbing as a pipeline, so no row below is third-party verified.)

VendorLatency claimTPN statusVerified live sportsSRT / HLS native
Camb.ai DubStreamsub-500msNot publicly disclosedTrophée des Champions, MLS NEXT Pro, Ligue 1, NASCAR pilot, FanCode CPL, AO, European Athletics, Eurovision Sport, Milano Cortina 2026 Paralympics (subtitling)Yes (via Ant Media and Videolinq partnerships)
Deepdub Live~125ms FTA published now, 10-15s at launchTPN Gold Shield + SOC 2 Type IINAB 2025 + IBC 2025 demosYes (SRT / HLS / MPEG-DASH)
AI-Media LEXI Voiceunder 10sInherited from AI-Media captioning legacyCaptioning legacy at TV broadcasters, LEXI Voice Encoder launched NAB 2026Yes
ElevenLabs + Speechmatics combo~200ms TTSVia partnersVia AI-Media LEXI integrationVia partners
OpenAI Realtime / gpt-realtime~300-500msNoNone for sportsNo native, REST-based
Cartesia Sonic-3.6sub-90ms TTS claimNoNone publishedNo native
OneMeta via NVIDIA HoloscanHoloscan-pacedNot disclosedAnnounced GTC March 2026, no live sports yetVia Holoscan
PapercupHybrid + human QANot disclosedNot real-timeNo native
SyncWords Live Dubbingunder 10sNot disclosedLive news / eventsYes
Resemble AIAudio-only focusNoNone publishedNo

The two serious choices for a buyer evaluating a sub-second live AI dubbing build in 2026 are Camb.ai (broadest sports references including the first tier-1 European football fixture, sub-second tier claim) and Deepdub Live (TPN Gold Shield and SOC 2 Type II, AWS Elemental MediaPackage native). Choose by deployment shape: if you need true sub-second conversational behavior (interpreter, two-way conference, voice agent on a live broadcast), Camb.ai. If you need broadcast-grade content security because a TPN-mandating studio is in your chain, Deepdub Live, and pin down its glass-to-glass number in the contract rather than the FTA figure on the page.

AI-Media LEXI Voice is the multi-vendor wrapper play (largest selection of translation engines and synthetic voices, runs on ElevenLabs Turbo TTS plus Speechmatics ASR underneath, and added a LEXI Voice Encoder with sound separation and noise removal at NAB 2026) and is the right choice if you already have AI-Media captioning in your pipeline. OneMeta via NVIDIA Holoscan is the embed-into-broadcast-infrastructure play. The rest of the field is adjacent (voice-agent tier, conference-tier, or VOD-tier rather than live-broadcast-tier).

Procurement reality at the live tier

The price step from on-demand AI dubbing to sub-second live AI dubbing is roughly an order of magnitude. On-demand VOD dubbing runs $1-3 per finished minute. Live dubbing typically runs 5-10x more on a per-minute basis because vendors charge for active-active redundancy, named-account engineering, and event-day Network Operations Center presence. Camb.ai DubStream Enterprise pricing is not public; industry pattern is six-figure annual commits for a meaningful live-events package. Treat that as a ballpark, not a quote.

Broadcast SLAs to negotiate.

  • Uptime. 99.95 percent (4.4 hours per year downtime) typical. 99.99 percent (52 minutes per year) for premium tier. Live sports buyers generally require 99.95+ for the event window, less stringent off-event.
  • Latency. Stage-by-stage commitments with end-to-end ceiling. Vendors quote median plus 95th percentile. Demand both.
  • Drift. Maximum audio-video desync over the event window. Lip-sync sensitivity is roughly 80 milliseconds; sports commentary is more tolerant, around 250 milliseconds.
  • Recovery time objective. Time to fail over from primary translation track. Most live-tier vendors run dual-region active-active.

Security and content protection.

For any studio-adjacent buyer (Disney, Netflix, Amazon, Warner Bros. Discovery, Paramount, Sony, Apple TV+), the TPN v5.3.1 Trusted Partner Network audit is required. Blue Shield (self-attested) and Gold Shield (third-party audited) tiers, with a 24-month assessment cadence. Deepdub Live publicly discloses TPN Gold Shield certification. Camb.ai still does not publicly disclose TPN status as of August 2026; ask directly during vendor evaluation. SOC 2 Type 2 and ISO 27001 are table stakes for any enterprise vendor.

Named-account engineering support.

  • Dedicated solutions engineer plus 24/7 NOC for live events.
  • Pre-event language tuning: sponsor names, athlete names, jargon glossaries. This is the single highest-impact support task in any live event.
  • Glossary management for proper nouns during the event window. Should be in every Statement of Work.

Voice library and rights.

  • Library voices versus clone-the-commentator. Different legal posture. Clone-the-commentator requires explicit consent, term, buyout structure, and sunset clauses. Post-SAG-AFTRA 2023 strike, commercial-grade voice cloning of real talent has contractual requirements that flow through to any vendor doing the cloning.
  • Output rights. Who owns the dubbed audio? Most vendors grant perpetual license to the buyer for the dubbed output; some retain training rights on the source. This is often the most-negotiated clause.

What's hard at the live tier

Ten failure modes broadcast buyers should test in a paid pilot before signing.

1. Voice consistency under context pressure. When the commentator's pace doubles (a goal, a crash, a wicket), TTS prosody can flatten. Camb.ai has explicitly benchmarked MARS5 and MARS6 against "prosodically hard and diverse scenarios like sports commentary," but the failure mode is still real.

2. Emotional tone preservation across languages. Cross-lingual prosody transfer (XLPT) is the term. Excitement in English commentary must manifest as appropriate excitement in Spanish or Hindi, with culturally-aware delivery rather than direct mapping. Deepdub's eTTS and Camb.ai's BOLI both claim emotion preservation; empirical third-party comparison is rare.

3. Accent and dialect handling. The ASR side: a Scottish commentator on a Ligue 1 broadcast translated to French is a hard input. The TTS side: Hindi has regional variants (Mumbai, Lucknow, Punjabi-tinged). Spanish has at least Castilian, Mexican, Argentine, Caribbean. Camb.ai's Hyperfusion deal in the UAE explicitly broke out Gulf, Levantine, Egyptian, and Modern Standard Arabic for this reason.

4. Proper-noun accuracy. Athlete names, sponsor names, place names. The most-quoted live AI failure mode. ASR mis-hears the name; MT mis-translates it; TTS mispronounces it. The fix is a pre-loaded glossary per event. Buyers should ask the vendor exactly how the glossary is managed and updated mid-event.

5. Profanity / banned-word filtering on a live feed. Live sports commentators occasionally swear under stress. The translated output must censor (or pass through) per the buyer's policy. Camb.ai and Deepdub do not publicly document their profanity policy in detail. Ask, get it in the SOW, test in pilot.

6. Multi-speaker disambiguation. Sports commentary is rarely one voice. Play-by-play, color, pitch reporters, sideline interviews all overlap. Speaker diarization must run in real time. Camb.ai claims automatic speaker diarization and multi-speaker emotion transfer. Test it specifically in pilot.

7. Music and voice mixing. Crowd noise, broadcast bumpers, sponsor stings, in-stadium music all pass through the translated audio mix. Most vendors decouple translation from mix and let the broadcaster's audio engineer ride the mix. Confirm this in your SOW so you do not end up with a vendor automix that fights the in-house mix.

8. Whisper-class ASR struggles in noisy environments. Crowd noise at 95+ dB SPL can defeat ASR. Vendors get around this by using a clean commentary feed (the press-box microphone) rather than the broadcast mix. Confirm the input plane in any SOW.

9. Language identification (LID) latency. When a commentator switches mid-sentence between English and Spanish (common in MLS, Liga MX, NASCAR Mexico), LID needs to run in parallel to ASR, not sequentially. A serial LID adds 200 milliseconds you cannot afford at the sub-second tier.

10. Latency drift over event duration. The ASR + MT + TTS pipeline can drift gradually over a 90-minute match as buffer queues fill. Best vendors run a periodic synchronisation heartbeat. Test for this across a full event window, not a 5-minute demo.

Broadcast-grade versus demo-grade: the smell test

The line between vendor reel and on-air production is observable. Markers of a vendor demo (do not trust as broadcast-grade): single clip on a vendor's home page, press-conference-only demo (Australian Open press conferences were initially demo-grade), no SRT / HLS integration disclosed, no multi-event reference customer, TPN status not disclosed, latency claimed but not measured at viewer endpoint. Markers of a broadcast-grade deployment: named partner with their own press release on their own URL, multiple events under the same partner, disclosed transport integration, TPN+ or equivalent security disclosure, named broadcaster engineering contact, stated latency budget and SLA.

The informal 12-events-or-more reference threshold is a fair smell test, not an industry standard. Camb.ai exceeds it: the Trophée des Champions, MLS NEXT Pro across multiple matches, Australian Open across two years, Ligue 1 multi-year, Eurovision Sport multi-event through to the Milano Cortina 2026 Paralympics, FanCode CPL plus broader cricket scope, European Athletics 2025 season, India Today Group at IBC 2025, plus the NASCAR pilot. Deepdub Live as of August 2026 has the product launch, IBC 2025 demos, TPN Gold Shield and SOC 2 Type II, but a verifiable list of 12+ named on-air live deployments is still harder to surface.

Sober point worth flagging for premium-buyer customers, and it needs stating more precisely than the category usually manages. The tier-1 line was crossed in January 2026 by the Trophée des Champions, a senior first-team fixture between the two biggest clubs in French football, dubbed live into Italian. What has not happened in public as of August 2026 is a full-season tier-1 commitment or a flagship US major-league broadcast: no NFL Sunday Night, no NBA national window, no Formula 1 Grand Prix live feed. The Trophée des Champions was one match, one target language, on a domestic OTT service rather than the primary broadcast feed. YES Network is testing through 2026 and has not gone on air.

So the buyer's actual question has a different answer than it had six months ago. It is no longer whether the tier has ever been trusted with a first-team match. It is whether anyone will commit to it for a whole season, and so far nobody has.

Buyer-side checklist for any live AI dubbing build

A pragmatic list grouped by category. Use this in any RFI or vendor evaluation.

Reference and proof of deployment. Name 10+ on-air production deployments in the last 12 months. For each, link to the partner's own press release on the partner's own URL (not the vendor's blog). Latency measured at the viewer endpoint, not at the lab endpoint, with date and event name. Specific failure incidents in the last 12 months and root cause.

SLA. Uptime commitment for event window versus off-event. Latency commitment with median, 95th, and 99th percentile. Audio-video drift commitment with ceiling and measurement method. Recovery Time Objective for primary track failure. Credits for SLA breach.

Deployment. Cloud, VPC, or on-prem options. Data residency in sovereign regions if needed (UAE, EU, India). Network ingress (SRT, RTMP, MPEG-TS, SDI / AES via gateway). Network egress (SRT, HLS, LL-HLS, MPEG-DASH, RIST). Multi-region active-active or active-passive default.

Security. TPN+ Blue Shield plus Gold Shield status, with date of last assessment. SOC 2 Type 2 and ISO 27001. KMS-backed key management. IP allowlists, private interconnect (Direct Connect, Cloud Interconnect). Audit logs accessible to the buyer. Content training opt-out (vendor does not train on buyer audio).

Voice library and rights. Library voice terms of use, exclusivity, removal options. Clone-the-commentator consent process, term, buyout structure, sunset. Output ownership in perpetuity for the buyer with no vendor reuse for training.

Language coverage and quality. Languages supported with quality tier (top-tier versus long-tail). Dialect coverage for target audience. Pre-event glossary support. Mid-event glossary update path.

Live tier specifics. Multi-speaker diarisation in real time. Profanity / banned-word policy (block, bleep, rephrase, pass-through). Sponsor-name handling (forced pronunciation list). Music duck and un-duck behaviour. Number / score / time-of-day handling (often a problem for MT).

Engineering. Named solutions engineer. 24/7 NOC during event window. Pre-event load test against simulated input. Direct backchannel to TTS / ASR / MT teams for incident escalation.

Pricing. Per-hour live event rate versus per-finished-minute VOD rate. Active-active redundancy cost. Multi-language add-on pricing. Annual commit versus event-based pricing.

Other deployments worth knowing about in 2026

Beyond MLS NEXT Pro, FanCode, and NASCAR, the live AI dubbing space added a meaningful set of references in 2024-2025:

  • Ligue 1 (LFP Media, France). Multi-year exclusive partnership with Camb.ai announced March 2025, 150 languages of live broadcast translation. It produced the Trophée des Champions Italian dub on Ligue 1+ on January 8, 2026, the reference deployment above.
  • Australian Open / Tennis Australia. Camb.ai in the AO StartUps 2024 cohort. Real-time multilingual dubbing of player press conferences and interviews in 2024 (Djokovic, Gauff, Medvedev) and 2025.
  • European Athletics. First major athletics association with a fully multilingual digital experience. 12 languages across event coverage, athlete profiles, competition updates, post-2025 SPAR European Cross Country Championships.
  • Eurovision Sport (EBU). Europe's first AI-powered real-time translated sports commentary at the 2024 World Athletics U20 Championships in Lima. French-to-Portuguese live dubbing.
  • Milano Cortina 2026 Paralympic Winter Games. Eurovision Sport and Camb.ai delivered live and on-demand subtitling across all six sports for the Games of March 6 to 15, 2026 (EBU). Real-time speech-to-text, so the subtitling tier rather than full dubbing, but it is a delivered multi-day event rather than an announcement.
  • India Today Group. First global news-media partnership for real-time multilingual translation. Hindi plus regional Indian languages plus diaspora-targeted English. Demonstrated at IBC 2025.
  • YES Network (Yankees / Nets regional sports network). Strategic agreement announced November 2025, testing in 2026. Not yet on-air.
  • Hyperfusion (MENA sovereign cloud). UAE-hosted sovereign MARS7 deployment for newsroom dubbing, presser dubbing, live broadcasting, trailer and VOD production. Arabic dialect packs.
  • Broadcom and Camb.ai on-device. November 2025 partnership to run translation directly on the neural processing unit inside Broadcom set-top-box silicon, a fleet of over 500 million devices. Text-to-speech on the NPU is now demonstrable; porting the real-time translation model to the same NPU is the stated next phase, not shipped. Broadcom's CES 2026 Wi-Fi 8 platform with the BCM4918 APU is the hardware this lands on.
  • IMAX. Camb.ai AI dubbing for IMAX original content globally, which Camb.ai now describes as real-time dubbing for immersive cinema rather than post-production only. Cinema rather than broadcast, so it is a credibility signal rather than a live-broadcast reference.
  • Comcast NBCUniversal. Camb.ai came through the Comcast NBCUniversal SportsTech programme, which is also how the NASCAR work reached air. A distribution relationship with a tier-1 US broadcast group, not a broadcast deployment of its own.
  • Dentsu and Nippon Cultural Broadcasting (Japan). Partnership concluded June 2025; from July 3, 2025 a Nippon Cultural Broadcasting radio programme streamed in the original Japanese alongside AI-translated English and Chinese (Dentsu). Radio tier, and the second radio reference in the set after NASCAR.
  • Videolinq. May 2025 partnership pairing Camb.ai dubbing with Videolinq live captioning and subtitle distribution, including delivery to audience mobile devices at live events. Relevant if your build needs captions and dubbing from one contract.

No verified tier-1 Western news broadcaster (BBC, Reuters, AP) live AI dubbing deployment as of August 2026. The BBC in particular is running a deliberately conservative AI programme aimed at summarisation and reformatting rather than audience-facing voice. India Today Group remains the closest analog. No verified live AI-dubbed earnings call from a Fortune 500 issuer, and no verified live AI dubbing at a major concert or awards show, as of the same date.

Keep reading

For the broader 2026 AI video stack including the on-demand AI dubbing tier, see AI dubbing and voice cloning and AI video generation 2026. For the provenance layer that studio-adjacent buyers increasingly ask about, see what breaking SynthID means for AI media proof, and for how the studios themselves are actually using this, see Hollywood is quietly using AI more than they say. When you are ready to scope a sub-second live AI dubbing build, a broadcast-grade procurement, or a vendor evaluation programme, our AI Content Production service covers exactly that, or book a strategy call to walk one specific scoping question live.

Frequently Asked Questions

Is the MLS x Camb.ai deployment on Apple TV's MLS Season Pass?+
No. The verified deployment is MLS NEXT Pro (the developmental league), not the MLS Season Pass first-team product on Apple TV. The first AI-driven multi-language broadcast was Orlando City B vs Inter Miami II on April 2, 2024 at IMG Academy, streamed free on MLSNEXTPro.com in four languages (English original plus AI-generated French, Spanish, Portuguese). Camb.ai continued the deployment through 2024-2025 MLS NEXT Pro coverage and contributed to MLS All-Star Game ancillary broadcasts. There is no public evidence Camb.ai live dubbing has run on the MLS Apple TV Season Pass product.
Has any tier-1 sports property deployed sub-second live AI dubbing on a flagship broadcast?+
Partly, and the answer changed in January 2026. The Trophée des Champions on January 8, 2026, Paris Saint-Germain against Olympique de Marseille, was dubbed live from French into Italian on Ligue 1+ and is the first European football match with AI-translated commentary. That is a senior first-team fixture, not a developmental league. What has still not happened as of August 2026 is a full-season tier-1 commitment or a flagship US major-league broadcast: no NFL, NBA, MLB or NHL national window. The other verified deployments sit below that line: MLS NEXT Pro (developmental MLS), NASCAR (single-event pilot via radio commentary), Australian Open press conferences. YES Network (Yankees / Nets regional sports network) announced a strategic agreement with Camb.ai in November 2025 with testing through 2026 and language selection planned inside the Gotham Sports App, but it is not on air. A full-season tier-1 deployment is now the most-watched gap in the space.
What latency should I expect at the viewer endpoint for a sub-second live AI dubbing build?+
Depends on transport. With SRT or WebRTC out, you can hit 800ms to 2 seconds glass-to-glass with a fully parallelised pipeline. With LL-HLS to viewer (the typical broadcast streaming path), you add 2 to 4 seconds of packaging and delivery latency on top of the AI pipeline, putting viewer-side at 3 to 6 seconds in best case. Camb.ai's sub-500 millisecond claim is the conversational tier or the TTS-first-byte measurement, not necessarily viewer-side end-to-end on a typical LL-HLS broadcast. Always ask the vendor to specify the measurement point.
How does Camb.ai's MARS8 family of TTS models actually work?+
MARS8 launched January 2026 as a family of four sizes: MARS-Flash (sub-150 ms TTS first-byte, optimised for real-time streaming), MARS-Pro (the workhorse for production-grade voice cloning), MARS-Instruct (director-level control for film and high-stakes media), and MARS-Nano (50M-parameter on-device model). MARS5 was released open-source in June 2024 (5-second voice cloning, two-stage AR-NAR pipeline). MARS6-turbo is a 70M-parameter encoder-decoder transformer. The MARS family handles cross-lingual voice cloning by fingerprinting the source speaker from a 2- to 5-second reference and preserving voice identity across 150+ language outputs. BOLI is Camb.ai's proprietary speech-translation engine that pairs with MARS for the speech-to-speech pipeline.
What's the price step from on-demand AI dubbing to sub-second live AI dubbing?+
Industry pattern is 5 to 10 times per minute. On-demand (asynchronous) AI dubbing runs $1 to $3 per finished minute. Live AI dubbing is typically priced higher because vendors charge for active-active redundancy across two regions, named-account engineering, 24/7 NOC during event windows, and event-day language tuning (sponsor names, athlete names, jargon glossaries). Camb.ai DubStream Enterprise pricing is not public; industry pattern is six-figure annual commits for a meaningful live-events package. Treat as ballpark, not a quote, and ask each vendor for a per-event-hour proposal against your actual event schedule.
What's the difference between Camb.ai DubStream and Deepdub Live for a broadcast buyer?+
Different positioning. Camb.ai DubStream is sub-500-millisecond claimed, with the broadest verified live sports references (12+ deployments including the Trophée des Champions, MLS NEXT Pro, Ligue 1, FanCode, the NASCAR pilot, Eurovision Sport through the Milano Cortina 2026 Paralympics, European Athletics and AO press conferences) and no public TPN status disclosure. Deepdub Live is TPN Gold Shield and SOC 2 Type II certified, AWS Elemental MediaPackage native, and SRT / HLS / MPEG-DASH compatible. Deepdub's published latency needs care: it launched on a 10-15 second end-to-end figure and now leads with ~125ms first time to audio, which measures a different point in the pipeline. For a studio-adjacent buyer with strict TPN requirements, Deepdub Live is the safer procurement. For breadth of on-air sports proof, Camb.ai. Either way, get the glass-to-glass number at the viewer endpoint in writing before you sign.
Will live AI dubbing replace human interpreters at conferences and events?+
Partially in 2026, more so over the next two to three years. The current pattern is meeting and conference tier (Zoom, Teams, Webex) is increasingly AI-dubbed via Wordly, Interprefy, JotMe, Kudo, Interactio, plus Microsoft Teams Interpreter agent (public preview, 9 languages, requires Microsoft 365 Copilot license). Conference plenary sessions and Davos / TED main rooms still use human interpreters in simultaneous-interpretation booths. AI dubbing at conference scale is real but not yet flagship. For business events, town halls, earnings calls, and most professional meetings, AI is now the default or co-default with human review on critical content.

Scoping a sub-second live AI dubbing build, a broadcast-grade vendor evaluation, or a multi-language sports / news production pipeline? We build the engineering layer end to end.

Explore AI Content Production

About the Author

Rajat Gautam

Rajat Gautam

AI Consultant & Founder

My work goes far beyond recommending tools - I design AI systems that integrate directly into your workflows, eliminate inefficiencies, and deliver measurable business impact. Every solution I build is tailored, practical, and built with long-term scalability in mind.

Need help with this?

Related Topics

AI Dubbing
Live Broadcast
Camb.ai
Deepdub
Sub-Second Latency
MLS
FanCode
NASCAR

Related Articles

Ready to transform your business with AI? Let's talk strategy.

Book a Free Strategy Call