Three seconds of someone's audio is now enough to produce roughly an 85 percent voice match (Source: Fortune, December 2025). More material gets you further, call it 30 to 60 seconds for something convincing, but three seconds is the floor people plan around.
That is the part of voice cloning that takes no time at all. The part that takes time is everything after it: whether the person agreed to this specific use, whether they get paid when the voice is reused, and whether you can prove both when someone asks. Those are the questions this post is about, and the short answer is that the exposure sits with whoever publishes, not with whoever ran the model.
[Average read time: 7 minutes]
What we concede to synthetic voice, without apology
The case for being careful is not a case against the technology, so give it its due first.
Synthetic voice earns its keep on volume and speed: IVR phone trees, accessibility read-outs for long documents, internal e-learning drafts nobody films, background and secondary lines that fill a scene, and scratch tracks that exist only to be replaced by a real recording later. Refusing to use it in those spots would not be judgment, it would be ideology.
The quality has moved too. Models from several vendors now carry credible emotion and intonation across a wide range of uses, and independent benchmark panels trade the top spot among them month to month (Source: MarkTechPost, May 2026). Human recordings still tend to win in careful attribute-match tests, so "indistinguishable from a person" is not a blanket truth. But the gap that defined the 2019 conversation has mostly closed.
Where it stops is narrower and more specific than the general anxiety suggests. It stops at brand voice, fiction, on-air talent, and anything where a real person's vocal identity is the asset you are actually buying.
The position: a voice is not a consumable
An actor's voice is part of who they are and how they make a living. It is not a free input sitting around to be scraped and recombined. That is a moral claim as much as a legal one, and we say it that way on purpose.
For a buyer it is also a reputation question. "We cloned a voice" is a sentence you do not want attached to your brand unless you can immediately show consent and payment behind it. If you cannot produce that paper trail, the exposure is yours, not your vendor's.
The rule that follows is short: no scraped or stolen voices, ever. A synthetic replica only comes from talent who agreed to that specific use, in writing, knowing what they agreed to.
What 2023 fixed, and what it did not
We would like to claim we invented this standard. The truth is the profession landed here, and the marker was 2023.
The Writers Guild walked out on 2 May 2023 and settled in late September; the tentative deal came on 24 September and members ratified it on 9 October 2023 (Source: Wikipedia, 2023 Writers Guild of America strike). SAG-AFTRA struck on 14 July 2023; that strike ran 118 days before ending on 9 November, and members ratified the new TV and theatrical contract on 5 December 2023 (Source: Wikipedia, 2023 SAG-AFTRA strike). The strike ending and the contract being ratified were not the same week, and the gap is worth keeping straight when someone quotes a date at you.
Two agreements matter for voice specifically. The 2023 TV/Theatrical agreement defines a digital replica as a replica of a performer's voice and/or likeness, so voice is explicitly in scope, and consent there has to be clear and conspicuous rather than buried in standard terms (Source: SAG-AFTRA, Digital Replicas 101). Because game and interactive work is overwhelmingly voice performance, the 2025 Interactive Media Agreement, ratified on 9 July 2025, is the one to read for synthetic-voice reuse: it requires separate, written, reasonably specific consent before a replica is created or used, and that consent is void if the use exceeds what was described (Source: SAG-AFTRA, 9 July 2025).
Here is the translation for a brand buyer. Those protections bind signatory employers and union members. Most corporate, e-learning, and industrial work is neither. The norm moved anyway, and buyers inherit the exposure whether or not a contract forces it, which is why the terms have to be written into non-union contracts by hand if you want them.
Statute is filling in around the union agreements, unevenly. California's AB 2602 has required reasonably specific consent for digital replicas of a living performer's voice or likeness since 1 January 2025, and AB 1836 extended protection to deceased performers from 1 January 2026 (Source: California AB 2602 and AB 1836). In the US there is no general federal right of publicity; protection is state by state. The exact reach of a law like AB 1836 will be tested in court before anyone knows what it covers.
Consent is a deliverable, not an assumption
Voice is uncomfortable in one specific way: anyone with hours of public audio is already a viable cloning target. Podcasters, executives, narrators, politicians. The training set already exists.
Public awareness of this has an origin story. In April 2018, BuzzFeed released a video conceived and voiced by Jordan Peele that put words in Barack Obama's mouth and ended with a warning about fake news (Source: BuzzFeed News, April 2018). It took off-the-shelf software and dozens of hours of processing. For a lot of people that clip was the moment the abstract risk became concrete. The harms since are no longer hypothetical: scam calls cloning a relative's voice, fabricated political audio, voice actors hearing themselves in ads they never recorded.
Not all consent is the same, and the distinction is the whole thing. Informed, scoped, time-bound consent means a performer agreed to a specific use, for a specific period, with the right to refuse the next one. A blanket clause that hands a company a voice forever, for any purpose it later invents, is a different arrangement wearing the same word.
So our standard is that consent is scoped, not blanket. Talent agree to specific use cases, markets, and content types, and that boundary is written into the contract rather than assumed. For any voice we use, we can name the talent, show the signed consent for that use class, and document the pay.
The four licensing shapes
The 2019 fear was replacement. The 2026 reality is a market, and the terms are where the negotiation actually happens. One performer's voice can be modulated or cloned to cover many characters, which makes some work more exposed than others: games, animation, e-learning, and dubbing sit closest to the change, because they need a lot of lines, in a lot of variations, often on short notice.
Four deal shapes have settled in. Know which one you are being sold.
- Per-use. The voice is licensed for one defined output, a single ad or one course module. Using it again means a new agreement.
- Per-project. The license covers everything inside one production, a game or a season, regardless of how many lines get generated.
- Term-limited. The voice can be used for a fixed window, say two years, after which the rights lapse unless renewed.
- Approval rights. The performer or rightsholder reviews and signs off on uses before they ship, which keeps the voice out of contexts they would object to.
These are also the mechanism for compensation, which is the question everyone skips to and nobody defines. A per-use or term-limited deal with approval rights keeps a performer in the loop and in the revenue. A blanket buyout does the opposite: it pays once for something that gets monetized for years. Whether the broader market follows the first model or races to the second is the open question of the next couple of years.
Our own arrangement is that compensation recurs. Talent are paid when their voice is reused, through a consenting talent network built over many years of doing this work the slow way. The scale behind it is real: our production archive holds more than 67,000 audio files across 42 languages, recorded over 21 years (Source: JBI archive study, June 2026).
We should be honest about one thing we are still building toward. A standing network of consenting talent with recurring pay on every replica reuse is the direction we are committed to rather than a box we have finished checking. We would rather tell you that plainly than oversell it.
Ask who cleared the voice, and under which agreement
Vendor labels get muddled here, so be precise. Of the well-known platforms, Replica Studios has a formal agreement with SAG-AFTRA. ElevenLabs and Respeecher run their own commercial licensing and payout models, which can be perfectly fair, but they are not union-approved, and nobody should market them that way.
If a vendor tells you a voice is "cleared," the follow-up is: cleared by whom, under which agreement. As synthetic voices start arriving inside vendor deliverables you did not commission, that question becomes part of intake rather than part of legal review.
Disclosure and the limits of watermarking
Two duties sit at the end of this, and they are different from consent.
The first is disclosure. The EU AI Act's transparency obligations for AI-generated content apply from 2 August 2026, and the duty to disclose falls on the deployer who publishes, not only the AI vendor. If you are a US brand publishing into the EU, that duty is yours. Our post on disclosure and provenance works through the detail, including the separate machine-readable marking duty and who it actually binds, which is not the same party.
The second is proving what happened. Watermarking tools such as Google's SynthID and Resemble's PerTh embed signals into synthetic audio that survive ordinary editing. They help. Every vendor describing them is careful to add that they are not a guarantee, because determined re-encoding or stripping can defeat them. Treat a watermark as probabilistic evidence rather than proof. Content-provenance standards like C2PA, which attach a signed history to a media file, are maturing alongside them and push in the same direction: a verifiable record of how a file was made (Source: C2PA specification).
Content Credentials and similar provenance tooling are a real option for audit trails, and we will be specific about the exact method used on a given project rather than wave at a logo. A label your auditors cannot verify is not worth much to them.
None of this substitutes for the thing underneath. Disclosure says the audio is synthetic. It does not say a real person agreed and got paid.
What to check before you sign
For any voice project, three things, in this order: who consented and to what specific use, how the performer is paid when the voice is reused, and whether the output has to be disclosed as synthetic where you are publishing it. Get all three in writing next to the file rather than in a conversation.
Our white paper, Localizing in the Age of Generative AI, has a full section on voice cloning, consent, and the law if you want the detail in one place. When you want to map the voice side of a project, we are glad to walk through it.
