The sample gets shorter
ElevenLabs says v4 and v4 Turbo generate speech faster, support more than 90 languages and can clone a voice from a 10-second sample. That last claim comes from the company; the report contains no independent test. Still, the product direction is clear: the input needed for cloning is shorter, and a voice agent can begin speaking before the language model behind it finishes its answer.[1]
A short sample cannot, by itself, tell the system whether the speaker authorized its use. The same 10 seconds could come from a recording supplied knowingly or from a video shared for another purpose. I am not alleging the latter happened here; the announcement leaves that distinction unmeasured. The technical ability to reproduce a voice and permission to use it for a particular purpose require separate checks.[1]
The person absent from the call
The company also aims to vary delivery with textual context and stack expression tags. That can make a listener’s simple comparison with one fixed voice sample less useful; better expression control is not evidence of deception on its own. The caller’s awareness is also distinct from the voice owner’s consent. If an agent imitates a human voice, whether the caller knows it and whether the speaker authorized that use are two different questions.[1]
The report on the new release gives concrete details about languages, latency and expression controls. It does not describe a method for checking the speaker’s permission, a document defining the permitted scope of use, or a way to contest a clone. Silence in this report does not establish that the company lacks such measures. The useful document would spell out whose approval is needed for a voice made from 10 seconds of audio, what uses that approval covers, and how it can be withdrawn.[1]