A face and a voice enter the studio
TechCrunch reporter Dominic-Madori Davis consented to the making of her digital likeness at Synthesia’s New York office. The team took many photographs and recorded 2 minutes of her voice. One version reads a script she provides; the interactive version answers questions about her article on venture fraud. When her parents asked things only the family would know, it redirected them to that article. The limit was visible to the people testing it.[1]
For the interactive likeness, speech becomes text, a language model processes the question, another model voices the answer, and Synthesia’s video model animates the face. Those four stages explain why a familiar face can sit over a tightly bounded conversation. Resemblance to Davis gives the system no access to everything Davis knows.[1]
The data path beyond the conversation
Consent is explicit in Davis’s case: she entered the studio and tried the likeness herself. Customers can choose voice and language components from Cartesia, ElevenLabs, Google, or OpenAI, and host an avatar in their own cloud or with Synthesia. Those choices change the path data takes in each installation. The one-article limit describes the conversation; this demonstration alone leaves the storage of photographs, the voice sample, and incoming questions unspecified.[1]
I would judge the conversational limit and the data path separately. Keeping the avatar on one subject can prevent an unexpected answer in the reporter’s voice; it does not set retention or access terms. Another installation could keep more components with one provider and reduce transfers. Davis’s test shows the subject limit working. Comparable clarity about the data path depends on the deploying organization disclosing its chosen models, hosting, and retention terms.[1]