The expected person

A participant entered a one-minute video call expecting to hear what another person looked forward to that year. In Tavus’s Griffin-Lite study, the partner’s face, voice and responses came from a model. After the call, 26 of 54 participants still believed they had met a human. Everyone learned the AI identity at the end. During that short exchange, the decision to enter a personal conversation rested on an incorrect assumption about the interlocutor.[1]

Tavus presents the approximately 48% result as an achievement in human likeness. The setup matters: participants were told they would meet another participant, giving them an expectation of a human alongside the natural-looking face. The study offers no basis for expecting the same rate in an unsolicited customer call. It does show how fragile appearance-based identity judgments can be under these conditions.[1]

What the face communicates

Griffin’s cues extend beyond lip movement. It processes gaze, expression, pauses and speech together, continuing to listen while responding. A reference image becomes a generated scene containing body movement, the chair, shadows and background. Waiting through a pause or nodding during speech supplies signals that the other party is listening. Identity judgment therefore involves the interaction as a whole, rather than spotting one visual defect.[1]

My reading is that more natural interaction calls for earlier identity disclosure. A person describing themselves on a video call should know whether the partner is a human or a model acting for a service. The experiment measures a setting where that information was withheld. Some human-identification responses may reflect the short conversation’s topic and the expectation established beforehand; that alternative still leaves disclosure timing central to real use.[1]

At the opening of the call

Tavus acknowledges the risk. Griffin-Lite is a research preview for selected trusted testers, rather than a product available to customers, and the company says it is developing disclosure features. That restriction prevents treating the study’s misleading introduction as an established customer practice. Between the preview and a customer conversation sits a consequential design choice: does disclosure arrive before the first personal answer, or after the exchange?[1]

The concrete signal to watch in a broader release is the opening identity notice and the user’s opportunity to choose whether to continue afterward. A human-looking face does not establish that the caller understands the service receiving their words. Among Griffin’s expressive gestures, the timing of its own identity disclosure may be the most consequential design element. People should be able to decide whether to share personal information after receiving that notice.[1]