Tavus claims Griffin AI model convinced 48% of users it was human in live video calls
Quick Look
- Tavus unveiled its Griffin AI model, which convinced 48% of participants in live video calls they were speaking with a real human, a significant improvement from its previous system's 2.4% success rate.
- The company says Griffin-Lite, tested on 54 people with 26 believing it was human, leads NVIDIA's VideoFDB benchmark in generation track but trails humans in perception.
- Tavus emphasizes the model requires safety measures before public release and is currently available only to trusted testers.
AI-generated summary
Why It Matters
Tavus previously had an AI system that scored only 2.4% in convincing people it was human during live video calls. The company has now developed Griffin, a Human Interaction Model designed to understand and generate face-to-face conversation by analyzing expressions, pauses, and words.
AI startup Tavus says its new model, Griffin, convinced 48% of the people who talked to it on a live video call that they were speaking with a real human.
The company unveiled it on October 1 and calls it the first Human Interaction Model, an AI built to understand and generate face-to-face conversation, paying attention to expressions and pauses as well as words. Its previous system scored 2.4% on the same test.
Griffin-Lite, the version Tavus tested, faced 54 people, and 26 of them said afterward they believed their partner was a real person. The older system faced 41 people and got exactly one believer.
Tavus explains that participants were told they would be matched with another person for a one-minute video call about what they were looking forward to this year. Only at the end were they asked whether it had crossed their mind that their partner might not be real.
That's not the classic Turing test, proposed in 1950 by British mathematician Alan Turing, where a judge talks to a hidden human and a hidden machine and has to work out which is which. Nobody on the call was told a bot might be on the other end.
The results come from Tavus' own research page, and participants were recruited through what it calls an independent research platform. A community note on X has already flagged that the results are not independently verified and do not follow a standard protocol. Tavus says the people who grew suspicious usually did so within 20 seconds.
AI companies building towards this has been a thing for a while. A UC San Diego study found OpenAI's GPT-4.5 convinced judges it was human in 73% of conversations, when prompted to play an introverted, internet-savvy young person. That test was text only.
Griffin adds a face and a voice, in real time.
On NVIDIA's VideoFDB benchmark, a test of live audio and video conversation, Tavus says Griffin-Lite ranks first. Its generation track, which grades how natural and expressive a model's responses are, gave Griffin-Lite 3.83. The next-best system got 2.80, and the human reference scored 3.92.
The perception track, which checks whether a model understands what it sees and hears, shows where it still trails people. Griffin-Lite scored 3.73 against 3.44 for the strongest baseline, while the human reference hit 4.20. Tavus says NVIDIA ran the evaluation independently.
Griffin is also full-duplex, meaning it listens, watches and talks at the same time, like a phone call instead of a walkie-talkie. In a Tavus demo, it coaches a man through a Rubik's cube based on what it sees in his hands, and waits when he goes quiet to think. Audio-to-video delay averages 0.43 seconds on NVIDIA H100 chips, the kind used in AI data centers, which Tavus says is half that of the next fastest method.
Why should anyone outside tech care? Well, for one thing, because scammers already work on video calls.
Back in January, North Korea-linked hackers used deepfakes, AI-made video that imitates a real person, on Zoom or Teams calls to pose as trusted contacts. Security researchers attribute the intrusion to BlueNoroff, a Lazarus Group subsidiary. Victims get talked into installing malware disguised as an audio fix.
David Liberman, co-creator of Gonka, a decentralized network for AI computing, said in that report that photos and video can no longer be trusted as proof that something is real. Those models were not as advanced as this one.
Companies already improvise their defenses. In 2025, Kraken flagged a suspected North Korean job applicant after its security team asked spontaneous questions, like requesting government ID and the names of local restaurants. The candidate struggled.
We tried Tavus’ models and the results were… disappointing. After some research, it turns out Griffin-Lite is not available to customers, only to select trusted testers as a research preview, and the company says it is working on disclosure features and with AI safety organizations.
This is because Tavus says Griffin needs safety measures before a public release.
Tavus raised a $40 million Series B in November 2025, led by CRV. The system that scored 2.4% stitched together three separate models, one each for visuals, dialogue and perception.
Trusted testers can request access to Griffin-Lite by submitting a form on the Tavus site.
What to Watch
AI outlook — possibilities, not facts
Tavus will release Griffin to the public after implementing safety features
Likely · Within months
Griffin or similar models will be used in social engineering attacks
Possible · Within months
Open Questions
- When will Griffin be available to the general public?
- What specific safety measures is Tavus implementing before public release?
- How independent was the research platform that recruited participants for the Griffin-Lite test?







