Tavus, a San Francisco startup built around what it calls human computing, has announced Griffin - which it describes as the first Human Interaction Model: a single full-duplex, video-to-video system that holds face-to-face conversations in real time. Unlike a voice assistant with a face bolted on, Griffin continuously sees, hears, interprets, speaks and reacts - adjusting its words, timing, tone and expressions to what is happening in the moment, whether that is being interrupted, knowing when to laugh, or waving at someone who just walked into the room.
The number that will travel is 48%. In a live study the company ran, 26 of 54 participants in the US and Europe believed they were speaking with a real person after a one-minute video call - participants were told they would talk with another participant about what they were looking forward to that year, and were only asked at the end whether their partner was real. For comparison, Tavus says just 1 of 41 participants (2.4%) mistook its previous Phoenix-4.5 stack for a human under the same protocol. The study was small and company-run, so the honest read is a strong signal, not a settled benchmark.
The third-party numbers carry more weight. NVIDIA scored Griffin-Lite in September on its Video Full-Duplex Benchmark, ranking it first on both tracks: a generation score of 3.83 out of 5 against a 3.92 human reference and 2.80 for the next-highest published system, and a perception score of 3.73 against 3.44 for the strongest reported baseline and 4.20 for humans. Tavus also reports the video component responds to incoming audio in an average of 0.43 seconds on H100 GPUs, roughly half the latency of the next-fastest published system it compared against - a company-reported figure. Both benchmark scores come from a language-model judge, which measures something narrower than a human's intuition.
Architecturally, Griffin replaces Tavus's previous three-model pipeline - Phoenix-4.5 for rendering, Raven-1 for multimodal perception and Sparrow-2 for conversational timing - with one model designed around the interaction itself, in which perception, conversational decision-making, speech and video generation all run continuously. Company demonstrations show it playing Simon Says, coaching someone through a Rubik's Cube and reacting to an object held up to the camera, and it generates full scenes rather than only animating a face. Tavus's research team says its work has been cited more than 90,000 times, and more than 150,000 developers and businesses use one or more of its component technologies.
What may matter most is what Tavus is not doing: shipping it. Griffin-Lite is available only as a research preview to selected trusted testers - not to customers - and the company says further alignment work is required precisely because a model this convincing can deceive. It is developing disclosure features and safety procedures before wider release. That caution lands in a world where deepfake video-call fraud is already an active criminal pattern, and where the same realism that powers mock job interviews and sales practice powers impersonation.
The company behind it is real and funded: Hassaan Raza co-founded Tavus with Quinn Favret in Y Combinator's Summer 2021 batch, and the company raised a $40 million Series B in November 2025 led by CRV with participation from Scale Venture Partners, Sequoia Capital, Y Combinator, HubSpot Ventures and Flex Capital. If Griffin's numbers hold up outside the company's own protocol, the real-time video Turing test has effectively been passed - and the first mover chose to keep the door mostly closed.
Comments (0)
Log in to join the discussion
Log InNo comments yet