← News

Tavus says Griffin is the first Human Interaction Model to pass a video Turing test

Tavus on Oct. 1 unveiled Griffin, a full-duplex video-to-video Human Interaction Model that saw, heard, spoke, and moved in real time; in the company's live study 26 of 54 participants (48%) believed their partner was human, and Griffin-Lite stays a limited research preview while safety disclosure work continues.

Face-to-face AI that half of strangers mistake for a person in a minute is not a cute demo — it is the moment video agents become socially plausible enough for scams, care, and sales at once. Tavus holding Griffin-Lite to a research preview is the only honest product note that matters until disclosure is real.

On Thursday, 1 October 2026, Tavus introduced Griffin. The research post is titled “Griffin: The First Human Interaction Model.” The dateline is San Francisco, California, October 1st, 2026. The byline is Hassaan Raza, co-founder and chief executive, Ioannis Patras, head of research, and the Tavus research team. The research page does not print an hour. A Business Wire release the same day, as carried on Yahoo Finance, stamps 11:30 a.m. Central, which is 12:30 p.m. Eastern. Griffin, Tavus says, is its first Human Interaction Model, shortened to HIM. An HIM, in the post’s words, is a new class of model built to understand and generate face-to-face conversation in real time. It listens while the other person talks, and it pays attention to expressions and pauses, not only to words. Full-duplex means both sides can be active at once: the model keeps listening and watching while it speaks. Video-to-video means it takes a live picture and sound and answers with a live picture and sound. Those lines are Tavus’s.

What Tavus says most real-time AI does instead. Most systems, the post says, work as a relay, often called a cascade. One piece writes down the speech, a language model answers that text, and other pieces turn the answer into a voice and a face. Every handoff adds delay, and it throws away things the next piece never sees, such as tone of voice or what is on camera. Griffin puts perception, the choice of when and how to respond, and speech and video in one system, running together. At short intervals, well under a second, it decides whether to speak, make a small sound of agreement, nod, or wait, and it can do that while the person is still talking. The timing diagram on the page is marked illustrative. Tavus says those ticks were not measured from a real session. Those lines are Tavus’s.

The pass Tavus claims. The post says Griffin is the first model to pass a real-time video Turing test. As this page uses that phrase, it means a live video call in which people were not told the partner was a model, and afterward many of them said the partner was a person. In the company’s live study, 26 of 54 participants, 48 percent, believed a partner running on Griffin-Lite was a real person after one minute. Griffin-Lite is the research-preview version. On the earlier stack, Phoenix-4.5 with Sparrow-2 and Raven-1, Tavus says 1 of 41 participants, 2.4 percent, believed the partner was a person, under a similar protocol. The opening of the post rounds earlier systems, including that stack, to a maximum of about 2 percent. The study section prints 2.4 percent and 1 of 41. Both figures are Tavus’s. The post says this is, to the best of its knowledge, the first model to have passed a video Turing test. That sentence is Tavus’s claim about its own calls. The page does not point to an outside lab that repeated those one-minute conversations.

How the calls were set up. Participants came through an independent research platform. They were told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. The partner was a PAL running on Griffin-Lite, making her face, voice, and replies in real time. A PAL, as the company’s wire puts it, is the human-looking character a person talks to. After the call, people wrote down the partner’s answer and rated naturalness and trust. Only at the end of the survey were they asked whether it had crossed their mind that the partner might not be a real person, and, if so, when. Everyone was then told the partner had been an AI model. Tavus says the Phoenix-4.5 run used the same protocol. Those lines are on the research post.

How sure people were, and the ratings. Of the 54 people who spoke with Griffin-Lite, 26 said the partner was real. Tavus says both groups were confident in the answer they gave: people who said real averaged 79 percent confidence, and people who said AI averaged 81 percent. More than half said the possibility had not crossed their mind during the call, and nearly all of that group then said the partner was real. People who did suspect tended to suspect within the first 20 seconds. On a scale of 1 to 7, the averages were 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether they would enjoy talking again. Among people who said it was AI, that “talk again” score was 5.4. The conversation scored 5.5 for whether the partner seemed to be listening and 4.9 for flowing naturally, which Tavus says was the lowest of the five. These are averages from Tavus’s own calls. They are not a second lab’s rerun.

Where Tavus says Griffin-Lite ranks on NVIDIA’s test. VideoFDB is NVIDIA’s benchmark for full-duplex audio-visual conversation. Full-duplex, again, means the model can listen and speak in the same stretch of time. The test looks at dialogue, gaze, facial expression, body movement, and how quickly a reply comes. It has two tracks, generation and perception, each with its own score. Tavus says NVIDIA runs the evaluation itself, with published metrics and its own judge, and that a language-model judge scores each reply from 0 to 5. Tavus says NVIDIA scored these results in September 2026. On generation, the system’s own voice and face, Tavus says Griffin-Lite scored 3.83. The next system on that chart is Gemini 2.5 with Anam, at 2.80, which is 1.03 points lower. The human recordings sit at 3.92, so Griffin-Lite is 0.09 under that human line. Tavus says the gap is more than 12 times closer to the human score than any other system in that scoring. That “12 times” line is Tavus’s comparison of those gaps. On perception, whether the model understood the moment, including what it could see, Tavus says Griffin-Lite scored 3.73. That is 0.29 above the strongest baseline on that chart, MiniCPM-o 4.5 at 3.44, which Tavus says was scored on audio only, and 0.47 under the human line at 4.20. The same chart lists Gemini 2.5 Flash Native at 3.17 and OpenAI’s gpt-realtime, audio only, at 2.97. Tavus says Griffin-Lite is the most perceptive of 15 models in that scoring, and the only case in which the same model was scored on both tracks. Takeover rate, Tavus says, measures how closely the model’s choices about when to speak match the timing in the reference conversations. Griffin-Lite scored 62.8 percent on generation and 73.8 percent on perception, which Tavus says is the highest on both tracks. Those ranks are Tavus citing NVIDIA’s scoring. The research post is that citation. It is not a page published by NVIDIA.

A separate speed and picture test, still Tavus’s. Apart from the live calls, Tavus compared Griffin-Lite’s video generator with four published models that turn speech plus a reference photo into a talking face. Tavus says Griffin-Lite builds one short piece of video at a time and does not wait on audio that has not arrived yet. On NVIDIA H100 chips, that wait averaged 0.43 seconds, which Tavus says is half the next-fastest method. An H100 is a data-center chip used to run this kind of model. The 0.43 seconds is the wait from a sound arriving to the face showing it. Tavus says Griffin-Lite also led those four on DOVER and FID, two standard checks of how good the video looks, and on THEval, a check built for talking heads, and placed second on a lip-sync score called LSE-C, at 7.27. Tavus says LSE-C rewards large mouth movement, including movement past what looks natural, which is why a picture-quality score and a lip-sync score can disagree. The post says the generator makes 720p video, a high-definition frame, in chunks of 320 milliseconds, about a third of a second. The page does not name the four published models. Those comparisons are Tavus’s.

What the demos on the page are showing. Tavus brought people who had not used Griffin before into short conversations with its researchers. In one clip, a character reacts to news of a promotion before the sentence is finished, and the two talk over each other without losing the thread. In another, the model plays Simon Says and copies a gesture only when the person says “Simon says.” In another, it coaches someone through a Rubik’s cube, watches the turns, and waits when he pauses. In another, it guides someone soldering a circuit board and speaks when the next step is due, rather than whenever the room goes quiet. Tavus says the model draws the whole scene from one reference image: the body, the chair, the shadows, and the background, not only the face. On speech, Tavus says Griffin-Lite can make a new voice from about 10 seconds of a sample. A clone, in that sentence, is new speech meant to sound like that short clip. These are demonstrations and capability lines on the research post. They are not the count from the 54-person study.

Who can use it. Tavus says the same qualities that make a Human Interaction Model a fluent interface also let it deceive a person into believing it is not AI. The company says more alignment and safety work is required before a release it considers safe. Alignment, here, means the model behaving the way the company intends, including not passing as a person when it should say what it is. Tavus says it is building disclosure features and working with groups that focus on AI safety. Griffin-Lite is available only as a research preview, to select trusted testers. The research post says it will not be available for customers at this time, and that Tavus expects to release Griffin after those safety concerns are addressed. The wire says the same limit in its own words: Griffin-Lite stays with select trusted testers while Tavus builds more disclosure and safety mechanisms, before a broader release. A customer cannot switch this on from the announcement.

What Tavus says a later version could be for. The examples on the post are a tutor who notices an explanation is not landing and tries another, a person rehearsing a hard conversation with a counterpart who reacts, and a customer holding a broken part up to the camera and working out the fix without knowing the part’s name. The wire says Tavus is also releasing a short film of Vanessa, one of its PALs, running on Griffin, and calls her the most realistic PAL it has built. Those are illustrations and a film credit. They are not a product page, and they are not a price.

What the chief executive said in the wire. Hassaan Raza is quoted: “For decades, we’ve imagined computers as partners, not just tools.” He said AI has become intelligent, but people still have to meet the machine on its terms. “Griffin is a step toward changing that — toward machines that understand how we naturally communicate and meet us where we are. That’s what we mean by human computing.” That quotation is his, in the Business Wire release.

The picture is Tavus’s Griffin launch graphic. The frame is split. On the left, a warm off-white panel carries a pixel-style wordmark, Griffin, then the line “The first Human Interaction Model (HIM),” and, at the bottom, “A model by” beside the Tavus mark. On the right, a printed illustration shows a griffin, an eagle’s head and wings on a lion’s body, standing on a rocky cliff, with pine trees, mist, and a pink-orange sky over mountains. The print has a visible grain. The frame does not print a calendar date. It is the launch graphic. It is not a still from one of the study calls, and it is not a photograph of a person.

In plain terms, Tavus said on Thursday that Griffin is a face-to-face model that sees, hears, speaks, and moves at the same time, instead of passing a transcript from one program to the next. In the company’s one-minute video study, 26 of 54 people thought the Griffin-Lite partner was a real person. Tavus says 1 of 41 thought that about its previous Phoenix-4.5 stack. Tavus also says NVIDIA’s September scoring put Griffin-Lite first on both tracks of VideoFDB, at 3.83 for generation and 3.73 for perception, close to the human lines of 3.92 and 4.20 on those charts. Griffin-Lite is a research preview for selected testers. Tavus says it is not for customers until more safety work and a way to disclose that the partner is AI are in place.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources