Tavus Griffin is a new research-preview system for face-to-face AI conversation, announced on 1 October 2026. Its striking headline is a Tavus-run one-minute study in which 26 of 54 callers thought the model was a person. The more useful story for businesses is what the launch does—and does not—prove about natural video agents, response timing, and disclosure.

Key takeaways

  • Limited access: Griffin-Lite is available to selected research testers; Tavus says it is not yet a customer product.
  • Company-run human-recognition result: 26 of 54 participants, or about 48%, judged a one-minute AI call to be human. This was not an independently replicated Turing test.
  • Independent benchmark: NVIDIA’s VideoFDB leaderboard lists Griffin-Lite at 3.83 out of five for generation and 3.73 for perception, while its timing measures remain behind human reference calls.
  • Business implication: The next challenge is trustworthy deployment: clear AI disclosure, consent, error handling, and evidence from longer real-world conversations.

What Tavus Griffin actually launched

Tavus, a San Francisco company building conversational video systems, introduced Griffin as what it calls a “Human Interaction Model.” The company says its system sees and hears a caller while producing speech and moving video, and can continue listening while it talks. That full-duplex design is the difference between a conversation and the familiar series of alternating voice messages.

The company has made Griffin-Lite available as a research preview to selected testers. Tavus explicitly says the preview is not available to customers yet. It has not announced general pricing or a date for broad release. Any business planning to deploy it in a support queue, sales call, or classroom should therefore treat those uses as prospective rather than available capabilities.

The launch matters because developers have been assembling AI video callers from separate components: speech recognition, a language model, speech synthesis, and an animated face. Each hand-off can lose information. A caller’s pause may signal uncertainty rather than the end of a turn; a facial expression may change the meaning of a sentence. Tavus says Griffin joins perception, conversational decisions, and audiovisual generation so the system can adjust within the same exchange. The architecture is described in Tavus’s dated research announcement.

That is a product claim and a design choice, not proof that every deployment will feel human. Real calls involve uneven connections, accents, interruptions, background noise, ambiguous visual context, and the obligations of the organization running the agent. These conditions require field testing beyond a polished demonstration.

Two approaches to video conversationA conventional relay hands audio through recognition, text generation, speech and animation. Griffin’s proposed loop receives audio and video continuously and generates voice and video while listening.How a video agent respondsRelay approachHear → transcribe → write reply → synthesize voice → animate faceGriffin’s claimed continuous approachListen + watch ↔ decide ↔ generate speech + video, with overlap

What the 48% result measures

In the experiment described by Tavus, 54 participants had a one-minute live video conversation with Griffin-Lite. They were told they would be paired with another participant. Afterward, 26 said the partner was a real person. The company compares that with one of 41 participants, roughly 2.4%, in a test of its prior video stack. Those are the company’s reported counts; the recruiting platform was not named, and the study has not been independently replicated.

A one-minute judgment is a narrow outcome. It does not show how a caller would assess a 20-minute support session, whether the model solves a problem correctly, or whether people still trust the organization after learning the caller was synthetic. Participant expectations also matter: a person who expects another human is being asked a different question from someone knowingly testing an AI system. CellCog’s independent examination points out that the study was run by Tavus and that participants were primed to expect another participant.

Calling the result a “passed Turing test” without qualification would turn a company-designed demonstration into a general scientific verdict. A fairer description is that Tavus achieved a substantial human-misidentification rate in its own short, specific video-call experiment. The company deserves credit for disclosing the sample sizes, but readers should wait for a published protocol and independent replication before drawing wider conclusions.

The measurement is still newsworthy. Remote support, interviews, tutoring and sales all depend partly on real-time social cues. Even a short call can show whether an interface is becoming more responsive. But the metric to watch in a commercial setting may be more practical: successful task completion, correct escalation, consent, user satisfaction after disclosure, and whether people can interrupt and redirect the agent reliably.

How NVIDIA’s VideoFDB changes the picture

Unlike the human-recognition experiment, NVIDIA’s VideoFDB leaderboard is a separate benchmark maintained by NVIDIA researchers. It uses 237 clips from real video conversations and assesses perception and generation with language-model judges. The published table lists Griffin-Lite at 3.83 out of five on the generation track, compared with 3.92 for the human reference and 2.80 for the next listed cascaded avatar. On perception, it lists Griffin-Lite at 3.73, compared with 4.20 for the human reference.

Those overall scores are a useful independent check on the company’s comparative claims. They do not show equivalence across every measure. On the generation track, Griffin-Lite’s nonverbal-cue appropriateness score is 2.83 against the human reference of 3.18. Its timing alignment is 62.8% against 78% for the human reference. The benchmark also reports a median generation latency of 1,892 milliseconds for Griffin-Lite against 900 milliseconds for the human reference.

That matters because Tavus separately cites a 0.43-second average audio-to-video response measure on its own hardware setup. The two numbers describe different tests, so one should not erase the other. The headline reaction measure is not the same as the latency on NVIDIA’s end-to-end conversational benchmark. XenoSpectrum’s independent analysis likewise highlights the timing and nonverbal-cue gaps rather than treating an overall score near the human reference as complete parity.

VideoFDB uses an automated judge and a curated sample, not a random sample of all customer calls. Its value is comparison under common conditions. A company considering a video agent should therefore combine benchmark results with tests on its own accents, cameras, bandwidth, products, and escalation rules.

NVIDIA VideoFDB scores and timingGriffin-Lite scores 3.83 versus human reference 3.92 in generation and 3.73 versus human reference 4.20 in perception. Its generation timing alignment is 62.8% versus human 78%.VideoFDB: close scores, visible timing gapNVIDIA leaderboard; scores are out of 5GenerationHuman 3.92Griffin 3.83PerceptionHuman 4.20Griffin 3.73Generation timing alignment: Human 78% · Griffin 62.8%Grey = human reference; red = Griffin-Lite. Not a customer-call outcome.

Why access is restricted

Tavus acknowledges that a natural-looking AI caller can deceive. It says it is working on disclosure and safety features before customer release. This is a material limit, not a minor footnote. An agent that is misidentified as a person in the company’s test creates foreseeable risks for impersonation, inappropriate reliance, and hidden automation. The appropriate response is to design transparent deployment rules before scale.

There is no evidence in Tavus’s announcement that Griffin already ships with a cryptographic watermark or a universal identity check. The earlier version of this article described such features as expected; that was speculative and has been removed. The company says it is developing disclosure mechanisms but has not announced their final specification. Any security claim should be checked against a released product rather than inferred from a roadmap.

Responsible testing can be concrete. Tell participants they are interacting with AI. Require consent to use a likeness or voice. Define which tasks the agent may complete and which require a person. Log incidents where it misreads a visual cue or continues after a caller tries to interrupt. Measure whether disclosure changes trust or task completion. These are editorial recommendations, not capabilities Tavus says the preview includes.

For a company in India, the immediate consequence is strategic rather than a ready-to-buy product. A research preview shows where service interfaces may be headed, but it does not establish local availability, pricing, or compliance terms. Indian businesses assessing video AI should seek a data-processing agreement, retention terms, consent flow, and performance results on their own users before procurement. Our coverage of ElevenLabs’ voice-AI valuation shows how investment in expressive interfaces is accelerating, while the Airbnb AI search rollout illustrates why geographic availability must be stated precisely.

What to watch next

The first milestone is a public or customer release with clear identity disclosure and terms. The second is independent human testing, preferably with longer calls, disclosed recruitment methods, and diverse accents and connection conditions. The third is repeatable task data: can a system give a correct answer, handle interruption, and transfer a difficult conversation to a person? Those measures would show whether a video agent improves service rather than merely seeming human for a minute.

One open technical question is whether Griffin’s unified approach keeps its advantage when connected to real company knowledge. A polished conversation still needs correct facts, permissions, and timely escalation. Another question is how it behaves when video is unavailable or delayed. In many mobile settings, the best interface may be audio or text rather than a synthetic face. The useful comparison will be against the same job performed through those alternatives.

The clearest reading of Tavus Griffin today is this: Tavus announced an ambitious full-duplex video model, its own short study found that 26 of 54 callers mistook the preview for a person, and NVIDIA independently lists strong—but not uniformly human-level—benchmark scores. Griffin-Lite remains a selected research preview while the company works on disclosure and safety. That distinction lets readers understand the advance without mistaking a launch demonstration for a broadly available or independently validated human replacement.

Frequently asked questions

Can businesses use Tavus Griffin now?

No general customer availability has been announced. Tavus says Griffin-Lite is available to selected research testers, while safety and disclosure work continues.

Did Tavus Griffin pass a video Turing test?

Tavus calls its result a pass. In the company’s one-minute study, 26 of 54 participants thought they had spoken to a person. It is a company-run, limited test, not an independent verdict on longer conversations.

How does Griffin compare on NVIDIA’s benchmark?

NVIDIA lists Griffin-Lite at 3.83 out of five in generation and 3.73 in perception. Human references score 3.92 and 4.20 respectively. Timing and nonverbal-cue measures show remaining gaps.

Why does disclosure matter?

When an AI system can be mistaken for a person, callers need to know who—or what—is speaking. Disclosure, consent, and human escalation are necessary for trustworthy customer use.

Sources: Tavus research announcement, 1 October; NVIDIA VideoFDB leaderboard; independent reviews from CellCog, XenoSpectrum, and Kingy AI. Performance claims from Tavus’s human-recognition study remain attributed to the company.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.