AI agents · OpenClaw · self-hosting · automation

Quick Answer

What Is Tavus Griffin? The Video Turing Test AI (2026)

Published:

The short answer

Tavus Griffin is a single model that does what used to take four — perceive the person on the call, decide when to speak, and generate voice and face video at the same time — and in Tavus’s October 1, 2026 study 48% of people who video-called it for a minute thought it was human, up from 2.4% for the company’s previous pipeline. It leads NVIDIA’s VideoFDB benchmark at 3.83 vs a 3.92 human reference. It is also vendor-tested, vendor-reported, and not publicly available: Griffin-Lite is limited to early testers pending safety work. Facts verified October 4, 2026.

What makes it different

Every conversational video agent before Griffin was a pipeline: speech-to-text → language model → text-to-speech → lip-synced avatar renderer, with separate models for turn-taking and for reading the user’s face. Each hop adds latency and loses information — the renderer never sees the user, the LLM never hears the tone.

Griffin collapses that into one full-duplex video-to-video model. “Full-duplex” means it listens and speaks at the same time, like a person: it can back-channel, be interrupted, and start answering before you finish. “Video-to-video” means the input is your camera and microphone and the output is its face and voice, end to end. Tavus’s label for the category is Human Interaction Model (HIM).

Griffin-LitePrevious Tavus stack (Phoenix-4.5 + Sparrow-2 + Raven-1)
ArchitectureSingle full-duplex video-to-video modelPipeline: rendering + turn-taking + perception models
Output720p video in 320 ms chunksReal-time rendering
Audio-to-video latency0.43 s average (H100)Higher (multi-hop)
Input to create personaOne reference imageReference video/image
”Thought it was human” after 1-min call48% (26/54)2.4% (1/41)
VideoFDB generation score3.83 (human ref 3.92)—
AvailabilitySelect early testersGA via Tavus CVI API

The numbers, with their caveats

The 48% study. Tavus ran it: 54 participants, one-minute live video calls with a persona (Tavus calls them PALs) powered by Griffin-Lite, asked only at the end of a survey whether they had spoken to a real person. 26 said yes. The baseline — 1 of 41 for the Phoenix-4.5 pipeline — is the more informative number, because it shows the jump is architectural rather than incremental. The limits are equally clear: vendor-designed, vendor-run, a one-minute window (the hardest cases for avatars are minutes 5–20), and no standardised protocol. Tavus claims the first model to “pass the video Turing test”; a fairer reading is the first to fool close to half of participants in a short call.

VideoFDB. This is NVIDIA’s benchmark for full-duplex audiovisual conversation, with public leaderboards. Griffin-Lite’s 3.83 generation score sits 0.09 below the 3.92 human reference and a full point above the next system (Gemini 2.5 + Anam at 2.80). On perception — how well the model reads the person — it scores 3.73 against MiniCPM-o 4.5 (3.44), Gemini 2.5 Flash (3.17) and OpenAI gpt-realtime (2.97). That is the independent-ish part of the story: Tavus submitted the model, NVIDIA scored it.

Why it is not public

Tavus is holding Griffin-Lite to select early testers while it finishes safety evaluations and builds disclosure features — ways to make clear to a person on a call that they are talking to an AI. That is the right call and also the obvious one: a model whose headline metric is “people could not tell” is a fraud and social-engineering tool until disclosure is solved. The company’s shipping products stay on the Phoenix/Sparrow/Raven stack.

Who this matters for

  • Customer-facing video agents (sales, support, tutoring, telehealth intake): Griffin is the first signal that a single-model agent can hold a natural face-to-face conversation. Plan for it; do not build on it yet.
  • Competitors — HeyGen LiveAvatar, Synthesia, Anam, Google’s Gemini Live avatar — now have a public benchmark to beat. See Gemini 3.8 Live Avatar vs HeyGen vs Tavus vs Synthesia.
  • Trust & safety teams: 48% in one minute from a single reference photo is the new baseline for impersonation risk on video calls. Verification processes that rely on “it looks like a real person” are done.

Related: best AI avatar and video agent tools, Claude watermark vs SynthID vs C2PA provenance.

Last verified: October 4, 2026. Study and latency figures are Tavus’s; VideoFDB scores are from NVIDIA’s public leaderboard as reported October 1–3, 2026.

Sources