I get asked a version of this question often: "If voice cloning is so convincing that family members can't tell the difference, how can Savi?"
It's a fair challenge. The short answer is that a real-time detector doesn't need to be perfect to be useful, and it doesn't work the same way a human listener does. We're not trying to recognize a specific person's voice. We're looking for a different class of signals entirely.
This is a technical article about what we actually measure and why. I'm writing it because I think understanding the detection approach helps users understand what Savi can and cannot do, which matters more to us than making the technology sound like magic.
What Synthetic Voices Actually Sound Like at the Signal Level
Text-to-speech and voice conversion systems have improved dramatically, but they still produce audio that differs from a natural human voice in ways that are detectable in the signal, even when they are not obvious to a human ear.
The most consistent differences are in spectral and temporal fine structure: the micro-variations in frequency, amplitude, and timing that characterize real human phonation. Natural speech is produced by a physical system with vocal tract resonances that shift continuously due to breathing, muscular tension, emotional state, and physical movement. Synthetic voices are generated by mathematical models that approximate those patterns but smooth over many of the irregular, non-periodic components.
Concretely, this shows up as: reduced aperiodic components in voiced sounds, unnaturally consistent formant transitions, flattened prosodic contours in unstressed syllables, and artifacts at phoneme boundaries where the generation model stitches together segments. The better modern models are, the smaller these differences are, but they have not disappeared entirely.
Additionally, voice conversion (taking one person's voice and mapping it to another's) introduces its own artifacts. The conversion layer must preserve intelligibility while transforming speaker characteristics, and it tends to produce characteristic distortions in higher-frequency bands, particularly in fricatives and affricates, that differ from what a natural voice produces on a phone codec.
The Latency Constraint: Detection Must Work on a Rolling Window
Detection at the signal level is reasonably tractable for a recorded sample with sufficient length. The hard part of what we do is that detection has to work on a rolling window of live audio, with a latency low enough to be useful during the call.
This eliminates approaches that require more than a few hundred milliseconds of audio to produce a reliable signal. It also means the detector has to handle varying network conditions, codec compression artifacts, and the acoustic environment on both ends of the call, because all of those introduce their own distortions that can look superficially similar to synthesis artifacts.
The architecture we settled on processes audio in short overlapping frames and computes a set of acoustic features that have been calibrated on a corpus of both genuine calls and calls with known synthetic voices, collected under realistic conditions (various codecs, background noise levels, call quality). The output is a running confidence score that updates continuously rather than a single binary decision.
What this means in practice: Savi does not wait for a full sentence before producing a signal. It updates its assessment every few hundred milliseconds. A call that looks synthetic from the first moment continues to accumulate evidence; a call that looks genuine initially but then introduces a voice conversion can be caught mid-call.
Why We Don't Rely Only on Acoustic Detection
Acoustic detection alone would produce too many false positives at the call quality levels users actually experience. A voice on a poor connection, or coming through a noisy environment, or from an elderly caller with a weakened voice can produce acoustic profiles that overlap with synthetic speech in ways that are hard to cleanly separate.
The second layer of detection is conversational structure. Scam calls follow recognizable patterns: the escalation sequence, the urgency injection, the payment demand, the instructions not to hang up. These structural patterns have been stable across thousands of documented cases, and they hold even as the specific scripts evolve.
Natural language processing on the call transcript runs in parallel with the acoustic analysis. When both layers are in agreement, the confidence is much higher than either alone. When they diverge, the combined score is conservative. We flag calls where the acoustic signal is ambiguous but the conversational structure is clearly high-risk, and we flag calls where the acoustic signal looks strongly synthetic even if the conversation hasn't yet reached a payment demand.
This two-signal approach also helps with the adversarial case: a caller who is aware of acoustic detection might alter their synthetic voice parameters to reduce detectable artifacts. But the conversational pattern is harder to change without making the scam less effective. The urgency and payment demand are structurally necessary to the scam. Removing them to avoid detection means the scam doesn't work.
What Savi Cannot Do
It's important to be clear about the limits.
We cannot reliably detect a high-quality voice clone when the call is very short and the audio quality is excellent. If a scammer uses a state-of-the-art voice cloning system, calls on a clean connection, and keeps the call to just a few seconds before a callback number is left, acoustic detection alone will not produce a confident result. In that scenario we rely on number reputation, caller pattern matching, and whatever transcript signal is available.
We also cannot verify identity. Detecting that a voice is synthetic tells you the voice is probably not a real person's natural speech. It does not tell you who the real person is, or confirm that a genuine-sounding voice is actually the person they claim to be. A human scammer calling with their own voice would not be detected by acoustic analysis at all. Our detection of that scenario relies entirely on conversational pattern recognition.
This is why we always describe Savi as a tool to use alongside good judgment, not instead of it. A call that passes every acoustic and conversational check can still be a scam. The phone code word strategy we recommend for families is effective precisely because it doesn't depend on detection: it creates a verification that a scammer cannot defeat even with a perfect voice clone.
How the Detection Improves Over Time
The model we deployed at launch was calibrated on the synthetic voice systems that were most common at that point: specific families of text-to-speech and voice conversion tools that dominated the market for accessible, low-cost voice cloning. That corpus is now larger, and we have been able to add labeled examples from reports of actual scam calls users have sent us.
The challenge is keeping up with new synthesis systems as they release. Each new generation of voice generation models requires a calibration pass to understand where its artifact profile sits relative to existing profiles and how much our existing feature set covers it. We maintain an evaluation pipeline that tests new synthesis systems against the detector periodically and triggers retraining when coverage drops below a threshold.
I want to be direct about what this means: there will always be a window between when a new synthesis system becomes widely available and when we have good coverage of it. That window is a real vulnerability. It's one reason the behavioral and conversational layers matter: even a synthetic voice the detector hasn't seen before will typically produce the same conversational pattern as every other phone scam.
Why This Is a Solvable Problem Over the Medium Term
I believe voice authenticity detection is a tractable problem, not a fundamental impossibility, for the following reason: the physical constraints of natural voice production are real, and they leave a signal in audio that synthetic systems cannot fully reproduce without physically simulating a human vocal tract. Better synthesis narrows the gap, but the cost and complexity of modeling all the micro-dynamics of natural voice production continues to favor detection over generation, especially at the quality levels accessible to mass-market scam operations.
The threat we worry more about in the medium term is not better synthetic voices in isolation. It is the combination of synthetic voice with better conversational AI, which could produce calls that look like natural conversation from start to finish rather than following a rigid script. That would reduce the effectiveness of the conversational detection layer and force more reliance on the acoustic layer. That's the arms race we are most actively preparing for.
What we have today is not a complete solution, but it is meaningfully protective for the call patterns that are actually being used in phone fraud right now. And the families who use it have something they didn't have before: a second pair of eyes on calls where the caller's voice is the main thing being trusted.