Instant voice cloning
10 s
of reference, the whole enrollment
Zero-shot cloning builds a voice from a short reference clip at request time, no training job, no per-voice model. Ours takes about ten seconds of audio.
Instant vs trained cloning
Trained cloning fits a model to hours of a speaker’s audio, per voice, often per language. Zero-shot conditions a single model on a reference clip inside the request: send ten seconds, get that identity back. New voice, no pipeline.
table 01
The two cloning disciplines, side by side
| Trained cloning | Zero-shot cloning | |
|---|---|---|
| Enrollment | hours of audio, per voice | ~10 seconds, in the request |
| Wait | a training job, hours to days | none, first request pays a cache fill |
| New language | often a new training run | same reference, 23 languages |
| Cost per voice | a per-voice fee or job | nothing marginal |
The cross-language test
The hard case is a language the reference never spoke, where the identity has to survive without any of the original phonemes. That held up in blind listening panels. One reference carries one identity across 23 languages.
Every term on this page is measurable. Take a key and read the numbers off your own requests.
Get a key and measure it yourself