Skip to content

Provenance and privacy applied at generation.

never

reference audio in a training set

A cloning engine ships with obligations. Here are the three we take on, applied at generation rather than promised in a PDF.

The three commitments

fig.

Applied at the engine, not appended to a policy

Fig

every sample

carries an inaudible watermark, applied at generation on every request

no training

customer text and reference audio never enter a training set

per key

every workload’s traffic auditable on its own key at GET /v1/usage

The watermark is part of synthesis. No configuration skips it, on any tier, for any voice.

What this gives an audit

Regulated callers end up explaining synthetic audio to someone. The answers are ready before the questions: the audio proves it is synthetic (waveform watermark, survives transcoding), the key that generated it is on record (GET /v1/usage), and the reference clip went nowhere.

What we do not claim

  • The watermark does not identify the cloned speaker. It marks audio as synthetic, nothing more.
  • Consent is enforced as a requirement on you, not detected by us: the speaker in a reference clip has agreed to be cloned.
  • Detection tooling changes. Where we apply the mark does not: at generation, on every request.

Notes

Can I verify the watermark myself?

Send us a sample and we run the detector on it and send back what it says, the same discipline as the latency readout. Detection is software; the mark itself rides in the waveform.

Is anything about my traffic used to improve the model?

No. Your text and reference audio never enter a training set. We fingerprint and cache references so that reusing the same clip skips re-cloning.

A key, one stream, your own script, nothing on it counted while you build.

Get a key, run your own script