Delivery is a request parameter.
1
seed, the same take, every render
Most engines give you a voice and a volume knob. This one exposes the prosody itself, melodic range, pacing, speed and gain, each a field on the ordinary request, and a seed so the take you liked is the take you keep.
The dial that decides how a line lands
temperature sets prosodic variation, pitch range and melody, from 0.1 to 1.2. Unlike a named preset it is continuous, so a read can sit anywhere between recited and performed rather than at one of eight fixed points.
We are not publishing a curve for it at the moment. The one this page carried was measured on the previous engine, and on our own re-measurement the effect is too small to separate from the variation between seeds. When we have a figure we can reproduce, it goes here.
The rest of the surface
The two that interact are temperature and cfg_weight: a bigger delivery wants more room, so raise the melody and loosen the guidance together, or the read just gets hurried. expressiveness is still accepted and no longer moves anything, it was remapped off the engine, and a controlled measurement puts its effect at +0.33 semitones, p=0.804.
- cfg_weight, guidance strength, which also sets pacing: 0.2 slow and spacious to 1.0 tight and brisk.
- speed, 0.6 to 1.5, applied after synthesis, so wording and voice are untouched.
- volume, 0.5 to 2.0, applied after mastering with a soft ceiling, so it never hard-clips.
- seed, the same seed, text, voice and dials return the same audio, so a take you like is a take you can keep.
What happened to the emotion field
This sheet used to document a body field that named an emotional read and set dial values under the hood. It was retired on 2026-07-29 and the engine now discards it: the same request at a fixed seed returns identical audio whichever emotion you name. It is still accepted, so nothing an old client sends will break.
What replaced it is better anyway, and it is the whole argument of this page: a named preset is eight or sixteen fixed points, and the dials are continuous. Write the line the way a person would say it, then set the melody and the pacing yourself.
And the detail you place by hand
These ride inline on every surface, bytes, SSE and WebSocket alike, and none costs a round trip. First audio holds its published behaviour with every field on this page set, and each response reports its own ttfa_ms.
- <spell>TKT4829XB</spell>, reads character by character, with letter and digit groups paced naturally. For confirmation codes, order IDs and serials.
- <break time="800ms"/> puts a real pause exactly where the read needs air. The tag is never spoken.
- pronunciation_dict, per-request sounds-like replacements for proper nouns and domain terms, so a brand name is right on the first call, not after a support ticket.
Notes
How much can the delivery actually change?
One voice, one sentence, a seed per setting: temperature moves pitch range and melodic shape continuously, so the read can sit anywhere between recited and performed, from one request field.
Is expression a separate model or a slower path?
Neither. Every control here is a field on the ordinary synthesis request, conditioning the same voice on the same endpoint. No expressive tier, no second model, and first audio holds its published behaviour with the fields set.
A key, one stream, your own script, nothing on it counted while you build.
Get a key, run your own script