> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.jambonz.org/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.jambonz.org/_mcp/server.

# Inworld

Inworld tuning is passed in `synthesizer.options`. Audio shaping lives under a nested `audioConfig` object, matching Inworld's own API.

```json
{
  "verb": "say",
  "text": "Thanks for calling, how can I help you today?",
  "synthesizer": {
    "vendor": "inworld",
    "language": "en",
    "voice": "Ashley",
    "options": {
      "temperature": 0.8,
      "audioConfig": {
        "pitch": 0.0,
        "speakingRate": 1.0
      }
    }
  }
}
```

## Options

**`temperature`** `number`

Sampling temperature. Defaults to `0.8`. Lower values give a more consistent, predictable read; higher values more variation between renders of the same text.

---

### audioConfig

**`audioConfig.speakingRate`** `number`

Speaking rate multiplier. Defaults to `1.0`. Values above `1.0` are faster, below are slower.

---

**`audioConfig.pitch`** `number`

Pitch adjustment. Defaults to `0.0`, the voice's natural pitch.

---

**`audioConfig.sampleRateHertz`** `number`

Output sample rate. jambonz already requests the rate that matches the call's codec, so leave this unset unless you have a specific reason — overriding it forces a resample.

---

**`audioConfig.bitRate`** `number`

Output bit rate.

---

## Models

Pick the model on the speech credential in the portal.

| Model                                          | Notes                                                                |
| ---------------------------------------------- | -------------------------------------------------------------------- |
| `inworld-tts-2`                                | Current generation, the best quality.                                |
| `inworld-tts-2-flash`                          | Same generation tuned for latency and cost — the fastest first byte. |
| `inworld-tts-1.5-max` / `inworld-tts-1.5-mini` | Previous generation.                                                 |
| `inworld-tts-1` / `inworld-tts-1-max`          | Deprecated, no word timestamps.                                      |

## Streaming

Inworld supports [TTS streaming](/guides/features/tts-streaming), including
token-by-token streaming for `say` with `stream: true` and for the `agent` verb,
where text is spoken as your LLM produces it.

`audioConfig.pitch` applies to non-streaming synthesis only; Inworld's streaming
API does not accept it.

## Word timestamps

On `inworld-tts-2`, `inworld-tts-2-flash`, `inworld-tts-1.5-mini` and
`inworld-tts-1.5-max`, Inworld returns word-level timing along with the audio.
jambonz uses it to track how much of a response the caller actually heard, which powers
[`trackTtsPlayout`](/guides/features/tts-streaming) and lets a barge-in trim the
conversation history to the words spoken before the interruption. No configuration
is needed beyond enabling that feature — jambonz requests the timing itself.

`inworld-tts-1` and `inworld-tts-1-max` return no timing, so playout tracking is
unavailable on them. Both are marked deprecated in the portal and are no longer
offered for new speech credentials; existing credentials that use them keep
working, and you can still select them if you need to.