> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.jambonz.org/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.jambonz.org/_mcp/server.

# Python SDK

The `jambonz-python-sdk` package lets you build jambonz voice applications in Python. It supports webhook and WebSocket transports, a REST API client for mid-call control, and chainable verb methods for building call flows.

**Source code:** [github.com/jambonz/python-sdk](https://github.com/jambonz/python-sdk)

> **Warning**
>
> **Experimental** — This SDK is under active development. APIs may change between releases. Please report issues on [GitHub](https://github.com/jambonz/python-sdk/issues).

## Installation

```bash
pip install jambonz-python-sdk
```

## Imports

The SDK provides three submodule imports:

```python
# Webhook apps (aiohttp, FastAPI, Flask, etc.)
from jambonz_sdk.webhook import WebhookResponse

# WebSocket apps
from jambonz_sdk.websocket import create_endpoint

# REST API client (mid-call control, outbound calls)
from jambonz_sdk.client import JambonzClient
```

## Webhook Transport

Use `WebhookResponse` to build verb arrays in response to HTTP webhooks. Methods are chainable and the response is serialized to JSON.

```python
from aiohttp import web
from jambonz_sdk.webhook import WebhookResponse

async def handle_incoming(request: web.Request) -> web.Response:
    jambonz = WebhookResponse()
    jambonz.say(text="Hello from jambonz!").gather(
        input=["speech", "digits"],
        actionHook="/handle-input",
        say={"text": "Press 1 for sales or 2 for support."},
    ).hangup()
    return web.json_response(jambonz.to_json())

async def handle_input(request: web.Request) -> web.Response:
    body = await request.json()
    speech = body.get("speech", {}).get("alternatives", [{}])[0].get("transcript", "")
    jambonz = WebhookResponse()
    jambonz.say(text=f"You said: {speech}").hangup()
    return web.json_response(jambonz.to_json())

app = web.Application()
app.router.add_post("/incoming", handle_incoming)
app.router.add_post("/handle-input", handle_input)
web.run_app(app, port=3000)
```

## WebSocket Transport

Use `create_endpoint` to build real-time WebSocket applications. This is the recommended transport for voice AI agents, as it enables bidirectional communication, event streaming, and mid-call updates.

```python
import asyncio
from jambonz_sdk.websocket import create_endpoint

async def main():
    make_service, runner = await create_endpoint(port=3000)
    svc = make_service(path="/")

    async def handle_session(session):
        async def on_gather_result(evt):
            transcript = (
                evt.get("speech", {})
                .get("alternatives", [{}])[0]
                .get("transcript", "")
            )
            session.say(text=f"You said: {transcript}").hangup()
            await session.reply()

        session.on("/gather-result", on_gather_result)

        session.say(text="Hello! Say something.").gather(
            input=["speech"],
            actionHook="/gather-result",
            timeout=10,
        ).hangup()
        await session.send()

    svc.on("session:new", handle_session)
    await asyncio.Future()

asyncio.run(main())
```

### .send() vs .reply()

* **`await session.send()`** — Use once for the initial verb array in response to `session:new`.
* **`await session.reply()`** — Use for all subsequent responses to actionHook events.

### Application Environment Variables

You can declare environment variables that are configurable in the jambonz portal UI:

```python
make_service, runner = await create_endpoint(
    port=3000,
    env_vars={
        "OPENAI_MODEL": {
            "type": "string",
            "description": "LLM model to use",
            "default": "gpt-4.1-mini",
        },
        "SYSTEM_PROMPT": {
            "type": "string",
            "description": "System prompt",
            "uiHint": "textarea",
            "default": "You are a helpful assistant.",
        },
    },
)

# Read values in session handler
async def handle_session(session):
    model = session.data.get("env_vars", {}).get("OPENAI_MODEL", "gpt-4.1-mini")
    # ...
```

## Audio Streams

When using the [listen](/verbs/verbs/listen) verb, `make_service.audio()` lets you handle both call control and audio on the same server:

```python
svc = make_service(path="/")
audio_svc = make_service.audio(path="/audio-stream")

async def handle_session(session):
    session.say(text="Listening...").listen(
        url="/audio-stream",
        sampleRate=8000,
        bidirectionalAudio={"enabled": True, "streaming": True, "sampleRate": 8000},
    )
    await session.send()

svc.on("session:new", handle_session)

async def handle_audio(stream):
    async def on_audio(pcm):
        # Process audio — feed to STT, record, etc.
        pass

    stream.on("audio", on_audio)

    # Send audio back
    stream.send_audio(pcm_buffer)

audio_svc.on("connection", handle_audio)
```

## REST API Client

Use `JambonzClient` for outbound calls and mid-call control:

```python
from jambonz_sdk.client import JambonzClient

async with JambonzClient(
    base_url="https://api.jambonz.us",
    account_sid="your-account-sid",
    api_key="your-api-key",
) as client:
    # Create an outbound call
    call_sid = await client.calls.create({
        "from": "+15085551212",
        "to": {"type": "phone", "number": "+15085551213"},
        "call_hook": "/incoming",
    })

    # Mid-call control
    await client.calls.mute(call_sid, "mute")
    await client.calls.redirect(call_sid, "https://example.com/new-flow")
```

## Verb Methods

Both `WebhookResponse` and WebSocket `Session` support the same chainable verb methods:

`.say()` `.play()` `.gather()` `.dial()` `.llm()` `.agent()` `.conference()` `.enqueue()` `.dequeue()` `.hangup()` `.pause()` `.redirect()` `.config()` `.tag()` `.dtmf()` `.listen()` `.transcribe()` `.message()` `.dub()` `.alert()` `.answer()` `.leave()` `.sip_decline()` `.sip_refer()` `.sip_request()`

All methods accept the same options as the corresponding [verb JSON schemas](/verbs/verbs/overview) and are chainable.

### Spec-Driven Verb Generation

The SDK does **not** hardcode verb method signatures. Verb methods are auto-generated at import time from [JSON Schema](https://github.com/jambonz/schema) files — the same schemas used by the Node.js SDK and the jambonz server. When the schema adds a new property to a verb, the SDK picks it up automatically with no code change needed.

## TTS Token Streaming

The WebSocket Session provides methods for incremental TTS token streaming, enabling low-latency voice AI experiences:

```python
async def on_llm_tokens(evt):
    tokens = evt.get("tokens")
    done = evt.get("done")

    if tokens:
        await session.send_tts_tokens(tokens)

    if done:
        session.flush_tts_tokens()

session.on("/llm-tokens", on_llm_tokens)
```

| Method                  | Description                                                                          |
| ----------------------- | ------------------------------------------------------------------------------------ |
| `send_tts_tokens(text)` | Send a chunk of text for TTS. Awaitable; resolves when jambonz acknowledges receipt. |
| `flush_tts_tokens()`    | Signal the end of a TTS token stream.                                                |
| `clear_tts_tokens()`    | Cancel all pending TTS tokens and reset state.                                       |

## LLM and Agent Updates

### Tool Output

When the LLM requests a tool/function call, respond with the result:

```python
async def on_tool_call(evt):
    tool_call_id = evt["tool_call_id"]
    name = evt["name"]
    args = evt["arguments"]
    result = await handle_tool(name, args)
    session.send_tool_output(tool_call_id, result)

session.on("/tool-call", on_tool_call)
```

### Agent Updates

Send mid-conversation updates to an active [agent](/verbs/verbs/agent):

```python
# Change the system prompt
session.update_agent({
    "type": "update_instructions",
    "instructions": "You are now a billing agent.",
})

# Inject context
session.update_agent({
    "type": "inject_context",
    "messages": [{"role": "user", "content": "Customer: Sarah, Gold tier."}],
})

# Replace tools
session.update_agent({"type": "update_tools", "tools": [...]})

# Trigger a new response
session.update_agent({
    "type": "generate_reply",
    "interrupt": True,
    "user_input": "Tell the customer about the flash sale.",
})
```

### LLM Updates

```python
session.update_llm({"instructions": "Switch to Spanish."})
```

| Method                                 | Description                                             |
| -------------------------------------- | ------------------------------------------------------- |
| `send_tool_output(tool_call_id, data)` | Send tool/function result back to the LLM or agent verb |
| `update_agent(data)`                   | Send an `agent:update` command                          |
| `update_llm(data)`                     | Send an `llm:update` command                            |

## Inject Commands

Inject commands execute immediately on an active call without affecting the verb stack:

```python
# Mute/unmute
session.inject_mute("mute")
session.inject_mute("unmute")

# Whisper to one party
session.inject_whisper({"verb": "say", "text": "The customer is a VIP."}, agent_call_sid)

# Control noise isolation
session.inject_noise_isolation("enable", vendor="krisp", level=80)
session.inject_noise_isolation("disable")

# Control recording
session.inject_record("startCallRecording", siprec_server_url="sip:siprec@recorder.example.com")
session.inject_record("pauseCallRecording")

# Send DTMF
session.inject_dtmf("1234")

# Redirect call flow
session.inject_redirect("/new-webhook")

# Tag the call with metadata
session.inject_tag({"priority": "high", "department": "billing"})
```

## Session Properties

| Property          | Type   | Description                                                     |
| ----------------- | ------ | --------------------------------------------------------------- |
| `call_sid`        | `str`  | Unique call identifier                                          |
| `from_number`     | `str`  | Caller phone number or SIP URI                                  |
| `to`              | `str`  | Called phone number or SIP URI                                  |
| `direction`       | `str`  | `'inbound'` or `'outbound'`                                     |
| `account_sid`     | `str`  | Account identifier                                              |
| `application_sid` | `str`  | Application identifier                                          |
| `call_id`         | `str`  | SIP Call-ID                                                     |
| `data`            | `dict` | Full call session data (includes `env_vars`, SIP headers, etc.) |

## Examples

See the [examples directory](https://github.com/jambonz/python-sdk/tree/main/examples) for runnable demos:

| Example       | Webhook | WebSocket | Description                  |
| ------------- | ------- | --------- | ---------------------------- |
| hello-world   | yes     | yes       | Minimal greeting             |
| echo          | yes     | yes       | Speech echo with gather      |
| ivr-menu      | yes     | —         | IVR menu with speech + DTMF  |
| voice-agent   | yes     | yes       | LLM pipeline with tool calls |
| dial          | yes     | —         | Outbound dial with fallback  |
| listen-record | yes     | yes       | Audio recording              |