> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.jambonz.org/sdks/python-sdk/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.jambonz.org/_mcp/server. # Python SDK The `jambonz-python-sdk` package lets you build jambonz voice applications in Python. It supports webhook and WebSocket transports, a REST API client for mid-call control, and chainable verb methods for building call flows. **Source code:** [github.com/jambonz/python-sdk](https://github.com/jambonz/python-sdk) > **Warning** > > **Experimental** — This SDK is under active development. APIs may change between releases. Please report issues on [GitHub](https://github.com/jambonz/python-sdk/issues). ## Installation ```bash pip install jambonz-python-sdk ``` ## Imports The SDK provides three submodule imports: ```python # Webhook apps (aiohttp, FastAPI, Flask, etc.) from jambonz_sdk.webhook import WebhookResponse # WebSocket apps from jambonz_sdk.websocket import create_endpoint # REST API client (mid-call control, outbound calls) from jambonz_sdk.client import JambonzClient ``` ## Webhook Transport Use `WebhookResponse` to build verb arrays in response to HTTP webhooks. Methods are chainable and the response is serialized to JSON. ```python from aiohttp import web from jambonz_sdk.webhook import WebhookResponse async def handle_incoming(request: web.Request) -> web.Response: jambonz = WebhookResponse() jambonz.say(text="Hello from jambonz!").gather( input=["speech", "digits"], actionHook="/handle-input", say={"text": "Press 1 for sales or 2 for support."}, ).hangup() return web.json_response(jambonz.to_json()) async def handle_input(request: web.Request) -> web.Response: body = await request.json() speech = body.get("speech", {}).get("alternatives", [{}])[0].get("transcript", "") jambonz = WebhookResponse() jambonz.say(text=f"You said: {speech}").hangup() return web.json_response(jambonz.to_json()) app = web.Application() app.router.add_post("/incoming", handle_incoming) app.router.add_post("/handle-input", handle_input) web.run_app(app, port=3000) ``` ## WebSocket Transport Use `create_endpoint` to build real-time WebSocket applications. This is the recommended transport for voice AI agents, as it enables bidirectional communication, event streaming, and mid-call updates. ```python import asyncio from jambonz_sdk.websocket import create_endpoint async def main(): make_service, runner = await create_endpoint(port=3000) svc = make_service(path="/") async def handle_session(session): async def on_gather_result(evt): transcript = ( evt.get("speech", {}) .get("alternatives", [{}])[0] .get("transcript", "") ) session.say(text=f"You said: {transcript}").hangup() await session.reply() session.on("/gather-result", on_gather_result) session.say(text="Hello! Say something.").gather( input=["speech"], actionHook="/gather-result", timeout=10, ).hangup() await session.send() svc.on("session:new", handle_session) await asyncio.Future() asyncio.run(main()) ``` ### .send() vs .reply() * **`await session.send()`** — Use once for the initial verb array in response to `session:new`. * **`await session.reply()`** — Use for all subsequent responses to actionHook events. ### Application Environment Variables You can declare environment variables that are configurable in the jambonz portal UI: ```python make_service, runner = await create_endpoint( port=3000, env_vars={ "OPENAI_MODEL": { "type": "string", "description": "LLM model to use", "default": "gpt-4.1-mini", }, "SYSTEM_PROMPT": { "type": "string", "description": "System prompt", "uiHint": "textarea", "default": "You are a helpful assistant.", }, }, ) # Read values in session handler async def handle_session(session): model = session.data.get("env_vars", {}).get("OPENAI_MODEL", "gpt-4.1-mini") # ... ``` ## Audio Streams When using the [listen](/verbs/verbs/listen) verb, `make_service.audio()` lets you handle both call control and audio on the same server: ```python svc = make_service(path="/") audio_svc = make_service.audio(path="/audio-stream") async def handle_session(session): session.say(text="Listening...").listen( url="/audio-stream", sampleRate=8000, bidirectionalAudio={"enabled": True, "streaming": True, "sampleRate": 8000}, ) await session.send() svc.on("session:new", handle_session) async def handle_audio(stream): async def on_audio(pcm): # Process audio — feed to STT, record, etc. pass stream.on("audio", on_audio) # Send audio back stream.send_audio(pcm_buffer) audio_svc.on("connection", handle_audio) ``` ## REST API Client Use `JambonzClient` for outbound calls and mid-call control: ```python from jambonz_sdk.client import JambonzClient async with JambonzClient( base_url="https://api.jambonz.us", account_sid="your-account-sid", api_key="your-api-key", ) as client: # Create an outbound call call_sid = await client.calls.create({ "from": "+15085551212", "to": {"type": "phone", "number": "+15085551213"}, "call_hook": "/incoming", }) # Mid-call control await client.calls.mute(call_sid, "mute") await client.calls.redirect(call_sid, "https://example.com/new-flow") ``` ## Verb Methods Both `WebhookResponse` and WebSocket `Session` support the same chainable verb methods: `.say()` `.play()` `.gather()` `.dial()` `.llm()` `.agent()` `.conference()` `.enqueue()` `.dequeue()` `.hangup()` `.pause()` `.redirect()` `.config()` `.tag()` `.dtmf()` `.listen()` `.transcribe()` `.message()` `.dub()` `.alert()` `.answer()` `.leave()` `.sip_decline()` `.sip_refer()` `.sip_request()` All methods accept the same options as the corresponding [verb JSON schemas](/verbs/verbs/overview) and are chainable. ### Spec-Driven Verb Generation The SDK does **not** hardcode verb method signatures. Verb methods are auto-generated at import time from [JSON Schema](https://github.com/jambonz/schema) files — the same schemas used by the Node.js SDK and the jambonz server. When the schema adds a new property to a verb, the SDK picks it up automatically with no code change needed. ## TTS Token Streaming The WebSocket Session provides methods for incremental TTS token streaming, enabling low-latency voice AI experiences: ```python async def on_llm_tokens(evt): tokens = evt.get("tokens") done = evt.get("done") if tokens: await session.send_tts_tokens(tokens) if done: session.flush_tts_tokens() session.on("/llm-tokens", on_llm_tokens) ``` | Method | Description | | ----------------------- | ------------------------------------------------------------------------------------ | | `send_tts_tokens(text)` | Send a chunk of text for TTS. Awaitable; resolves when jambonz acknowledges receipt. | | `flush_tts_tokens()` | Signal the end of a TTS token stream. | | `clear_tts_tokens()` | Cancel all pending TTS tokens and reset state. | ## LLM and Agent Updates ### Tool Output When the LLM requests a tool/function call, respond with the result: ```python async def on_tool_call(evt): tool_call_id = evt["tool_call_id"] name = evt["name"] args = evt["arguments"] result = await handle_tool(name, args) session.send_tool_output(tool_call_id, result) session.on("/tool-call", on_tool_call) ``` ### Agent Updates Send mid-conversation updates to an active [agent](/verbs/verbs/agent): ```python # Change the system prompt session.update_agent({ "type": "update_instructions", "instructions": "You are now a billing agent.", }) # Inject context session.update_agent({ "type": "inject_context", "messages": [{"role": "user", "content": "Customer: Sarah, Gold tier."}], }) # Replace tools session.update_agent({"type": "update_tools", "tools": [...]}) # Trigger a new response session.update_agent({ "type": "generate_reply", "interrupt": True, "user_input": "Tell the customer about the flash sale.", }) ``` ### LLM Updates ```python session.update_llm({"instructions": "Switch to Spanish."}) ``` | Method | Description | | -------------------------------------- | ------------------------------------------------------- | | `send_tool_output(tool_call_id, data)` | Send tool/function result back to the LLM or agent verb | | `update_agent(data)` | Send an `agent:update` command | | `update_llm(data)` | Send an `llm:update` command | ## Inject Commands Inject commands execute immediately on an active call without affecting the verb stack: ```python # Mute/unmute session.inject_mute("mute") session.inject_mute("unmute") # Whisper to one party session.inject_whisper({"verb": "say", "text": "The customer is a VIP."}, agent_call_sid) # Control noise isolation session.inject_noise_isolation("enable", vendor="krisp", level=80) session.inject_noise_isolation("disable") # Control recording session.inject_record("startCallRecording", siprec_server_url="sip:siprec@recorder.example.com") session.inject_record("pauseCallRecording") # Send DTMF session.inject_dtmf("1234") # Redirect call flow session.inject_redirect("/new-webhook") # Tag the call with metadata session.inject_tag({"priority": "high", "department": "billing"}) ``` ## Session Properties | Property | Type | Description | | ----------------- | ------ | --------------------------------------------------------------- | | `call_sid` | `str` | Unique call identifier | | `from_number` | `str` | Caller phone number or SIP URI | | `to` | `str` | Called phone number or SIP URI | | `direction` | `str` | `'inbound'` or `'outbound'` | | `account_sid` | `str` | Account identifier | | `application_sid` | `str` | Application identifier | | `call_id` | `str` | SIP Call-ID | | `data` | `dict` | Full call session data (includes `env_vars`, SIP headers, etc.) | ## Examples See the [examples directory](https://github.com/jambonz/python-sdk/tree/main/examples) for runnable demos: | Example | Webhook | WebSocket | Description | | ------------- | ------- | --------- | ---------------------------- | | hello-world | yes | yes | Minimal greeting | | echo | yes | yes | Speech echo with gather | | ivr-menu | yes | — | IVR menu with speech + DTMF | | voice-agent | yes | yes | LLM pipeline with tool calls | | dial | yes | — | Outbound dial with fallback | | listen-record | yes | yes | Audio recording | > Experimental: Build jambonz voice applications with jambonz-python-sdk