> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.jambonz.org/tutorials/voice-ai-examples/google-gemini-live/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.jambonz.org/_mcp/server. # Google Gemini Speech to Speech > **Note** > > The jambonz application referenced in this article can be > found [here](https://github.com/jambonz/gemini-s2s-example). This is an example jambonz application that connects to the Google Gemini Live API and illustrates how to build a Voice-AI application using jambonz and Google Gemini. The example covers: * a weather agent using Gemini [function calling](https://ai.google.dev/gemini-api/docs/function-calling) * an optional MCP server integration, so the same agent can expose its tools through the [Model Context Protocol](https://modelcontextprotocol.io/) ## Prerequisites * a [jambonz.cloud](https://jambonz.cloud/) account (or a self-hosted jambonz deployment on 10.1.0 or later) * a Google Cloud Platform account with the [Gemini API](https://ai.google.dev/) enabled * a carrier and virtual phone number of your choice ## Running instructions ### Set environment variables ```bash GOOGLE_API_KEY="your-gemini-api-key" PORT=3000 # Optional — only required when testing the MCP integration MCP_SERVER_URL=http://your-host:3001/sse ``` | Environment variable | Value | | :------------------- | :------------------------------------------------------------------------------------------------------------------------------------- | | `GOOGLE_API_KEY` | A Google API key with access to the Gemini Live API. You can create one in [Google AI Studio](https://aistudio.google.com/app/apikey). | | `PORT` | The port your Express server listens on. Defaults to `3000`. | | `MCP_SERVER_URL` | Optional. When set, the agent's tools are discovered from an MCP server instead of declared inline. | ### jambonz setup 1. [Create a carrier](/guides/using-the-jambonz-portal/basic-concepts/creating-carriers) entity in the jambonz portal. 2. [Add your speech provider](/guides/using-the-jambonz-portal/basic-concepts/creating-speech-credentials) of choice. Gemini Live handles speech-to-speech end to end, but jambonz still needs a speech credential configured on the account. 3. Create a new jambonz application under the [*Applications*](https://jambonz.cloud/internal/applications) tab. Point both the `Calling webhook` and `Call status webhook` at your server: ``` ws://your-example-domain.ngrok.io/google-s2s ``` 4. [Provision a phone number](/guides/using-the-jambonz-portal/basic-concepts/creating-phone-numbers) on your carrier and associate it with the application. ### Run the app ```bash npm install GOOGLE_API_KEY= npm start ``` To run with MCP tools, open two terminals: ```bash # Terminal 1 — MCP server MCP_SERVER_PORT=3001 npm run mcp-server # Terminal 2 — jambonz app GOOGLE_API_KEY= MCP_SERVER_URL='http://:3001/sse' node app.js ``` Call your virtual number and ask Barbara about the weather. ## How the `llm` verb is wired up The application calls `session.llm({...})` with `vendor: 'google'` and a Gemini Live model. The `llmOptions.setup` object is forwarded verbatim to Google's [BidiGenerateContentSetup](https://ai.google.dev/api/live#bidigeneratecontentsetup) message: ```js session.llm({ vendor: 'google', model: 'models/gemini-2.0-flash-live-001', auth: { apiKey: process.env.GOOGLE_API_KEY }, actionHook: '/final', eventHook: '/event', toolHook: '/toolCall', llmOptions: { setup: { generationConfig: { speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } } } }, systemInstruction: { parts: [{ text: 'You are a helpful agent named Barbara that can only provide weather information.' }] }, tools: [{ functionDeclarations: [{ name: 'get_weather', description: 'Get the weather for a location', parameters: { type: 'object', properties: { location: { type: 'string', description: 'The location to get the weather for' }, scale: { type: 'string', enum: ['celsius', 'fahrenheit'] } }, required: ['location'] } }] }] } } }); ``` See the full route in [lib/routes/weather-agent.js](https://github.com/jambonz/gemini-s2s-example/blob/main/lib/routes/weather-agent.js). ### Proactive greeting ("speak first") For outbound calls — or any scenario where you want Gemini to speak first — add a `greeting` to `llmOptions`. jambonz sends it immediately after setup so the caller hears the agent within the first second: ```js llmOptions: { setup: { /* ... */ }, greeting: 'Greet the caller warmly and ask how you can help.' } ``` The value is an **instruction to the model**, not the literal greeting text. Use `"Say exactly: Hello, thank you for calling Acme."` if you need a scripted line. > **Info** > > This also works on `models/gemini-3.1-flash-live-preview`. On the 3.1 preview, Google restricted `clientContent` to seeding history only, so jambonz uses `realtimeInput.text` under the hood — the `greeting` field is the portable way to trigger a first turn across all Gemini Live models. ### Session resumption Gemini Live sessions can be resumed across websocket reconnects. Opt in by passing `sessionResumption: {}` in `llmOptions`. Each `llm_event` hook delivers a `sessionResumptionUpdate` containing a fresh `newHandle` — store the latest handle, then reconnect with `sessionResumption: { handle: '' }` to continue the conversation. ## Function calling The `toolHook` fires when Gemini wants to call one of the declared functions. Respond with `session.sendToolOutput`: ```js session.sendToolOutput(tool_call_id, { toolResponse: { functionResponses: [ { id, response: { output: { temperature: 20, unit: 'celsius' } } } ] } }); ``` Gemini's native tool format uses `functionCalls` (inbound) and `functionResponses` (outbound) — jambonz passes them through without reshaping, so the payloads match the [Gemini Live tool use](https://ai.google.dev/gemini-api/docs/live-api/capabilities#tool-use) docs exactly. ## Interruption handling When the caller speaks over Gemini, the module emits `output_audio.playback_stopped` with `completion_reason: "interrupted"` on the event hook, and the queued audio is discarded so the caller hears their own voice, not stale agent audio. No application code is required — interruption handling is built in. #### A note on actionHook Like every jambonz verb, the `llm` verb fires `actionHook` when the session ends, including a `completion_reason`: * Normal conversation end * Connection failure * Disconnect from remote end * Server failure * Server error ## Resources * [Google Gemini Live API documentation](https://ai.google.dev/gemini-api/docs/live-api) * [BidiGenerateContent protocol reference](https://ai.google.dev/api/live) * [Example application source](https://github.com/jambonz/gemini-s2s-example) * jambonz documentation: * the [`llm`](/verbs/verbs/llm) verb * the [`dial`](/verbs/verbs/dial) verb * step-by-step [guides](/guides/telephony-integrations) for adding carriers to jambonz > Using jambonz to connect custom telephony to the Google Gemini Live API