> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.jambonz.org/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.jambonz.org/_mcp/server.

# Google Gemini Speech to Speech

> **Note**
>
> The jambonz application referenced in this article can be
> found [here](https://github.com/jambonz/gemini-s2s-example).

This is an example jambonz application that connects to the Google Gemini Live API and illustrates how to build a Voice-AI application using jambonz and Google Gemini.

The example covers:

* a weather agent using Gemini [function calling](https://ai.google.dev/gemini-api/docs/function-calling)
* an optional MCP server integration, so the same agent can expose its tools through the [Model Context Protocol](https://modelcontextprotocol.io/)

## Prerequisites

* a [jambonz.cloud](https://jambonz.cloud/) account (or a self-hosted jambonz deployment on 10.1.0 or later)
* a Google Cloud Platform account with the [Gemini API](https://ai.google.dev/) enabled
* a carrier and virtual phone number of your choice

## Running instructions

### Set environment variables

```bash
GOOGLE_API_KEY="your-gemini-api-key"
PORT=3000

# Optional — only required when testing the MCP integration
MCP_SERVER_URL=http://your-host:3001/sse
```

| Environment variable | Value                                                                                                                                  |
| :------------------- | :------------------------------------------------------------------------------------------------------------------------------------- |
| `GOOGLE_API_KEY`     | A Google API key with access to the Gemini Live API. You can create one in [Google AI Studio](https://aistudio.google.com/app/apikey). |
| `PORT`               | The port your Express server listens on. Defaults to `3000`.                                                                           |
| `MCP_SERVER_URL`     | Optional. When set, the agent's tools are discovered from an MCP server instead of declared inline.                                    |

### jambonz setup

1. [Create a carrier](/guides/using-the-jambonz-portal/basic-concepts/creating-carriers) entity in the jambonz portal.
2. [Add your speech provider](/guides/using-the-jambonz-portal/basic-concepts/creating-speech-credentials) of choice. Gemini Live handles speech-to-speech end to end, but jambonz still needs a speech credential configured on the account.
3. Create a new jambonz application under the [*Applications*](https://jambonz.cloud/internal/applications) tab. Point both the `Calling webhook` and `Call status webhook` at your server:
   ```
   ws://your-example-domain.ngrok.io/google-s2s
   ```
4. [Provision a phone number](/guides/using-the-jambonz-portal/basic-concepts/creating-phone-numbers) on your carrier and associate it with the application.

### Run the app

```bash
npm install
GOOGLE_API_KEY=<your key> npm start
```

To run with MCP tools, open two terminals:

```bash
# Terminal 1 — MCP server
MCP_SERVER_PORT=3001 npm run mcp-server

# Terminal 2 — jambonz app
GOOGLE_API_KEY=<your key> MCP_SERVER_URL='http://<your host>:3001/sse' node app.js
```

Call your virtual number and ask Barbara about the weather.

## How the `llm` verb is wired up

The application calls `session.llm({...})` with `vendor: 'google'` and a Gemini Live model. The `llmOptions.setup` object is forwarded verbatim to Google's [BidiGenerateContentSetup](https://ai.google.dev/api/live#bidigeneratecontentsetup) message:

```js
session.llm({
  vendor: 'google',
  model: 'models/gemini-2.0-flash-live-001',
  auth: { apiKey: process.env.GOOGLE_API_KEY },
  actionHook: '/final',
  eventHook: '/event',
  toolHook: '/toolCall',
  llmOptions: {
    setup: {
      generationConfig: {
        speechConfig: {
          voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Aoede' } }
        }
      },
      systemInstruction: {
        parts: [{ text: 'You are a helpful agent named Barbara that can only provide weather information.' }]
      },
      tools: [{
        functionDeclarations: [{
          name: 'get_weather',
          description: 'Get the weather for a location',
          parameters: {
            type: 'object',
            properties: {
              location: { type: 'string', description: 'The location to get the weather for' },
              scale: { type: 'string', enum: ['celsius', 'fahrenheit'] }
            },
            required: ['location']
          }
        }]
      }]
    }
  }
});
```

See the full route in [lib/routes/weather-agent.js](https://github.com/jambonz/gemini-s2s-example/blob/main/lib/routes/weather-agent.js).

### Proactive greeting ("speak first")

For outbound calls — or any scenario where you want Gemini to speak first — add a `greeting` to `llmOptions`. jambonz sends it immediately after setup so the caller hears the agent within the first second:

```js
llmOptions: {
  setup: { /* ... */ },
  greeting: 'Greet the caller warmly and ask how you can help.'
}
```

The value is an **instruction to the model**, not the literal greeting text. Use `"Say exactly: Hello, thank you for calling Acme."` if you need a scripted line.

> **Info**
>
> This also works on `models/gemini-3.1-flash-live-preview`. On the 3.1 preview, Google restricted `clientContent` to seeding history only, so jambonz uses `realtimeInput.text` under the hood — the `greeting` field is the portable way to trigger a first turn across all Gemini Live models.

### Session resumption

Gemini Live sessions can be resumed across websocket reconnects. Opt in by passing `sessionResumption: {}` in `llmOptions`. Each `llm_event` hook delivers a `sessionResumptionUpdate` containing a fresh `newHandle` — store the latest handle, then reconnect with `sessionResumption: { handle: '<stored handle>' }` to continue the conversation.

## Function calling

The `toolHook` fires when Gemini wants to call one of the declared functions. Respond with `session.sendToolOutput`:

```js
session.sendToolOutput(tool_call_id, {
  toolResponse: {
    functionResponses: [
      { id, response: { output: { temperature: 20, unit: 'celsius' } } }
    ]
  }
});
```

Gemini's native tool format uses `functionCalls` (inbound) and `functionResponses` (outbound) — jambonz passes them through without reshaping, so the payloads match the [Gemini Live tool use](https://ai.google.dev/gemini-api/docs/live-api/capabilities#tool-use) docs exactly.

## Interruption handling

When the caller speaks over Gemini, the module emits `output_audio.playback_stopped` with `completion_reason: "interrupted"` on the event hook, and the queued audio is discarded so the caller hears their own voice, not stale agent audio. No application code is required — interruption handling is built in.

#### A note on actionHook

Like every jambonz verb, the `llm` verb fires `actionHook` when the session ends, including a `completion_reason`:

* Normal conversation end
* Connection failure
* Disconnect from remote end
* Server failure
* Server error

## Resources

* [Google Gemini Live API documentation](https://ai.google.dev/gemini-api/docs/live-api)
* [BidiGenerateContent protocol reference](https://ai.google.dev/api/live)
* [Example application source](https://github.com/jambonz/gemini-s2s-example)
* jambonz documentation:
  * the [`llm`](/verbs/verbs/llm) verb
  * the [`dial`](/verbs/verbs/dial) verb
  * step-by-step [guides](/guides/telephony-integrations) for adding carriers to jambonz