# Gemini 3.5 Transcribe API: Recorded Audio, Live Captions, and a Safe First Build

> Use Files and Interactions for recordings, or the Live API for raw PCM captions. Vocabulary biasing and recorded speaker or word annotations require separate configurations.

- URL: https://blog.laozhang.ai/en/posts/gemini-3-5-transcribe-api
- Published: 2026-08-27
- Updated: 2026-10-07
- Author: LaoZhang AI Team (https://blog.laozhang.ai/en/about)
- Topic: API Guides
- Tags: Gemini 3.5 Transcribe, Gemini API, speech-to-text, Live API, audio transcription

---
Use `gemini-3.5-transcribe` with the Files API and Interactions API for a recording such as an MP3. Use `gemini-3.5-transcribe-live` with the Live API when captions must update as speech arrives. The recorded route can return speaker labels and word offsets; the live route returns interim and finalized text without those annotations. These are the dedicated model IDs in the current [model documentation](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe).

The first implementation decision is what your application must retain. A readable transcript, a transcript linked to playback, and a changing caption preview need different configurations and storage. In particular, **recorded custom vocabulary cannot be combined with speaker diarization or word timestamps**. The examples below keep those requests separate.

The API examples follow the current documentation. We checked Python syntax and the application helpers offline, including synthetic PCM and transcription events; we did not execute SDK clients, upload audio, call the models, or measure recognition quality. A service response and a reviewed audio sample remain necessary before production.

## Start with the output, then choose the route

| Product requirement | Recorded transcription | Live captions |
|---|---|---|
| Model ID | `gemini-3.5-transcribe` | `gemini-3.5-transcribe-live` |
| Input path | Upload a file, then pass its URI to Interactions | Send raw audio over a persistent Live connection |
| Maximum duration | 1 hour; 30 minutes with diarization or word timestamps | 10-minute session |
| Speaker labels and word offsets | Available in a compatible verbatim configuration | Not available |
| Application text state | Save the response and extract the required annotations | Replace the interim preview; append finalized segments |

These limits come from the [dedicated model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe). Split longer recordings before submitting them, retaining each chunk's original start time if playback needs a common timeline. A ten-minute Live limit requires an application plan for the next session; a longer-lived credential does not extend it.

![Recorded and live transcription routes with their input and output differences](https://blog.laozhang.ai/posts/en/gemini-3-5-transcribe-api/img/endpoint-decision.webp)

Choose recorded transcription first for interviews, podcasts, meeting archives, or uploaded voice notes. Live is useful when seeing words immediately matters, such as captions during a call. A finished meeting summary still needs a separate processing step. These dedicated STT models do not provide native function calling, answer arbitrary audio questions, or speak back as a Live Agent would.

Google's release labels are inconsistent as checked on October 6: the [August 26 changelog](https://ai.google.dev/gemini-api/docs/changelog#august-26-2026) calls both models generally available, while the [launch announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/) still describes public preview. Use the documented identifiers and verify access in your project; neither label establishes availability for every account.

## Verify one recorded file before adding features

The [recorded transcription guide](https://ai.google.dev/gemini-api/docs/transcribe) uses an audio item directly in `input`, with the uploaded file's URI and MIME type. This is an Interactions request, not a normal `generateContent` prompt.

For the reader-facing examples, use Python with the Google Gen AI SDK installed and an API key already configured on a trusted machine. [Google's key guide](https://ai.google.dev/gemini-api/docs/api-key) explains creating and configuring a key in AI Studio. This complete recorded example takes a local MP3 filename, prints the merged text, and saves the full response:

```python
import json
import sys
from pathlib import Path
from google import genai

client = genai.Client()
uploaded = client.files.upload(file=sys.argv[1])
interaction = client.interactions.create(
    model="gemini-3.5-transcribe",
    input=[{
        "type": "audio",
        "uri": uploaded.uri,
        "mime_type": uploaded.mime_type,
    }],
)
payload = interaction.model_dump(mode="json", exclude_none=True)
Path("recorded-response.json").write_text(
    json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
)
print(interaction.output_text)
```

Save it as `recorded.py`; its invocation is `python recorded.py sample.mp3`. Upload only audio you are entitled to process under the account's applicable data terms. Other supported recorded formats include WAV, FLAC, OGG, and WebM; take the MIME type from the actual upload rather than guessing from a renamed extension.

For REST clients, the corresponding method is `POST https://generativelanguage.googleapis.com/v1beta/interactions`, authenticated with the `x-goog-api-key` header. After uploading through Files, send the same `model` and flat audio `input` fields shown above. The returned file URI is the input reference, not a public download link.

A useful first success check is a completed interaction, nonempty `output_text`, and words that correspond to the supplied recording. Retain the raw response, model, input duration, MIME type, request status, and available usage information. Nonempty text alone does not establish accuracy or prove that speakers and timestamps were produced.

## Clean prose and traceable transcripts require different modes

Recorded transcription defaults to verbatim. Use that mode when false starts, repeated words, and self-corrections matter. For cleaned dictation, recorded smart mode is the **string** `"smart"`; its configuration is `{"mode": "smart"}`. It is not `{"mode": {"type": "smart"}}`, and it cannot provide speaker diarization or word timestamps. The configuration belongs under `generation_config.transcription_config` in the [recorded guide](https://ai.google.dev/gemini-api/docs/transcribe).

### Choose vocabulary biasing or structured annotations

For names and technical terms, use a vocabulary configuration without the annotation options:

```python
vocabulary_config = {
    "language_codes": ["en-US"],
    "custom_vocabulary": ["BigQuery", "Kubernetes", "Acme Q7"],
}
```

For playback, subtitles, or speaker attribution, use a verbatim **object**, with no custom vocabulary:

```python
annotated_config = {
    "mode": {
        "type": "verbatim",
        "diarization_mode": "speaker",
        "timestamp_granularities": ["word"],
    },
}
```

Both dictionaries are alternatives for this replacement call in `recorded.py`, after `uploaded` has been created:

```python
chosen_config = annotated_config  # Or vocabulary_config; never merge them.
interaction = client.interactions.create(
    model="gemini-3.5-transcribe",
    input=[{
        "type": "audio",
        "uri": uploaded.uri,
        "mime_type": uploaded.mime_type,
    }],
    generation_config={"transcription_config": chosen_config},
)
```

If language is unknown or the recording switches languages, omit `language_codes` or use an empty list for automatic detection. Vocabulary lists allow up to 1,000 entries, with Google recommending up to 100 for typical best results. That is a recognition hint, not a guarantee that every listed name will be correct. Diarization is documented for up to eight speakers, but attribution with **three or more is experimental**; word timestamps can reduce transcription accuracy. Those qualifications and the vocabulary incompatibility are on the [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe).

When you need annotations and polished prose, preserve the annotated verbatim response first, then create a separate cleaned view downstream. Do not overwrite the words associated with offsets and imply they still align exactly to the audio. Two separate model calls also produce separate results; vocabulary-biased text cannot simply inherit another call's timing.

### Preserve WordInfo before generating subtitles

`output_text` is convenient for display, but the documented structured path is `steps[].content[].annotations[]`. Filter annotations with `type == "word_info"`; their fields include `text`, `speaker`, `start_offset`, and `end_offset`. Offsets are duration strings such as `"0.100s"`, not already numeric seconds. Saving the complete response protects information your first renderer may not use. See the [structured output examples](https://ai.google.dev/gemini-api/docs/transcribe).

![The full response retains word_info annotations with speaker and duration offsets beside merged display text.](https://blog.laozhang.ai/posts/en/gemini-3-5-transcribe-api/img/output-anatomy.webp)

This standalone parser reads the JSON saved by the recorded example. It preserves each original word annotation and adds numeric offsets for application use:

```python
import json
import re
import sys
from decimal import Decimal
from pathlib import Path

def seconds(value):
    if value is None:
        return None
    if not isinstance(value, str) or not re.fullmatch(r"\d+(?:\.\d{1,9})?s", value):
        raise ValueError(f"Invalid nonnegative duration: {value!r}")
    return Decimal(value[:-1])

def extract_words(payload):
    words = []
    for step in payload.get("steps", []):
        for content in step.get("content", []):
            for annotation in content.get("annotations", []):
                if annotation.get("type") != "word_info":
                    continue
                start = seconds(annotation.get("start_offset"))
                end = seconds(annotation.get("end_offset"))
                if start is not None and end is not None and end < start:
                    raise ValueError("Word ends before it starts")
                words.append({
                    "annotation": dict(annotation),
                    "start_seconds": start,
                    "end_seconds": end,
                })
    return words

if __name__ == "__main__":
    payload = json.loads(Path(sys.argv[1]).read_text(encoding="utf-8"))
    print(json.dumps(extract_words(payload), ensure_ascii=False, default=str))
```

An offset of zero is valid. Missing offsets remain `None`, and missing speaker labels stay missing; the parser does not invent them. An empty word list means no matching annotations were found, so inspect the request and response before building subtitles. To create SRT or VTT, group the actual timed words into readable cues and convert their seconds into that format's time notation. Add the original chunk start only when the offsets belong to a split recording; account for overlapping speakers rather than forcing all words onto one sequential speaker track.

## Live transcription is a media state machine

Live expects **raw signed 16-bit PCM, 16 kHz, mono, little-endian**. At this rate, 100 ms is 1,600 samples or 3,200 bytes. The [Live transcription guide](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe) recommends 100 ms chunks and `audio/pcm;rate=16000`. An MP3, a WAV header, or WebM/Opus microphone output cannot become PCM by changing its MIME label: decode it and, when needed, resample and downmix before sending.

Keep two text states. Each `interim_input_transcription` replaces the current preview. Each `input_transcription` appends a finalized segment and clears that preview. Process both fields if an event contains both, and ignore events without `server_content`. Do not globally deduplicate equal final text: a speaker may legitimately say the same phrase twice.

The following SDK example streams a local, already compatible WAV at its natural pace. It removes the container by reading frames, implements caption storage, and bounds the connection, sends, and EOF drain. It uses default automatic VAD. The eight-second drain is an **application deadline**, not a documented service completion promise. The JSON output preserves received text and keeps `complete=False`: a deadline, an error, or the end of a receive iterator does not prove that the whole input was finalized.

```python
import asyncio
import contextlib
import json
import sys
import wave
from dataclasses import asdict, dataclass, field
from pathlib import Path
from google import genai
from google.genai import types

@dataclass
class Captions:
    preview: str = ""
    finals: list[str] = field(default_factory=list)

    def apply(self, content):
        if content is None:
            return
        interim = getattr(content, "interim_input_transcription", None)
        final = getattr(content, "input_transcription", None)
        if interim is not None and interim.text is not None:
            self.preview = interim.text
        if final is not None and final.text is not None:
            self.finals.append(final.text)
            self.preview = ""

def validate_wav(audio):
    if (audio.getnchannels(), audio.getsampwidth(), audio.getframerate(),
        audio.getcomptype()) != (1, 2, 16000, "NONE"):
        raise ValueError("Use uncompressed mono 16-bit 16 kHz WAV")
    duration = audio.getnframes() / 16000
    if not 0 < duration <= 540:
        raise ValueError("This example accepts up to nine minutes of audio")
    return duration

async def receive(session, captions):
    async for response in session.receive():
        captions.apply(getattr(response, "server_content", None))
    return "receive_iterator_ended"

async def stream_file(client, config, audio, captions):
    async with client.aio.live.connect(
        model="gemini-3.5-transcribe-live", config=config
    ) as session:
        receiver = asyncio.create_task(receive(session, captions))
        try:
            while chunk := audio.readframes(1600):
                if receiver.done():
                    await receiver  # Surface receive errors immediately.
                    raise RuntimeError("Receive stream ended before audio EOF")
                if len(chunk) % 2:
                    raise ValueError("Incomplete 16-bit PCM sample")
                await asyncio.wait_for(session.send_realtime_input(
                    audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")
                ), timeout=5)
                await asyncio.sleep(len(chunk) / 32000)
            await asyncio.wait_for(
                session.send_realtime_input(audio_stream_end=True), timeout=5
            )
            try:
                return await asyncio.wait_for(asyncio.shield(receiver), timeout=8)
            except asyncio.TimeoutError:
                return "eof_drain_deadline"
        finally:
            receiver.cancel()
            with contextlib.suppress(asyncio.CancelledError):
                await receiver

async def main(filename):
    captions = Captions()
    result = {"status": "not_started", "complete": False}
    try:
        with wave.open(filename, "rb") as audio:
            duration = validate_wav(audio)
            client = genai.Client()
            config = types.LiveConnectConfig(
                response_modalities=["TEXT"],
                input_audio_transcription=types.AudioTranscriptionConfig(
                    language_codes=[], mode="VERBATIM"
                ),
            )
            result["status"] = await asyncio.wait_for(
                stream_file(client, config, audio, captions), timeout=duration + 30
            )
    except Exception as error:
        result.update(status="error", error_type=type(error).__name__)
    finally:
        result.update(asdict(captions))
        Path("live-captions.json").write_text(
            json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8"
        )
        print(result["status"], "final segments:", len(captions.finals))

if __name__ == "__main__":
    asyncio.run(main(sys.argv[1]))
```

Save this as `live_file.py`; the invocation is `python live_file.py sample-16k-mono.wav`. The nine-minute input cap leaves headroom inside the model's ten-minute limit. This intentionally finite example is a starting point for evaluating a stream, not a microphone capture or a reconnection implementation. Inspect `finals`, `preview`, and `status` separately. A final segment means the model finalized that segment, not that the whole recording has been human-verified or that every last sample was finalized before the deadline.

Live modes use uppercase enums, `"VERBATIM"` and `"SMART"`, inside `input_audio_transcription`. Do not copy the lowercase recorded mode or its verbatim object into this configuration. For push-to-talk, the documented manual VAD path disables automatic activity detection and sends `activity_start` and `activity_end`; implement that alternative deliberately instead of mixing activity signals into the default path above.

For production, assign your own session and segment sequence numbers, retain the source audio position, and decide how users see a disconnect or incomplete ending. Those are application identifiers. Equal text across a reconnect cannot prove whether an event is a duplicate, and this example offers no cross-session exactly-once guarantee.

### Keep browser credentials separate from session limits

A browser or mobile client connecting directly should obtain a constrained ephemeral token from an authenticated backend, rather than carry a permanent API key. The current [ephemeral-token guide](https://ai.google.dev/gemini-api/docs/live-api/ephemeral-tokens) limits this mechanism to the Live API on `v1beta`. Its defaults allow one use, one minute to start a new session, and 30 minutes for connection messaging. Those credential clocks do not override the Transcribe Live model's ten-minute session limit. Bind the token to the exact transcription model and text configuration, and handle renewal at your application's authenticated backend.

## Limits and price are architecture inputs

The [Gemini Developer API pricing page](https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-transcribe), checked October 6, lists these paid rates:

| Dedicated route | Audio input per 1 million tokens | Text output per 1 million tokens | Google's rounded combined estimate |
|---|---|---|---|
| Recorded Transcribe | $2.00 | $12.00 | About $0.005/minute |
| Transcribe Live | $3.50 | $21.00 | About $0.009/minute |

The estimate assumes **25 audio tokens per second and 175 output text tokens per minute**, not a generic audio-model token rate. Under those assumptions, one recorded minute costs `1,500 × $2 / 1,000,000 + 175 × $12 / 1,000,000 = $0.0051`; one live minute costs `$0.008925`. Thus 100 hours, or 6,000 minutes, is **$30.60 recorded or $53.55 live** before other costs. Using Google's rounded minute estimates instead gives roughly $30 and $54. These are calculations, not invoices or a fixed per-minute tariff.

For actual budgeting, calculate audio input tokens times the input rate plus output text tokens times the output rate, divided by one million. Include retries, media decoding, storage, streaming infrastructure, downstream summaries, and human correction. The model page does not offer a Batch API path for these dedicated models, so do not assume a generic batch discount.

### Free pricing and file expiry do not settle data handling

The pricing table includes free-of-charge input and output rows. Quotas, project access, and applicable service terms still matter. The [Gemini API terms](https://ai.google.dev/gemini-api/terms) say general unpaid-service content can be used for product improvement and human review, and warn against submitting sensitive or confidential information. Paid Services have different data-use rules, including the active Cloud Billing project condition, while limited security, abuse, and legal processing can still occur. The terms also apply Paid Services data-use treatment in the EEA, Switzerland, and the UK, and require Paid Services for API clients serving users there. Language choice does not establish a user's region or eligibility.

The [Files API guide](https://ai.google.dev/gemini-api/docs/files) documents automatic file deletion after 48 hours, a 2 GB per-file limit, and 20 GB per project. Keep your own original audio when needed: Files does not let you download the uploaded content back. File expiry is the lifecycle of that file resource, not a promise that interaction outputs or all provider logs share the same retention period.

## Evaluate the fields that can hurt the business

Before release, compare real output against a reviewed transcript from your actual workload. Include accents, background noise, fast speech, overlapping speakers, code-switching, and your production codecs. Check names, order IDs, amounts, dates, and domain terms separately: a low overall word error rate can still hide a costly entity mistake.

For annotated recordings, inspect speaker attribution and whether clicking a word seeks to the right audio. For Live, measure time to the first partial, preview revisions, final latency, incomplete endings, and reconnect behavior. A parser passing synthetic fixtures confirms how your code treats fields; it does not show that the model recognizes the right words or produces accurate offsets.

Google's [August 26 launch report](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/) cites Artificial Analysis WER figures of 4.0% for streaming and 2.6% for recorded transcription. Treat these reported benchmarks as context for your evaluation, not a result for your microphones, speakers, vocabulary, or deployment.

## Diagnose the contract before changing prompts

| Symptom | First check | Next action |
|---|---|---|
| Vocabulary plus speakers or timestamps is rejected | Recorded configuration combines incompatible options | Choose the vocabulary-only or annotated-verbatim branch |
| Text exists but speakers or word offsets are missing | You saved only `output_text`, used smart mode, or called Live | Use recorded annotated verbatim and inspect `word_info` annotations |
| Recorded smart mode fails schema validation | `mode` was sent as a smart object | Send the lowercase string `"smart"` |
| Live text repeats during speech | Every interim update was appended | Replace the preview; append final segments without global text deduplication |
| Live audio fails or produces implausible text | Bytes are a compressed container or the wrong PCM format | Decode and check channel count, rate, sample width, byte order, and chunking |
| End-of-file hangs or loses the last text | No bounded receive policy after `audio_stream_end` | Drain with a deadline, save partial state, and report incomplete output |
| A transcript cannot summarize or call tools | Dedicated STT is being used as an agent | Pass the transcript to a separate suitable processing step |

Start with one eligible, non-sensitive recorded sample. Add the compatible features your product needs, preserve the full response, then implement Live separately if immediate captions justify it. Promote the workflow only after real audio, failure handling, cost, and data-use checks meet your application's acceptance criteria.

## Frequently asked questions

### Can Gemini 3.5 Transcribe read MP3 files?

Yes, the recorded route supports MP3: upload the file through Files and pass its URI and MIME type to the Interactions API. Live takes raw PCM rather than MP3 bytes, so decode the audio first if you are using that route. See the [recorded input formats](https://ai.google.dev/gemini-api/docs/transcribe) and [Live audio format](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe).

### Can I get live captions with speaker names and word timestamps?

The current dedicated Live model does not return diarization or word timestamps. Use recorded `gemini-3.5-transcribe` with annotated verbatim mode for those fields, within its 30-minute structured-input limit. Diarized labels also do not automatically identify a person's real name. The [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe) documents these feature boundaries.

### Is Gemini 3.5 Transcribe free?

The Developer API pricing page lists free input and output rows, alongside paid token rates; that does not promise unlimited use or universal project access. For a paid estimate, Google's rounded combined figures are about $0.005 per recorded minute and $0.009 per live minute under its published token assumptions. Check [current pricing](https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-transcribe) and the applicable [data-use terms](https://ai.google.dev/gemini-api/terms) before sending audio.

### Should I use a model name ending in `-preview`?

Use the exact Developer API identifiers documented for this integration: `gemini-3.5-transcribe` and `gemini-3.5-transcribe-live`. A similarly named model in another platform or a search result is not evidence of a Developer API alias. Confirm your model and route against the [current model documentation](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe).

## Sources

External pages this guide links to, in the order they appear. Last updated 2026-10-07.

- [model documentation](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe) (ai.google.dev)
- [August 26 changelog](https://ai.google.dev/gemini-api/docs/changelog) (ai.google.dev)
- [launch announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/) (blog.google)
- [recorded transcription guide](https://ai.google.dev/gemini-api/docs/transcribe) (ai.google.dev)
- [Google's key guide](https://ai.google.dev/gemini-api/docs/api-key) (ai.google.dev)
- [Live transcription guide](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe) (ai.google.dev)
- [ephemeral-token guide](https://ai.google.dev/gemini-api/docs/live-api/ephemeral-tokens) (ai.google.dev)
- [Gemini Developer API pricing page](https://ai.google.dev/gemini-api/docs/pricing) (ai.google.dev)
- [Gemini API terms](https://ai.google.dev/gemini-api/terms) (ai.google.dev)
- [Files API guide](https://ai.google.dev/gemini-api/docs/files) (ai.google.dev)
