ElevenCreative Review 2026: 10 Free Minutes, Then $6/Mo
ElevenCreative gives you 10,000 free credits a month, then $6 to $99. Studio, Flows, Music v2 and Dubbing v2 reviewed, plus where the credit system bites.
Read Article →
Gemini 3.5 Transcribe is the most precise speech-to-text model Google has shipped, by its own benchmarks and by independent testing, and it entered public preview on August 26, 2026. The model ships as two separate APIs: a real-time streaming option for live voice apps and a pre-recorded option that adds speaker labels and word-level timestamps for podcasts, meetings, and interviews.
Google built the new model around a smart transcription mode that catches mid-sentence corrections, drops filler words, and formats the output automatically, on top of the traditional word-for-word transcript most speech tools default to. Independent testing from Artificial Analysis puts word error rate at 2.6% for pre-recorded audio and 4.0% for live streaming, alongside a 70% cut in time-to-final-transcript compared with Chirp 3, the model it replaces. The rollout already reaches Gboard, Google Antigravity, AI Studio, and the Gemini app on macOS.
ElevenLabs already bundles production-grade text-to-speech, voice cloning, dubbing, and its own Scribe transcription model into one workspace, live today rather than in preview.
Try ElevenLabs Free →Gemini 3.5 Transcribe splits into two APIs built for different jobs. The Live API, gemini-3.5-transcribe-live, streams audio continuously in both directions and returns text with sub-second latency, for voice agents and other apps that can’t wait for a pause in speech. The Interactions API, gemini-3.5-transcribe, processes pre-recorded audio and adds speaker attribution with word-level timestamps, the format podcast editors and meeting notes need.
Both APIs share the same behavior around messy real speech. The default mode, verbatim, transcribes exactly what was said, filler words included. A separate smart mode resolves the kind of self-correction that shows up in ordinary conversation, like someone saying “let’s meet Tuesday, no, Wednesday,” strips out “ums” and “ahs,” and formats the result instead of leaving a flat wall of text.
Custom vocabulary is supported too. Feed it jargon, product names, or unusual spellings ahead of time and it adapts instead of guessing. The Interactions API can also identify separate speakers and timestamp each one.
Speaker identification is reliable for up to three people in the same recording. Google labels support for four or more speakers as experimental, so larger group recordings and panel discussions may need manual cleanup.
Independent testing backs up Google’s accuracy claim. Artificial Analysis measured a 2.6% word error rate on pre-recorded audio through the Interactions API and 4.0% on live streaming through the Live API, with the model holding up well in noisy environments and correctly capturing alphanumeric strings like postal codes and order IDs.
On the FLEURS multilingual benchmark, which tests accuracy across a wide language set, the model scored 5.04% word error rate for non-streaming audio and 5.50% for streaming.
Chirp 3, Google’s prior transcription model, is the baseline for that 70% speed figure. Google says the newer model brings new capabilities on top of a lower word error rate, but the clearest practical difference is speed: Artificial Analysis clocked a 70% improvement in time-to-final-transcript against Chirp 3, which matters most in the Live API’s real-time use cases, where lag is felt immediately.
Availability splits by audience. Developers can test Gemini 3.5 Transcribe now through the Gemini API in Google AI Studio and inside Google Antigravity. Enterprises get it in public preview through the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience listed as coming soon. Consumers already have it running inside four products:
Android keyboard feature, live in select countries and languages, that turns rambling speech into cleanly formatted text, with voice edits to what's already transcribed
Uses screen context and, with permission, chat history to improve accuracy on details like file names
Voice-driven vibe-coding, dictating changes to a project instead of typing them
English-only for now, with voice commands that read screen context to summarize local files or generate an image at the cursor
Chrome support is coming soon, for talking instead of typing into any web form or text field. On macOS, the model can also hand a task to a different Gemini model through function calling: asking it to generate an image or analyze a file mid-conversation routes the request automatically, with Gemini 3.5 Transcribe acting as the interface rather than doing that work itself.
For a working creator, the practical uses line up with tools already in daily rotation: turning a raw podcast recording into a searchable transcript, drafting captions before a video export, dictating a rough script instead of typing one, or feeding a clean transcript into one of the best AI text-to-speech tools to produce a dubbed voiceover in another language. Developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents already support it, and early users like Vivo, Intellitek Health, and Lingopal cited latency, accuracy, and language coverage in feedback Google published at launch.
Google didn’t give the August 26 launch a standalone product page or a price list. It shipped the model as a layer inside things people already use: a keyboard, a coding tool, an operating system’s built-in assistant. That’s a different pitch than the one dedicated voice AI platforms have been making.
ElevenLabs built its business the other way, as a platform where transcription is one feature among several rather than a background utility folded into other apps. Its own speech-to-text model, Scribe, sits inside a workspace that also handles text-to-speech, voice cloning, and dubbing, aimed at people who want one tool built specifically for audio production. A developer wiring a voice agent together needs something different than a video editor who just wants a finished interface, not an API to learn.
What Gemini 3.5 Transcribe changes is the baseline. When a model this accurate ships free inside a keyboard app, decent transcription stops being something creators pay a separate subscription for by default, and platforms competing for that budget need a clearer reason than accuracy alone. Google has moved fast on Gemini features this quarter, from a watermark toggle for AI-generated media to this transcription launch. Same pattern both times: capability shipped as infrastructure, not sold as a product on its own. It’s worth watching alongside our AI news feed.
A clean transcript is the first step in a dubbing pipeline. ElevenLabs turns that script into voiceovers across dozens of languages, with voice cloning that keeps the original speaker's tone intact.
Start Dubbing with ElevenLabs →Google has not published public-preview pricing for Gemini 3.5 Transcribe. Developers can test it now through the Gemini API in Google AI Studio and Google Antigravity, and consumer access through the Gemini app on macOS and Rambler on Gboard doesn't carry a separate charge, but the announcement lists no API rates for the preview.
Gemini 3.5 Transcribe measures a 2.6% word error rate on pre-recorded audio and 4.0% on live streaming audio, according to independent benchmarks from Artificial Analysis. On the FLEURS multilingual benchmark across top languages, word error rate comes in at 5.04% for non-streaming and 5.50% for streaming, and Google says the model holds up well in noisy environments and captures alphanumeric strings like postal codes and order IDs accurately.
Gemini 3.5 Transcribe supports more than 85 languages with automatic language detection built in. The model also handles regional accents and can switch languages mid-stream during a live conversation without being told which language is coming next.
Rambler is a Gboard feature for Android that uses Gemini 3.5 Transcribe to turn rambling, unstructured speech into cleanly formatted text. It's built for dictation that doesn't come out as a tidy sentence on the first try, and it supports voice-based edits to text that's already been transcribed instead of retyping by hand.
Gemini 3.5 Transcribe replaces Chirp 3 as Google's transcription model with new capabilities, a lower word error rate, and meaningfully faster response times. Artificial Analysis measured a 70% improvement in time-to-final-transcript compared with Chirp 3, alongside smart-transcription features like filler-word removal and self-correction handling that Chirp 3 didn't offer.
Gemini 3.5 Transcribe suits people already working inside Google's surfaces and developers wiring the API into their own apps, since transcription comes bundled with tools like Gboard, Antigravity, and AI Studio. A dedicated platform fits production audio work better: ElevenLabs pairs its own transcription model, Scribe, with text-to-speech, voice cloning, and dubbing in one workspace, built for teams whose output is finished audio rather than a raw transcript.