Analyst: Google releases Gemini 3.5 speech transcription model, with accuracy ranking in the top five
2026-08-27 12:09:15
According to CoinMeta, Google has released the Gemini 3.5 speech transcription model, which is currently the most accurate speech transcription model offered by Google. For the first time, a smart mode has been incorporated into a dedicated transcription model. This model is available in both real-time and recorded transcription versions, with the real-time version now open for public testing. Based on evaluations, the WER (word error rate) for non-streaming transcription is 2.6%, while for the real-time version, it is 4.0%. The smart mode can automatically remove catchphrases and organize spoken text into paragraphs, lists, dates, and numbers. The model supports over 85 languages and allows for the addition of custom professional vocabulary. The cost of recorded transcription is approximately $0.005 per minute, and the real-time version costs about $0.009 per minute. This model has already been used in Google Translate's Rambler and macOS versions, and will be integrated into the Chrome in the future.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.