AINewsnow

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.

Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports 4.0% wo…

Read the full story at MarkTechPost ↗

Timeline · 3 reports

  1. 2026-08-29 14:30 · MarkTechPost
    Google AI Releases Gemini Omni 1.1 Flash: 40-Second Scene Extension, First/Last Frame Control, and 4K Upscaling
  2. 2026-08-28 05:39 · r/machinelearningnews
    Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
  3. 2026-08-28 05:00 · MarkTechPost
    Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

More stories

  1. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  2. Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use — r/machinelearningnews
  3. AI skills — r/AI_Agents
  4. Gemini 4 Pro vs Fable 5 vs GPT6 Astra — r/GeminiAI
  5. AI models are not hacking “autonomously” — r/artificial
  6. Plugin4Shell and NIST IR 8587, days apart: what actually authorizes an AI agent’s action? — r/AI_Agents
  7. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  8. Gemini 2.5 pro model disappeared in AI Studio — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →