Eigen RadarAI
Analysis

Google's new transcription model cleans up speech while it listens

Gemini 3.5 Transcribe turns speech in more than 85 languages into formatted text while removing filler words and speaker corrections. Particle's Radar converts more than 130,000 podcasts into an index that agents can query. Amazon's book-scanning operation in Las Vegas shows the uncertainty that appears when physical publications are turned into a data source.

Artificial Intelligence··Morning
Conceptual scene of a microphone filtering messy coloured speech into clean, blank structured transcript segments.

Gemini 3.5 edits speech as it listens

The streaming and non-streaming versions of Gemini 3.5 Transcribe are generally available through the Gemini API. The model turns speech in more than 85 languages into formatted text, removing filler words and speaker corrections while labelling up to three speakers. Artificial Analysis measured an average word error rate of 4.0 per cent for streaming and 2.6 per cent for non-streaming. On FLEURS, the rates were 5.50 per cent and 5.04 per cent respectively. Google says time to a final transcript is 70 per cent shorter than with its previous model. Custom vocabulary accepts up to 1,000 terms, while support beyond three speakers remains experimental.[1]

Radar opens podcasts to agent queries

Particle's Radar transcribes podcast episodes with speaker labels and extracts highlights along with the people, companies, brands and products discussed. Its index covers more than 130,000 podcasts, including every show in Apple's Top 200 lists across 135 categories, and takes in 20,000 new episodes each day. A web search serves individual users, while the company's API and MCP server let AI agents query the material directly. Particle names journalists, researchers, data resellers, AI search platforms and hedge funds among its customers.[2]

Books are discarded after scanning at an Amazon facility

A worker interviewed by 404 Media said machines at Amazon's VGT3 facility in Las Vegas cut the spines from incoming books, roughly 20 to 25 high-speed scanners digitise the loose pages, and the paper is then discarded. The material included new sealed books, volumes from library liquidations, German, Russian and Japanese publications, and University of London documents. A tracking device that the outlet had previously placed in a rare-book shipment also reached the same facility. 404 Media could not identify which AI company would receive the scanned data, and Amazon did not comment for the report.[3]

References

  1. News sourceGoogleGoogle's new transcription model strips the filler words while it is still listening↩
  2. News sourceTechCrunchA podcast index built for agents opens 130,000 shows through an API and MCP↩
  3. News source404 MediaInside the Amazon warehouse that cuts up books to scan them for AI training↩