EmbeddingGemma 2 brings images and sound into local file search
Google DeepMind released EmbeddingGemma 2 to represent text, code, images, audio and video in a shared search space. Applications can load only the components they need and retrieve local content by meaning. Downloadable weights and development examples support on-device use, while the model shares components with Gemma 4 for retrieval-assisted answers.
Artificial Intelligence··Midday
Text and media become searchable by meaning
Google DeepMind, Google’s AI research group, released EmbeddingGemma 2 on 6 October. The model converts text, source code, images, audio and video into a shared numerical representation, letting applications retrieve related content by meaning. It is built on Gemma 4 and makes its weights available under the Apache 2.0 licence for developers to run themselves.[1], [2]
Local execution is intended to let an application search files without sending their contents to an outside server. Google provides development examples for media libraries and for locating particular moments within video. The shared representation also accepts mixtures of input types, so a search system can work with text and media inside the same context.[1]
Applications can load only the encoders they need
The complete model contains 740 million parameters. Its text and code component uses 270 million, with optional encoders adding 170 million for vision and 300 million for audio. An application confined to text can use that component without loading the media encoders; applications handling images or sound can add the relevant part of the model.[1]
EmbeddingGemma 2 produces 768-dimensional vectors. Matryoshka Representation Learning, a method for retaining useful information in shorter representations, allows vectors of 512, 256 or 128 dimensions. Google says the shorter choice can reduce local vector-database storage by up to six times. These vectors are the numerical representations a retrieval system stores and compares when searching for relevant files.[1]
A larger context joins retrieval with Gemma 4
The input capacity is 8,192 tokens, four times that of the earlier EmbeddingGemma. Google’s examples translate that space into 29 images, 58 video frames or roughly five and a half minutes of audio. Tokens are the units a model processes; the different media examples describe ways to fill its available input capacity rather than separate storage quotas.[1]
The model shares its text tokenizer and audio encoder with Gemma 4. Developers can combine retrieval with answer generation using information from local files, and obtain the weights through Hugging Face and Kaggle. Google also says performance can vary across the more than 100 supported languages. The release includes no safety tuning or output moderation, with mitigation instead applied to training data.[1], [2]