Updates·October 6, 2026, 19:36
Google launches EmbeddingGemma 2: Open AI model for text, images, audio and video
AI-generated and checked against the sources listed below.
Google has released EmbeddingGemma 2, an open model that can understand text, images, audio and video and run locally on a phone. The full version uses around 567 MB of RAM on a Pixel 11 Pro.

Google has launched EmbeddingGemma 2, an open AI model that can convert text, images, audio and video into numbers that a computer can compare. The model is small enough to run directly on a phone or an ordinary computer without data being sent to the internet.
What is an embedding?
An embedding is a kind of numerical fingerprint of a piece of content. If two things are similar in meaning, they get similar fingerprints. It is used for search that understands meaning rather than just words. It is also used for RAG, where an AI looks up relevant documents before it answers.
With EmbeddingGemma 2, text, images, audio and video are placed in the same space. In principle, you can therefore search images or audio clips with ordinary text.
Small and modular
The model is built on Google's Gemma 4 architecture. It comes in several sizes, and you only download what you need:
- 270 million parameters for text and code - 170 million extra for images - 300 million extra for audio - up to 740 million parameters with all parts
Parameters are the settings a model has learned during training. The fewer there are, the smaller and faster the model typically is.
On a Google Pixel 11 Pro, according to the sources I have seen, the text version uses about 191 MB of active RAM. The full multimodal version uses around 567 MB when compressed (quantized). That is little compared with many other AI models.
Longer context than its predecessor
The context window is 8,000 tokens, four times as much as in the first EmbeddingGemma. According to the coverage, that corresponds to up to 5.5 minutes of audio, 29 images or 58 video frames at a time.
What does it mean for you?
Because the model is open, developers and companies can download it for free and use it in their own apps. It can, for example, help search private photos, recordings and documents on the phone without sending them to a server. That is good for both privacy and speed.
The model can be downloaded from Hugging Face and Kaggle and works with MediaPipe, LiteRT, Ollama, LM Studio and llama.cpp, among others.
What it means for you
If you have many photos, recordings and documents on your phone, apps with the model may eventually make it easier to find them by typing what they are about. It happens locally, so your private files do not have to leave the phone. You will not use the model directly yourself, but if you are a developer or work with IT in a company, you can download it for free from Hugging Face or Kaggle and try it out in a tool like Ollama or LM Studio.
Sources
Get the week's AI news in your inbox
Choose your level, topics and length. One email a week, unsubscribe at any time.
Subscribe to Promptly Newsletter



