Google says that on a Pixel 11 Pro, EmbeddingGemma 2 requires 191MB of active RAM for text-only weights with quantisation.
Google EmbeddingGemma 2 is built on the Gemma 4 architecture
Photo Credit: Google
Google has announced EmbeddingGemma 2. This latest lightweight embedding model is designed to process text, code, images, audio, and video in a single, unified embedding space. The new model is built on the Gemma 4 architecture, and Google claims EmbeddingGemma 2 is optimal for on-device inference. EmbeddingGemma 2 has 740 million parameters. It is available on Hugging Face and Kaggle and will soon be available on the Gemini Enterprise Agent Platform Model Garden. Google has also released Nano Banana 2.1, an image generation and editing model. It is based on Gemini 3.6 Flash.
In a blog post, Google has detailed EmbeddingGemma 2. As mentioned, it is built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license. Google says this model has 740 million parameters and is suitable for on-device use. It will help users to find a specific video clip from a voice memo or search through audio recordings using a text query.
Google says the EmbeddingGemma 2 brings cross-modal search. It is said to offer a 9.92-point improvement in code performance and is suitable for local codebase indexing, semantic code search, and coding agent retrieval.
The new model is claimed to have leading scores among multimodal embedding models with fewer than one billion parameters across benchmarks including MTEB Code and MAEB. Google says it requires 270 million parameters for text-only workloads, with optional vision and audio encoders adding 170 million and 300 million parameters, respectively.
EmbeddingGemma 2 uses Matryoshka Representation Learning (MRL), allowing developers to reduce its 768-dimensional output vectors to 512, 256, or 128 dimensions. Google states that this will reduce storage requirements and memory needs for local vector databases by up to six times.
Google says the model is optimised for on-device performance, and the text-only version requires 191MB of active RAM on a Pixel 11 Pro, while the full multimodal model requires about 567MB. EmbeddingGemma 2 gets an 8K-token context window, which Google says is four times larger than that of EmbeddingGemma. This allows it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or combinations of these inputs on local hardware.
Google says EmbeddingGemma 2 is suited for local codebase indexing, semantic code search, and coding agent retrieval. When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines.
The EmbeddingGemma 2 weights are available through Hugging Face and Kaggle. Google confirmed that the Gemini Enterprise Agent Platform Model Garden availability is coming soon. It is deployable via cross-platform apps with Google AI Edge MediaPipe, LiteRT, transformers.js, and WebGPU. It is compatible with MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio.
Meanwhile, Google has also launched Nano Banana 2.1 as an update to Nano Banana 2 (Gemini 3.1 Flash Image). The new image generation and conversational editing model is claimed to deliver significant improvements in visual quality, prompt adherence, multi-turn character consistency, and text rendering. It has 1,31,072 input tokens and 32,768 output tokens.
Nano Banana 2.1 is based on Gemini 3.6 Flash. Google states that Nano Banana 2.1 maintain the speed and cost efficiency of Nano Banana 2. It provides improved visual quality and realism across 1K, 2K, and 4K output resolutions. It is also said to resolve tiling artefacts on wide and panoramic aspect ratios at 2K and 4K resolutions. It supports up to 14 reference images. It is now accessible in the Gemini App, Google AI Studio, Gemini API, Google Search AI Mode, Google Ads, Google Flow and Google Stitch.
Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.