Google Unveils Gemini Embedding 2, Its First AI Model to Map Text, Images and Video Together

Gemini Embedding 2 understands text, images, and videos in the same language for easier retrieval.

Advertisement
Written by Akash Dutta, Edited by Ketan Pratap | Updated: 11 March 2026 11:14 IST
Highlights
  • Gemini Embedding 2 is available in public preview via API and Vertex AI
  • It also captures semantic intent across over 100 languages
  • The model can process up to six images per request

Gemini Embedding 2 can also understand interleaved input across multiple modalities

Photo Credit: Google

Google released its first fully multimodal embedding model on Tuesday. Dubbed Gemini Embedding 2, the artificial intelligence (AI) model maps text, images, audio, and videos into a single, unified embedding space. This means it uses an architecture to understand concepts whether they are written as words, spoken aloud, or shown in an image or a video. The Mountain View-based tech giant says this new system will simplify the way a large language model (LLM) understands information and will allow it to perform more complex actions.

Google's First Multimodal Embedding Model Is Here

In a blog post, the tech giant detailed the new AI model. It is the successor to the text-only embedding model that was released last year, and it captures semantic intent across more than 100 languages. Gemini Embedding 2 is currently available in public preview via the Gemini application programming interface (API) and Vertex AI.

Advertisement

AI models typically have different digital file cabinets to store text, photos, videos, and audio files. Whenever a user requests information in a specific format, it begins looking into that specific cabinet. Usually, an LLM treats a "cat" in a text document and a "cat" in a video as two completely different things. And to make matters more complex, the method to obtain information differs with each format.

Gemini Embedding 2 solves this problem by creating a new architecture that can only use a single cabinet for all kinds of information. This allows it to process a document that has both text and images at the same time, as humans do. Google says this new system simplifies “complex pipelines and enhances a wide variety of multimodal downstream tasks.” Some of these include Retrieval-Augmented Generation (RAG) and semantic search, sentiment analysis, and data clustering.

Advertisement

Coming to the AI model's capabilities, it has a text context window of up to 8,192 input tokens. It can also process up to six images per request in PNG and JPEG formats, and supports up to 120 seconds of video input in MP4 and MOV formats. Additionally, it can natively process and map audio data without needing text transcriptions. Further, it can also embed up to six-page-long PDFs.

The Gemini Embedding 2 can also understand interleaved input, so users can send across multiple modalities (such as text and image) in the same request. Google says this capability allows the model to gain a more accurate understanding of complex, real-world data.

 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Advertisement

Related Stories

Popular Mobile Brands
  1. Moto Pad 70 Set to Launch on This Date in India
  2. Oppo K15 With Dual 50-Megapixel Rear Cameras Arrives at This Price
  3. WhatsApp Introduces Redesigned Message Bubbles for iOS: Report
  4. Samsung Galaxy F70 Pro Appears on Geekbench Ahead of Anticipated Launch
  5. OnePlus N6x Confirmed to Launch in India With This Battery
  6. HMD Touch AI Reportedly Unveiled With AI Tools, 3.2-Inch Touchscreen
  7. iPhone 18 Pro Max Price Leak Hints at Apple's Costliest Non-Folding iPhone
  8. Xbox Insiders Can Now Stream Select Owned Games for Free With Ads
  1. Anthropic Expands Claude Voice Mode With Opus, Sonnet AI Models
  2. Samsung Galaxy F70 Pro Appears on Geekbench Ahead of Anticipated Launch
  3. Crypto Hacks Hit AFX and Verus Protocol Within Hours, Draining $31.6 Million
  4. OnePlus N6x Battery Capacity Revealed Week Before Launch in India
  5. Memory and Storage Chip Prices Tipped to Crash Next Year, But the Shortage Could Last Longer
  6. Moto Pad 70 India Launch Date Announced; Teased to Feature Dimensity 6400 SoC, 10,200mAh Battery
  7. HMD Touch AI Reportedly Launched in China With Doubao AI, 3.2-Inch Touchscreen
  8. Exynos 2700 Leak Suggests Samsung is Planning a Big Performance Upgrade
  9. Google Chrome Could Make Notification Prompts Less Annoying on Android
  10. Bitcoin Slips Below $66,000 as Middle East Tensions Pressure Crypto Market
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.