Elon Musk’s xAI Unveils Grok 1.5 Vision AI Model in Preview, To Compete With GPT-4 Vision and Gemini Pro 1.5

xAI said Grok 1.5 Vision can process a wide variety of visual information, including documents, diagrams, charts, and more.

Advertisement
Written by Akash Dutta, Edited by Manas Mitul | Updated: 15 April 2024 14:09 IST
Highlights
  • This is xAI’s first-generation multimodal AI model
  • Grok 1.5 Vision has a context length of 1,28,000 tokens
  • Recently, Grok AI was made available in open-source

xAI said Grok 1.5 Vision outperforms existing AI models in the company’s new RealWorldQA benchmark

Photo Credit: xAI

Elon Musk's artificial intelligence (AI) firm xAI has unveiled a new AI model dubbed Grok 1.5 Vision. This large language model (LLM) is an enhanced version of the recently released Grok 1.5 model. With this upgrade, the AI model is now equipped with computer vision, making it capable of accepting visual media as input. It can process images and answer questions about it. Notably, the announcement came just days after OpenAI introduced its own computer vision-powered GPT-4 model.

The announcement was made by the official X (formerly known as Twitter) account of xAI. The firm shared a blog post detailing the new AI model and shared some of its benchmark scores. Since the vision capabilities were added to the recently unveiled Grok 1.5 model, most of the details remain the same. It has the same context window of 1,28,000 tokens and the general benchmark scores are also likely to remain the same.

xAI also shared benchmark scores of Grok 1.5 Vision tested on a benchmark developed by the company. The AI firm calls it the RealWorldQA benchmark and it measures “real-world spatial understanding”. It also tested the model in several other benchmarks such as MMMU, Mathvista, ChartQA, and more. While Grok outperformed OpenAI's GPT-4 with Vision and Gemini 1.5 Pro in RealWorldQA, it scored less in MMMU and ChartQA.

Advertisement

For the unversed, computer vision is a branch of computer science that deals with equipping computers (and AI models) with the ability to identify and understand objects in the real world using images and videos. This is designed to help computers see and process visual signals the way humans do. With the rise of multimodal AI models, many firms are now focusing on developing vision-focused models. Google's Gemini 1.5 Pro and OpenAI's GPT-4 with Vision both have this capability.

Advertisement

This technology also offers a wide range of applications. The Indian calorie tracking and nutrition feedback platform Healthify recently added a feature called Snap where users can click a picture of a food item or cuisine, and GPT-4 with Vision-powered AI chatbot suggests how the recipe can be made healthier, and how much exercise one needs to do to burn the extra calories. In future, AI models with computer vision can assist in the diagnosis of diseases, building self-driving cars, and more.


Is the Samsung Galaxy Z Flip 5 the best foldable phone you can buy in India right now? We discuss the company's new clamshell-style foldable handset on the latest episode of Orbital, the Gadgets 360 podcast. Orbital is available on Spotify, Gaana, JioSaavn, Google Podcasts, Apple Podcasts, Amazon Music and wherever you get your podcasts.
Affiliate links may be automatically generated - see our ethics statement for details.
 

Get your daily dose of tech news, reviews, and insights, in under 80 characters on Gadgets 360 Turbo. Connect with fellow tech lovers on our Forum. Follow us on X, Facebook, WhatsApp, Threads and Google News for instant updates. Catch all the action on our YouTube channel.

Further reading: xAI, Elon Musk, Grok, X, Artificial intelligence, AI
Advertisement

Related Stories

Popular Mobile Brands
  1. Google Expands Android Theft Protection With New Security Features
  2. WhatsApp's Claims Its New Feature Can Protect Against Cyberattacks
  3. Amazfit Active Max With 1.5-Inch AMOLED Display Launched in India: See Price
  4. Vivo X200T Launched in India With These Features
  5. Nothing Phone 4a Pro's  Battery, Durability, Charging Details Revealed
  6. HP HyperX Omen 15 Gaming Laptop With RTX 5060 GPU Launched in India
  7. Border 2 Revives "Sandese Aate Hain": Sunny Deol Returns
  8. Xiaomi 17, Xiaomi 17 Ultra Global Variants' RAM, Storage and Colours Leaked
  9. Here's How Much the iQOO 15R Might Cost in India
  10. The Conjuring: Last Rites OTT Release Date: When and Where to Watch it Online?
  1. Samsung Galaxy Z TriFold Goes on Sale in the US Starting January 30: Price, Specifications
  2. Wonder Man Now Available for Streaming Online: Where to Watch it Online?
  3. Gulab Jamun Streaming Now on Manorama Max: Know Everything About This Malayalam Drama Film
  4. The Rabbit House OTT Release Date: When and Where to Watch This Chilling Mystery Online?
  5. The Wheel of Fortune Now Available for Streaming on SonyLIV: What You Need to Know About Akshay Kumar-Starrer Game Show
  6. Apple Could Pay More for iPhone RAM as Samsung, SK Hynix Said to Increase Prices by Up to 100 Percent
  7. Xiaomi 17, Xiaomi 17 Ultra RAM, Storage and Colourways Leaked as Company Gears Up for Global Launch
  8. Samsung Galaxy A57 Design Spotted in Leaked Renders; Might Feature Triple Rear Camera Setup
  9. Google Expands Android Theft Protection With New Security Features
  10. WhatsApp Announces Strict Account Settings for Protecting At-Risk Individuals Against Sophisticated Cyberattacks
Gadgets 360 is available in
Download Our Apps
Available in Hindi
© Copyright Red Pixels Ventures Limited 2026. All rights reserved.